Universitat Internacional de Catalunya

Big Data and Artificial Intelligence

Big Data and Artificial Intelligence
3
15815
1
First semester
OB
Main language of instruction: Catalan

Other languages of instruction: English, Spanish,

Introduction

The Big Data and Artificial Intelligence subject is designed to provide students with a
comprehensive understanding of two complementary pillars of modern data-driven
development: machine learning and large-scale data engineering.
In the Artificial Intelligence component, students will build a solid theoretical and
practical foundation in predictive modelling. Starting from linear models and statistical
learning theory, the course progresses through Support Vector Machines, K-Nearest
Neighbours, Decision Trees, Random Forests, and Neural Networks, with a dedicated
module on model interpretability to ensure that predictions can be understood, audited,
and trusted in real-world deployments.
In the Big Data component, students will explore the tools and processes used to
manage and process large volumes of data in distributed environments. Through the
study of technologies such as Hadoop, Spark, and Kafka, they will learn how to build
scalable, resilient systems capable of handling massive datasets with high availability
and fault tolerance. The course will also cover ETL (Extract, Transform, Load)
processes, efficient data storage, and real-time processing using tools such as Apache
Flink and Spark Streaming. Finally, it will delve into key concepts such as data security
and governance, with a particular focus on regulatory frameworks like GDPR and
CCPA, offering a holistic view of the challenges and solutions within the Big Data
ecosystem.

Objectives

Build a theoretical foundation in predictive modelling: Understand the statistical
learning framework, bias–variance trade-off, and model-selection principles that
underpin all supervised and unsupervised algorithms.
 Master core ML algorithms: Gain working knowledge of linear models, SVMs,
KNN, Decision Trees, Random Forests, and Neural Networks, including their
assumptions, strengths, and limitations.
 Develop interpretability skills: Learn to explain and audit model predictions
using techniques such as feature importance and SHAP values, so that ML
systems can be deployed responsibly.
 Understand Big Data architecture: Learn the core technologies (Hadoop, Spark,
Kafka) and how to build scalable, fault-tolerant distributed systems.
 Master data processing techniques: Gain expertise in ETL processes,
MapReduce, Apache Spark, and real-time data processing with Spark
Streaming and Apache Flink.
 Explore Big Data storage solutions: Study efficient storage architectures that
support fast access to large datasets.

 Implement data security and governance: Understand data privacy and
compliance with regulations like GDPR and CCPA.
 Apply practical knowledge: Design and implement real-world Big Data solutions
that are scalable, secure, and efficient.

Competences/Learning outcomes of the degree programme

  • CN01 - Describe the advanced aspects of bioengineering related to human health based on specific books on the subject along with scientific publications at the frontier of knowledge.
  • CN04 - Compare the different approaches to computing and data analysis in biomedical signal management.
  • CN05 - Explain the principles of bioengineering used in the design and manufacturing of new personalized medicine therapies.
  • CN06 - Identify the necessary steps and developments for the correct processing of medical devices based on available medical data.
  • CP01 - Interpret scientific results (both theoretical and practical) to make judgments with critical reflection, taking into account social, scientific, or ethical concepts.
  • CP02 - Work proactively in a multidisciplinary team, being aware of the role to be developed.
  • CP06 - Produce technologies and products applying the unique principles of bioengineering for each specific healthcare use, respecting fundamental rights of equality between men and women, and promoting human rights and values of peace and democratic culture; using language that avoids androcentrism and stereotypes.
  • HB01 - Differentiate new methods and theories in bioengineering to enhance versatility, adaptability to new situations, and problem-solving through critical reasoning.
  • HB02 - Validate results, calculations, studies, reports, work plans, and other similar works obtained through scientific experimentation
  • HB03 - Evaluate the social and environmental impact of technical solutions through the analysis and application of quality principles and methods.
  • HB04 - Apply bioengineering terminology in a multilingual and multidisciplinary environment, with an adequate oral and written level of English.
  • HB06 - Relate social problems to health issues, using a balanced and compatible approach of technique, technology, economy, and sustainability.
  • HB07 - Utilize data and information processing in the bioengineering field for a later critical assessment of the results.
  • HB08 - Solve problems arising in bioengineering in the health field through the application of multidisciplinary concepts.
  • HB09 - Examine the influence of bioengineering-related topics on the specific needs or characteristics of genders: biological and medical aspects.
  • HB10 - Apply knowledge in the use of advanced computational methods to improve the quality of the healthcare system.
  • HB12 - Discriminate relevant information from different sources, books, and/or scientific articles related to bioengineering.

Syllabus

1. Introduction to Machine Learning
2. Linear Models
3. Statistical Learning
4. Support Vector Machines
5. K-Nearest Neighbours
6. Decision Trees and Random Forest
7. Model Interpretability
8. Neural Networks
9. Big Data architectures and ecosystems (Hadoop, Spark, Kafka)
10. Scalability and fault tolerance in distributed systems
11. Extract, transform, and load (ETL) processes
12. Data storage in Big Data
13. Massive data processing (MapReduce and Apache Spark)
14. In-memory processing and RDD (Resilient Distributed Datasets)
15. Real-time data processing (Spark Streaming and Apache Flink)
16. Data security and governance (GDPR and CCPA regulations)

Teaching and learning activities

In person



The subject will be divided into 1) interactive theoretical classes, and 2) practical labs
and hands-on exercises.
● Study materials: Access to pre-prepared content (documents, videos, and
recommended readings) via Moodle, guiding independent learning.
● Autonomous work: Students are responsible for learning by studying the
materials, seeking additional information, and completing practical activities.
● Tutorials: Available online or in person for doubt resolution.

Evaluation systems and criteria

In person



1 st Call:

• Theoretical Knowledge (Interactive Classes) (40%)
• Practical Labs and Hands-on Exercises (60%)

2 nd Call: Only the note of the 2 nd call exam will be taken into account. In addition, there
will be no option for honors distinction.
● A minimum score of 5.0 out of 10 is required to pass.
● If the mid term exam is not passed. The final exam will count for 70% of the
final mark.
Other Calls
Final grade (2nd–6th attempts): 100% final or recovery exam, requiring a minimum of
5.0 out of 10 to pass.
RECOVERY EXAM
● Same format as the “FINAL EXAM.”

Bibliography and resources

● “An Introduction to Statistical Learning” – James, Witten, Hastie, Tibshirani,
Taylor, Springer, 2023.
● “The elements of statistical learning” – Hastie, Tibshirani, Friedman, Springer,
2017.
● "Big Data: Principles and Best Practices of Scalable Real-Time Data Systems"
– Nathan Marz, James Warren, Manning Publications, 2015.
● "Designing Data-Intensive Applications" – Martin Kleppmann, O'Reilly Media,
2017.

● "Fundamentals of Data Engineering" – Joe Reis, Matt Housley, O'Reilly Media,
2022.
● "Streaming Systems" – Tyler Akidau, Slava Chernyak, Reuven Lax, O'Reilly
Media, 2018.