Universitat Internacional de Catalunya
Basic Computational Skills for Bioinformatics
Other languages of instruction: Catalan, Spanish
Teaching staff
Questions may be discussed in person or by videoconference with the course coordinator or the corresponding instructor. Students should contact the instructor in advance to arrange an appointment.
Introduction
This course introduces essential computational skills for bioinformatics and the practical use of biomedical databases. Students learn to work in command-line environments, organise biological data, write basic scripts in Python and Bash, and develop clear and reproducible computational workflows.
The course also provides practical experience in retrieving, evaluating and connecting information from major biomedical databases. Students work with nucleotide, gene, expression, protein, genetic variation, disease and drug-related resources and apply them to realistic biomedical questions.
Guided exercises and case-based activities help students translate biological questions into computational tasks, select appropriate data sources, interpret results and document their analytical process.
AI-assisted programming tools may be used to support code generation, explanation and debugging. Particular emphasis is placed on defining the analytical objective, critically reviewing generated code, testing its behaviour, validating the results and documenting the final workflow.
The course contributes primarily to SDG 3 (Good Health and Well-being), SDG 4 (Quality Education) and SDG 9 (Industry, Innovation and Infrastructure) by strengthening computational literacy, reproducible research practices and the responsible use of digital tools in biomedical research.
Pre-course requirements
Students are recommended to have previously completed the courses Introduction to Bioinformatics and Biomolecular Interactions.
No previous programming experience is required. However, students should be familiar with basic computer use, including working with files, folders and standard web-based applications.
Objectives
- Provide students with the foundational computational knowledge required to work with biological and biomedical data using command-line tools and basic scripting.
- Develop introductory programming skills in Python and Bash through guided practical exercises that emphasise readable, modular, well-documented and reproducible code.
- Familiarise students with major biomedical databases and help them understand their organisation, scope, limitations and appropriate use in biomedical research.
- Strengthen critical thinking and problem-solving skills by guiding students to design, document, verify and evaluate computational workflows and their outputs.
- Promote responsible and collaborative computational practices, including the critical use of AI-assisted programming tools, and demonstrate how computational approaches can connect molecular data with clinical and translational research.
Competences/Learning outcomes of the degree programme
- CN14 - Identify the principles of biomedical sciences related to health, as well as the basic concepts and tools that have an impact on Biomedical Sciences and allow them to work in any of its fields (biomedical companies, bioinformatics labs, research laboratories, clinical analysis companies, etc.).
- CP05 - Apply biological foundations in the search for practical solutions to health problems, following ethical standards and scientific rigour and respecting fundamental equal rights between men and women, and the promotion of human rights and the values inherent in a peaceful society of democratic values that includes inclusive, non-discriminatory language without stereotypes.
Learning outcomes of the subject
By the end of the course, students will be able to:
- Use command-line tools to work in computational environments, manage and process biological data, and document the individual steps of a reproducible workflow.
- Write, adapt, debug and test basic scripts in Python and Bash for clearly defined data-processing tasks, including the critical review of AI-generated code.
- Select appropriate biomedical databases and retrieve data using web interfaces, common file formats and basic API queries.
- Interpret and combine information from nucleotide, gene-expression, protein, genetic-variation, disease and pharmacological resources to answer defined biomedical questions.
- Organise a small bioinformatics project using clear file structures, version control, appropriate documentation, and principles of collaboration and reproducibility.
- Evaluate the provenance, quality, limitations and consistency of biomedical data and computational outputs, and communicate conclusions within the relevant biomedical context.
Syllabus
Short description:
The syllabus is organised into two complementary parts. Part I develops foundational skills in command-line use, Python and Bash programming, biological file handling, APIs, version control, reproducibility and the critical use of AI-assisted programming tools. Part II applies these skills to the exploration and evaluation of major biomedical databases through guided exercises and case-based activities.
CHAPTER 1: PROGRAMMING AND COMPUTATIONAL FOUNDATIONS (15 HOURS)
1.1 Computational environments and the command line (3 hours)
-
Navigating file systems and managing files and directories
-
Permissions and access control
-
Pipes, command chaining and text-processing tools: cat, grep, awk and sed
-
Working with local and cloud-based data environments
1.2 Programming fundamentals in Python and Bash (3 hours)
-
Basic Python syntax and data types
-
Variables, conditions and loops
-
Lists, dictionaries and basic shell scripts
-
Breaking down biomedical questions into computational tasks
1.3 Biomedical file formats, data processing and APIs (3 hours)
-
Common biomedical and data formats: FASTA, GTF/GFF, VCF, CSV/TSV and JSON
-
Reading, writing and transforming tabular data
-
Introduction to APIs and remote data access
-
Retrieving data through basic API queries
1.4 Modular and reliable code (3 hours)
-
Defining and using functions
-
Organising code into modules
-
Naming, commenting and documenting code
-
Debugging and error handling
-
Basic testing and validation of results
1.5 Reproducible and AI-assisted workflows (3 hours)
-
Project organisation and documentation
-
Introduction to Git and version control
-
Critical use of AI-assisted tools for code generation, explanation and debugging
-
Testing, validating and documenting AI-generated code
-
Group mini-project: development of a reproducible data workflow
CHAPTER 2: BIOMEDICAL DATABASES AND DATA SOURCES (15 HOURS)
2.1 Understanding biomedical databases (3 hours)
-
Main types of biomedical data
-
Database organisation, identifiers and cross-references
-
Data curation, provenance, versioning, quality and limitations
-
Introduction to the major portals: NCBI, EMBL-EBI, Ensembl and UniProt
2.2 From nucleotides to genes and genomes (3 hours)
-
Nucleotide and sequence resources: GenBank, RefSeq and ENA
-
Gene resources: NCBI Gene, Ensembl and GeneCards
-
Genome exploration using the UCSC Genome Browser
-
Gene nomenclature using HGNC
-
Sequence retrieval and cross-referencing between databases
2.3 Gene expression and regulation (3 hours)
-
Gene-expression resources: GEO and Expression Atlas
-
Single-cell data using the Single Cell Expression Atlas
-
Regulatory and tissue-expression resources: GTEx and ENCODE
-
Interpretation of expression profiles and their biological context
2.4 Protein sequences, functions and interactions (3 hours)
-
Protein sequence resources: UniProt and RefSeq
-
Functional annotation: InterPro, Pfam and PROSITE
-
Protein–protein interaction resources: STRING, IntAct and BioGRID
-
Critical interpretation of predicted and experimentally supported interactions
2.5 Genetic variation, diseases and pharmacological resources (3 hours)
-
Genetic-variation resources: ClinVar, gnomAD and dbSNP
-
Disease-related resources: OMIM and Orphanet
-
Pharmacogenomic and drug resources: ClinPGx, DrugBank and PubChem
-
Integration of molecular, disease and pharmacological information in a biomedical case
-
Evaluation of evidence, data quality and database limitations
Teaching and learning activities
In person
In person
Fully in-person classroom modality
Lectures – 18 hours: Interactive classroom sessions in which the instructor introduces core concepts and demonstrates computational methods, tools and databases using worked examples. The sessions may include short guided exercises to reinforce the practical application of the content.
Case Method (CM) – 12 hours: Working in groups, students address practical problems based on biomedical and bioinformatics cases using the data, scripts and databases introduced during the course. Each group identifies appropriate resources, develops a documented workflow, interprets the results, and presents and justifies its conclusions. The instructor guides the discussion, provides feedback and introduces additional concepts when necessary.
Evaluation systems and criteria
In person
Fully in-person classroom modality
Students in the first examination period:
-
Midterm multiple-choice examination: 30%
-
Final multiple-choice examination: 40%
-
Case Method activities: 30%
The examinations may assess conceptual understanding, code interpretation, identification of errors, selection of appropriate file formats and databases, workflow design, and interpretation of computational and biomedical results.
Students in the second or subsequent examination periods:
The grade obtained for the Case Method activities will be retained. The final examination will account for 70% of the final grade.
Students repeating the course who wish to retake the midterm examination in the third or fifth examination period may do so, provided that they notify the course coordinator in advance.
General evaluation criteria:
-
A minimum grade of 5 out of 10 must be obtained in the final examination for the weighted average to be calculated. The minimum overall grade required to pass the course is 5 out of 10.
-
Multiple-choice questions will have four possible answers. One point will be awarded for each correct answer, 0.33 points will be deducted for each incorrect answer, and unanswered questions will receive zero points.
-
Case Method activities will be evaluated according to the suitability of the selected tools and data sources, the correctness and reproducibility of the workflow, the quality of the analysis and critical interpretation, the clarity of the documentation and presentation, and the student’s contribution to the group work.
-
Due to the continuous-assessment nature of the course, students must attend at least 75% of the the total scheduled classroom hours.
-
Plagiarism, the unauthorised sharing of code or results, and the undeclared or unauthorised use of AI tools in assessed activities may be treated as academic misconduct. Any use of AI-assisted tools must follow the instructions established for the activity and must be declared when required.
-
Electronic devices may only be used for educational purposes during class. Recording or sharing images, audio or video of students or instructors without permission is not allowed and may result in removal from the session and the application of the relevant university regulations.
Bibliography and resources
-
Python Software Foundation. The Python Tutorial – Python 3 Documentation. https://docs.python.org/3/tutorial/
-
Python Software Foundation. The Python Standard Library. https://docs.python.org/3/library/
-
Matthes, E. (2022). Python Crash Course: A Hands-On, Project-Based Introduction to Programming (3rd ed.). No Starch Press. https://nostarch.com/python-crash-course-3rd-edition
-
Software Carpentry. Programming with Python. https://swcarpentry.github.io/python-novice-inflammation/
-
GNU Project. Bash Reference Manual. https://www.gnu.org/software/bash/manual/
-
Chacon, S., & Straub, B. (2014). Pro Git (2nd ed.). Apress. https://git-scm.com/book/en/v2
-
Biopython Project. Biopython Tutorial and Cookbook. https://biopython.org/wiki/Documentation
-
NCBI. Tutorials and Educational Resources. https://www.ncbi.nlm.nih.gov/home/tutorials/
-
EMBL-EBI Training. Introductory Bioinformatics Pathway. https://www.ebi.ac.uk/training/online/courses/introductory-bioinformatics-pathway
-
Ensembl. Tutorials and Worked Examples. https://www.ensembl.org/info/website/tutorials/
-
UniProt. Help and Documentation. https://www.uniprot.org/help/
-
Buffalo, V. (2015). Bioinformatics Data Skills: Reproducible and Robust Research with Open Source Tools. O’Reilly Media.
Additional documentation and tutorials for the databases covered during the course will be provided through Moodle.