Data Science & Analytics · Machine Learning · MSc Bioinformatics
Turning biological and scientific data into models, pipelines and tools people can actually use.
I'm from Ponte de Lima and studied at the University of Minho, where an MSc in Bioinformatics was, in practice, a training in data: programming, machine learning, statistics and databases, applied to problems where the dataset is large, noisy and rarely behaves the way the textbook says.
What I enjoy is the full path from raw data to something usable — exploring and cleaning the data, choosing and validating a model, and then building the interface or pipeline that puts the result in someone's hands. Most of my work sits at the intersection of data science and software engineering: not just a notebook with an accuracy score, but code that is documented, tested and reproducible.
The project that taught me the most was extending KEGGCharter, a published open-source tool used by other researchers. Working inside someone else's codebase, on software people depend on, forces a discipline that writing from scratch never does.
Python · Biopython · data visualisation · open source
KEGGCharter is a published open-source tool that maps genomic and transcriptomic data onto KEGG metabolic pathways. Its output was static images: accurate, but impossible to explore — taxonomy and gene expression were drawn as separate layers, so you could not read microbial identity and expression in a single view.
I extended the tool with an interactive charting interface, adding multi-level taxonomic representation, on-click access to the underlying numeric values, and cross-references to external databases. Developed as an MSc project in collaboration with the Centre of Biological Engineering (CEB, University of Minho), under the guidance of the tool's author.
→ Project write-up: Mapping Omics datasets on KEGG Metabolic Pathways
Python · scikit-learn · clustering · dimensionality reduction
End-to-end machine learning analysis on the GDSC1 dataset (IC50 responses for 208 drugs across ~1000 cancer cell lines). Exploratory analysis, feature engineering and missing-value treatment, followed by unsupervised learning (dimensionality reduction and clustering) and supervised models predicting drug response from a gene-expression profile and a compound's SMILES representation.
R · Bioconductor · differential expression · machine learning
Analysis of gene expression data from the GDC Data Portal: preprocessing, descriptive statistics and visualisation, univariate analysis and differential expression, then clustering, dimensionality reduction, predictive modelling with model comparison, and gene selection by importance. Reports generated with R Markdown.
Python · NumPy · pandas · algorithm implementation
Implementation of core machine learning algorithms from first principles using only NumPy and pandas, following a common scikit-learn-style API. Built to understand what happens under .fit() rather than to call it.
MySQL · MongoDB · Neo4j · data modelling
Migration of a hospital management relational database into two non-relational paradigms — document-oriented (MongoDB) and graph-based (Neo4j) — including schema redesign for each paradigm, query implementation and a critical comparison of performance and modelling trade-offs against the original relational system.
Python · unit testing · type hinting · documentation
Original implementations of motif discovery (Gibbs sampling), evolutionary computation, pattern matching (finite automata, tries, suffix trees), Burrows-Wheeler alignment, graphs and biological networks. The repository is deliberately written as production-style code: documented, type-hinted and covered by unit tests.
More repositories
- Flux Balance Analysis — constraint-based metabolic modelling of Chlamydomonas reinhardtii with COBRApy and MEWpy, predicting metabolic flux under different environmental and genetic conditions.
- Linear regression in R — statistical modelling and diagnostics.
- Bioinformatics laboratories — systematic gene analysis pipelines with Biopython.
- Web scraping + MySQL — data collection and database population in Python.
- Databases and NoSQL exercises — SQL and NoSQL practice.
MSc in Bioinformatics — University of Minho, School of Engineering · 2023–2026
Programming, machine learning, statistics, databases and systems biology.
Dissertation: metabolic prediction in the gut microbiome using rule-based pipelines (RetroRules), supervised by Prof. Miguel Rocha.
BSc in Applied Biology — University of Minho, School of Sciences · 2020–2023
Final project at CBMA: microplastic contamination in water bodies using an artificial intelligence approach — data processing and analysis in Python, with statistical analysis in GraphPad Prism and IBM SPSS.
Also: elected course representative for the MSc cohort (2023–2025) and part of the organising committee of the Bioinformatics Open Days for two editions.
📄 Certificates — training, workshops and event participation.
A first professional role working with data — data analytics, data science, data engineering or machine learning engineering.
📍 Porto · Braga · Aveiro · Guimarães · Lisboa, or remote within Europe
🗓️ Available immediately