M.Sc. Bioinformatics — Saarland University
Coursework: machine learning, neural networks, genomic data analysis & NGS, single-cell bioinformatics, microbiome data analysis.
Bioinformatics Master’s student, Saarland University
Saarbrücken, Germany
I am a Master’s student in Bioinformatics at Saarland University, currently working on my thesis: a genomic language model that predicts antibiotic resistance directly from bacterial DNA, built with an industry partner. Before switching into bioinformatics I spent three years as a backend developer, which is probably why I care as much about whether a pipeline actually runs end-to-end as I do about the model behind it. I’m currently looking for part-time machine learning work alongside my thesis.
Coursework: machine learning, neural networks, genomic data analysis & NGS, single-cell bioinformatics, microbiome data analysis.
My thesis is on predicting antimicrobial resistance from genotype data — instead of running a lab test to see which antibiotics a bacterial strain resists, the goal is to predict it from its genome. I’m building a transformer-based model for this with Databiomix, co-supervised by their CEO and by a professor at Saarland University, working from curated resistance data benchmarked against real lab results.
I led a 4-person team building a model to predict which phages (viruses that infect bacteria) would work against which bacterial strains — useful for phage therapy as an alternative to antibiotics. We used graph neural networks and self-supervised learning to make the most of limited labeled data, and I built the Snakemake pipeline that took raw data through to evaluation. We placed 3rd at Startup Weekend Saarbrücken 2024.
I worked on predictive models for infertility risk and disease patterns (TB, hepatitis, polio) in Afghanistan, using public health data to help identify where healthcare resources were most needed.
I built backend systems and REST APIs in Python/Django for an e-commerce platform, and helped migrate it from a multi-page to single-page architecture. I also pulled sales and product insights out of SQL data using Redash and did the database design work.
I built StockOCR, an internal inventory tool that used Tesseract OCR to read serial numbers and MAC addresses straight off equipment photos, plus the check-in/check-out workflow around it.
An ML pipeline I built using K-means clustering and logistic regression to automate candidate profiling for a recruitment scenario, with asynchronous data pulled from GitHub and Stack Overflow and real-time feedback via code-evaluation checks.
Pipeline diagrams for the two projects above, in more architectural detail.
Python, SQL, R, Bash
PyTorch, scikit-learn, graph neural networks, self-supervised learning, statistical modeling
Biopython, Snakemake, genomic/NGS data analysis
Git, Docker, Django, Linux
I’m open to part-time ML roles alongside my thesis, and to conversations about genomic ML, AMR, or phage therapy.