Open research projects
The projects listed below are part of the upcoming call for applications, opening on August 4, 2026.
Abstract
Foundation models for gene regulation
Abstract
Genomes encode instructions for cells to regulate gene activity in response to their environment. However, despite its importance for biology, medicine, and biotechnology, the understanding of the underpinning regulatory code remains incomplete for any organism. We have secured a 10-million Euro grant (ERC Synergy) to derive the first comprehensive sequence-based model of eukaryotic gene regulation. EPIC combines innovative high-throughput technologies (Pelechano, Karolinska Institute) to probe gene expression regulation at an unprecedented scale across 100 species (with Benz, TUM) and multiple conditions with synthetic biology to massively test regulatory sequences (Verstrepen, VIB, Louvain). Deep learning on these data allows us to build predictive models and unravel complex regulatory instructions (Gagneur).
One PhD position is open. It focuses on unraveling the genetic instructions that determine mRNA boundary choice through evolutionary time scales. Ultimately, this code will enable the effective and accurate annotation of the genomes of hundreds of thousands of species.
The project will have 4 phases.
1. DNA language models [1,2] will be designed and trained to learn sequence representations upon which supervised learning models trained on public data will provide a first version of the model
2. EPIC will generate a massive dataset of multiomics assays ranging from chromatin accessibility to protein abundance via RNA boundaries, synthesis and decay and rates, for 100 species in multiple conditions. This “labelled data” will be used to train and improve supervised learning models, e.g. as in [3], improving step 1.
3. You will develop methods to design which sequences to test at high throughput to best improve models. EPIC will test nearly one million artificial genes.
4. EPIC core data focuses on fungi, species for which the throughput and length of the synthetic DNA sequences will allow us to achieve a complete code. However, we expect the modeling strategies to generalize to higher eukaryotes, including humans, which you will investigate.
You will work in collaboration with 3 other PhD students within the Gagneur lab.
Foundation models for gene regulation
Domain: Life Sciences
Supervisors: Julien Gagneur, TUM/Helmholtz Munich, J. Philipp Benz, TUM
Abstract
Beyond Goodhart: Non-Gameable Metrics for Trustworthy Biomedical AI — adversarially robust evaluation proxies for health question answering and AI co-scientists
Abstract
AI systems are increasingly used as autonomous agents in scientific research, proposing hypotheses and designing, running, and screening experiments with little human oversight. Most automated research loops reduce a task to a single metric that stands in for the real objective: a benchmark score, a reward model, or increasingly an LLM-based judge. By Goodhart’s law, once such a proxy becomes the optimization target, it stops measuring the property it was meant to track. The system learns to hill-climb the proxy rather than improve at the underlying task, a failure mode known as reward hacking. This matters in high-stakes biomedical settings such as answering public health questions or running an “AI co-scientist” that proposes and screens hypotheses: systems can score well while performing unreliably, because fluency, benchmark scores, and automated judgements diverge from correctness.
To make such systems reliable, metrics must be good proxies for the underlying objective and difficult to reward-hack. This project develops evaluation proxies that are as hard to game as possible. We first construct small, expensive, expert-annotated reference datasets that sit close to the real task and define a measure of how far a given metric is from the underlying question. We then test evaluation adversarially, applying methods developed to expose worst-case LLM failures (embedding-space, continuous, and reinforcement-style attacks) to inflate a proxy without improving the task, and train the proxies to resist those manipulations. We report the residual gameability of each proxy as an explicit metric.
The hardened proxies we receive can be used both to fine-tune models and to publish a benchmark and public leaderboard, such that training targets the real objective rather than the proxy. We validate on two testbeds: public health question answering and an AI co-scientist reasoning task, grounded in curated public-health corpora, clinical-trial text, and biomedical knowledge graphs.
The main contribution is methodological and transferable beyond biomedicine: we generate insights for the reliable design of agentic research loops: a framework for building, attacking, and validating evaluators. Biomedicine supplies the high-stakes testbeds and expert ground truth. The project trains a doctoral researcher in both adversarial machine learning and biomedical application.
Beyond Goodhart: Non-Gameable Metrics for Trustworthy Biomedical AI — adversarially robust evaluation proxies for health question answering and AI co-scientists
Domain: Medicine & Health
Supervisors: Leo Schwinn, Helmholtz Munich, Sebastian Lobentanzer, Helmholtz Munich
Abstract
AI Augmented Target Discovery in Translational Kidney Disease
Abstract
Single-cell genomics has revolutionized our ability to measure biological processes across human tissues and model systems. This technology has delivered cellular reference atlases that characterize human tissues in health and disease, and has enabled high-throughput perturbation screens that measure the cellular effects of hundreds of drug compounds in a single experiment. Together, these innovations have the potential to usher in a new paradigm for drug discovery, characterizing pathological cellular states in atlases and measuring which drugs can reverse these in perturbation screens. Kidney disease is a particularly enticing domain to pilot this innovation, where advanced organoid models are available and disease burden of chronic disease with limited treatments is high. Yet, open challenges around model system fidelity, translatability of perturbation responses to patients, modeling of high-throughput screens, and extension to low-cost highly-scalable enabling technologies remain.
This project is part of a wider effort on single-cell-based drug discovery in kidney disease. In this project, we will model and analyse drug perturbation effects in organoid models and map these to single-cell reference atlases of kidney disease in human donors and mouse models. We will extend existing reference mapping approaches to accommodate targeted sequencing technologies on single-cell and bulk level, and use these to enrich kidney disease atlases with perturbation data, enabling the modeling and prediction of perturbation effects towards identifying disease targeting drug effects. The project will involve close collaboration with kidney disease, organoid, computational biology, and perturbation screening experts from the industry side, as well as PhDs and Postdocs building reference atlases and perturbation modeling tools as part of this project. Overall, we aim to establish a comprehensive translational kidney disease atlas
that connects organoid and mouse model data with human disease biology, providing a shared resource for both experimental and computational researchers, guiding therapeutic discovery by predicting interventions that restore cellular health.
AI Augmented Target Discovery in Translational Kidney Disease
Domain: Medicine & Health
Supervisors: Fabian Theis, Helmholtz Munich/TUM
Abstract
Scalable Multimodal Imaging and AI for Phenotypic Drug Screening in Kidney Organoids
Abstract
Kidney organoids provide a scalable human model system for studying kidney disease and therapeutic perturbations. We propose to develop artificial intelligence models trained on longitudinal multimodal imaging datasets from kidney organoids exposed to genetic and compound perturbations. By integrating imaging with complementary molecular profiling from the broader collaboration, this project aims to predict how perturbations affect organoid health, maturation, toxicity, and disease-relevant phenotypes.
The project will establish computational methods for robust image analysis, quality control, longitudinal phenotyping, and representation learning. These models will be used to identify perturbations that shift disease-associated organoid phenotypes toward healthier states, prioritize candidates for deeper characterization, and support the development of improved disease models. The broader goal is to establish a scalable multimodal AI framework that connects organoid phenotypes with molecular response states for translational kidney disease research.
Scalable Multimodal Imaging and AI for Phenotypic Drug Screening in Kidney Organoids
Domain: Medicine & Health
Supervisors: Tingying Peng, Helmholtz Munich, Malte Lücken, Helmholtz Munich
Abstract
Secure Conversational AI Companions for Patient-Reported Outcome Capture, Multimodal Digital Twins, and Guideline-Concordant Clinical Decision Support
Abstract
Faced with a critical diagnosis and confined to a hospital bed, patients often turn to unsupervised sources such as generic chatbots for medical information and treatment advice. Yet although clinicians remain patients' primary and most trusted point of contact, time pressure often prevents them from eliciting the full picture of a patient's symptoms and concerns during brief consultations. This MUDS project develops a patient-facing conversational AI assistant for inpatients from the ground up, and progressively extends it into a closed-loop architecture connecting patients and clinicians.
The project proceeds in three stages. First, a safety-gated, German-language patient assistant will be built using retrieval-augmented generation (RAG) over certified patient education material and clinical guidelines. From the outset, it will be hardened both against adversarial manipulation (prompt injection, jailbreaks, data-poisoning and membership-inference risks on sensitive clinical data) and against reliability failures under benign input variation, such as distribution shifts, misspellings, paraphrased or vaguely framed questions, and out-of-scope
requests. Second, the assistant will be extended with robust pipelines to extract structured patient-reported outcomes from open conversation, enabling their use in downstream clinical risk models. Third, these patient-reported outcome signals will be fused, together with laboratory, imaging, wearable and EHR data, into a multimodal patient digital twin.
In parallel, a clinician-facing counterpart will aggregate and visualise the captured patient-reported outcome signals for the treating physician and retrieve guideline- and literature-grounded treatment suggestions (e.g., Onkopedia for hematology, ADA guidelines for diabetes), with hard safety constraints and mandatory human review before any suggestion reaches a clinician. The resulting end-to-end pipeline (secure patient assistant → structured PRO/digital twin → clinician assistant → guideline-grounded suggestion) will be developed and validated across two clinical domains, hematology inpatient care and diabetes management, testing the generalisability of the architecture beyond a single disease area.
The project is embedded in the M1 Munich Medicine Alliance between LMU and TUM (together with their university hospitals and Helmholtz Munich), leveraging shared clinical infrastructure, data governance frameworks and the Bavarian health-data cloud initiative for multi-site evaluation.
Methodologically, the project contributes new approaches at the intersection of adversarial robustness for medical LLMs, multimodal representation learning for digital twins, and retrieval-augmented, guideline-grounded clinical reasoning — advancing safe, trustworthy conversational AI as a bridge between patients and clinicians in routine care.
Secure Conversational AI Companions for Patient-Reported Outcome Capture, Multimodal Digital Twins, and Guideline-Concordant Clinical Decision Support
Domain: Life Sciences
Supervisors: Leo Schwinn, Helmholtz Munich/TUM, Carsten Marr, Helmholtz Munich/LMU Clinic
Abstract
AI-Driven Therapeutic Design Using Generative Models and Optimization
Abstract
The discovery pipeline for new therapeutics remains slow and costly, and antibiotic resistance is projected to become one of the leading causes of death worldwide. Generative AI models have recently shown promise for de novo therapeutic design. However, such models typically explore chemical space without efficiently optimizing for the multiple, often competing properties, such as activity, toxicity, synthesizability, required for a viable drug candidate, and they rely on sample inefficient strategies.
This project addresses that gap at the level of the generative modeling problem itself, by developing an optimization guided generative framework that couples generation with efficient, property driven sampling for therapeutic design. This general framework will be applied across two different therapeutic modalities: small molecules and peptides. For small molecules, the project incorporates a complementary target identification pipeline developed by the Sieber lab, which combines structure based methods with the discovery of novel small molecule inhibitors against Gram negative pathogens. Peptide design will target the generation of therapeutic peptides that bind specific receptors, supporting noncanonical amino acids, and will be applied initially to targeted drug delivery. One example application will be developed in collaboration with the Ertürk lab, aiming at peptides targeting T cells. Future extensions to additional therapeutically relevant targets such as GLP-1 receptor family members are also planned.
By combining AI expertise from the Szczurek lab with chemistry expertise from the Sieber lab, this project builds on established wet lab collaborations for experimental validation, with the Sieber lab validating small molecules and the Ertürk lab validating peptide binders, to apply optimization guided generative AI methods within a closed loop therapeutic design pipeline.
AI-Driven Therapeutic Design Using Generative Models and Optimization
Domain: Medicine & Health
Supervisors: Ewa Szczurek, Helmholtz Munich, Stephan A. Sieber, TUM
Abstract
Tri-Modal Multi-Omics Resolves Immune Deviations in Subjects with Mitochondrial Diseases
Abstract
Mitochondria are increasingly recognized as key regulators of immune function, yet we still understand little about how mitochondrial dysfunction reshapes the human immune system. This PhD project will address that question using a cutting-edge tri-modal single-cell multi-omics approach that integrates gene expression, nuclear chromatin accessibility, and mitochondrial chromatin accessibility in immune cells from patients with mitochondrial disease. The student will join an international collaboration between Charité (Germany) and INSERM Bordeaux (France) and work at the interface of computational biology, immunology, and translational medicine. Contributing to our work in scvi-tools and the scverse ecosystem, the project will develop new generative modeling methods for the joint analysis of RNA, ATAC, and mitochondrial features, with the goal of identifying clinically relevant immune endotypes. The candidate will establish lineage tracing in human immune cells using mitochondrial information captured directly in single-cell assays. The candidate will work on a unique and deeply characterized patient cohort. Molecular profiles generated in this project will be linked with clinical phenotypes, vaccine response, and metabolic measurements from the consortium. This provides a strong opportunity to connect computational method development with clinically relevant questions in rare disease biology. This PhD project offers the opportunity to develop reusable computational tools, contribute to open-source software, analyze a unique patient cohort, and help define how mitochondrial dysfunction contributes to immune dysregulation. The results are expected to support precision medicine approaches in mitochondrial disease and contribute to the growing field of immunometabolism.
Tri-Modal Multi-Omics Resolves Immune Deviations in Subjects with Mitochondrial Diseases
Domain: Life Sciences
Supervisors: Can Ergen, Helmholtz Munich, Emmanuel Saliba, HIRI
Abstract
Statistics and data science for large-scale chemical genomics screens
Abstract
The advent of high-throughput plate-reader and automation platforms has recently enabled chemical genomics screens at unprecedented scale, generating thousands of bacterial time series per experiment, including optical-density growth curves and luminescence readouts of promoter activity in genetically engineered strains. DGrowthR, an R package and desktop application recently developed by the Müller group and collaborators, already provides an end-to-end framework for preprocessing, exploratory functional data analysis, Gaussian-process modeling, growth-parameter extraction, and differential-growth testing based on Bayes factors and permutation p-values. This doctoral project will extend these statistical workflows into a broader statistical platform for dynamic microbial phenotyping. First, the student will add spline-based curve fitting, including penalized and shape-aware spline models, as a transparent and computationally efficient complement to Gaussian processes. Second, the project will generalize the DGrowthR workflow from optical-density growth curves to promoter luminescence curves, where engineered bacterial strains report transcriptional activity over time. These data raise new statistical questions because signal timing, amplitude, burst-like activation, and growth-normalized promoter output must be modeled jointly and compared across genetic and environmental perturbations. Third, the student will investigate whether permutation p-values derived from Bayes-factor statistics can be replaced or complemented by calibrated alternatives based on asymptotic theory, parametric approximations, score/Wald-type statistics, or resampling schemes that preserve experimental structure. The outcome will be a rigorously validated open-source R software extension, benchmarked on existing DGrowthR growth datasets and new growth-plus-luminescence experiments from engineered bacterial strains. The project sits at the interface of biostatistics, functional data analysis, microbial systems biology, and reproducible scientific software, and will train a doctoral researcher to translate modern statistical methodology into practical tools for life-science discovery.
Statistics and data science for large-scale chemical genomics screens
Domain: Life Sciences
Supervisors: Göran Kauermann, LMU, Christian L. Müller Helmholtz Munich
Abstract
Data-driven translation of a single-cell multi-omics biomarker into a minimal flowcytometry panel for non invasive diagnosis of coronary artery disease
Abstract
Coronary artery disease (CAD) is a leading cause of death worldwide, yet its diagnosis still relies on invasive catheter-based angiography or CT imaging, both of which expose patients to radiation, contrast agents and procedural risk. In prior work, we identified a cell-type-resolved mRNA biomarker that predicts CAD status with high accuracy, matching the diagnostic performance of CT angiography. However, mRNA based biomarkers are not fit for routine clinical use. This doctoral project will translate this biomarker from the transcript to the protein level, making it deployable on standard flow cytometry platforms available in any clinical laboratory. The doctoral researcher will use feature selection methods to develop a minimal panel of antibody markers from thousands of candidate transcripts, explicitly modelling the uncertainty between transcript and protein measurements using large reference CITE-seq (matched protein + RNA) datasets and dedicated patient CITE-seq pilot data. The resulting panel will be tested and refined in a pilot cohort and then validated in a clinical cohort of 300 patients undergoing CT or coronary angiography at LMU University Hospital. This project combines method development in computational biology and machine learning with translational cardiovascular medicine, and is jointly supervised by a computational biology PI (Matthias Heinig) and a clinical cardiology PI (Konstantin Stark). Beyond its clinical value as a cost-effective, non-invasive triage test for CAD, the project will deliver a generalizable computational framework for translating high-dimensional single-cell/omics biomarkers into lowdimensional, clinically deployable assays - a methodological contribution of interest well beyond cardiology.
Data-driven translation of a single-cell multi-omics biomarker into a minimal flowcytometry panel for non invasive diagnosis of coronary artery disease
Domain: Medicine & Health
Supervisors: Matthias Heinig, Helmholtz Munich, Konstantin Stark, LMU Universtity Hospital
