Project case study
Genome curation
Predicting neighbouring genome fragments and the ends that connect them, using contact images and graph learning.
This project is part of my collaboration with Tree of Life, through my work at the Wellcome Sanger Institute.
From contact evidence to compatible paths
Selected real-data examples, not benchmarks. Eight displayed fragments are an excerpt of a larger model computation.
1. Contact evidence
2. Candidate graph
3. Learned join scores
4. Compatible paths
On mobile, swipe diagrams or focus with Tab and use the arrow keys.
Recorded validation
- Join F1
- 0.631
- Validation samples
- 98
- Threshold
- 0.5
Candidate scoring, not assembly accuracy. Main validation force-includes eligible reference joins. This single-seed, validation-selected result is not independent test performance. Retrieval-subset scores use augmented-graph predictions and are not independently verified inference performance.
My contribution
Data preparation → image features → graph models → training and evaluation → iterative experiments → stakeholder feedback.
Experiments were tracked in MLflow and compared run by run to guide model and data changes. Progress was presented regularly to curators from Darwin Tree of Life and the Vertebrate Genomes Project, bioinformaticians and AI engineers at Rockefeller University, and the Google AI genomics team, and their feedback shaped the next steps.
Engineering details
- Traceable data: canonical pair IDs, versioned records, atomic writes, and recorded exclusions.
- Model design: five ordered image views, graph message passing, and pair/multi-view baselines. Pretrained ViT-MAE, Hugging Face Transformers, PyTorch, and PyTorch Geometric remain third-party components.
- Evaluation: train-only normalisation; separate join and conditional-end metrics; stop on out-of-memory errors rather than silently dropping candidates.