Postdoctoral Fellow
Listed on 2026-09-12
-
Research/Development
Research Scientist, Data Scientist, Biomedical Science
Are you interested in studying metagenomic derived protein families and developing methods to interrogate vast collections of proteins, to determine pockets of interesting novel families? Metagenomics is transforming our understanding of the microbial world by uncovering enormous numbers of novel proteins. MGnify, one of the largest metagenomics resources, contains ~6 billion unique protein sequences. At the same time, artificial intelligence (AI) methods like Alpha Fold and ESMfold can accurately predict protein structures directly from sequence.
The Alpha Fold database (AFDB) contains models for >200 million Uni Prot proteins, while the ESM Atlas comprises >600 million models for MGnify proteins, with many more expected soon. A new joint initiative with Christine Orengo's team based at University College London, we will co-develop scalable strategies for classifying the MGnify protein databases into protein families based on structure and investigate the distribution of novel protein families across taxa and environments.
group
The Finn Research Group currently comprises two PhD students and two post doctoral fellows. This team is closely aligned with the Microbiome Informatics team responsible for producing MGnify, the AMR portal and the microbial data in Ensembl. The Finn Research covers a range of different research themes, from computational tool development to deep dives into data driven research topics such as exploring the human skin and gut microbiomes.
The tool development takes on a number of different forms, from algorithmic development to the application of emerging AI technologies. These tools are typically designed to work at scale, with a view that many of these will be utilised in the data resources produced by the Microbiome Informatics team.
You will report directly to Research Group lead, Rob Finn.
Your roleYou will co-develop strategies that can deal with the classification of the vast protein database provided by MGnify into protein families based on structure. This will include clustering, functional labelling and developing selection criteria for producing structures. You will help update the current MGnify database with taxonomic information, based on a range of sources from within MGnify. You will also expand the biome information, based on an emerging tool produced within the wider team.
Using these pieces of information, you will conduct an investigation looking for correlations between biome and taxonomy and protein family distributions, relating this to functions. Expanding on prior research, you will undertake a specific task aimed at trying to identify bacteriophage encoded bacterial anti-defence systems and use the functional and structural classification to propose potential mode of actions. The research will undertake both methodological approaches (including the adopting of AI-based approaches) as well as data analysis at scale.
have
- PhD in the biological sciences, computational biology, bioinformatics, computer science or a related field, and proven research experience in a relevant field.
- A strong background and understanding in microbiology and/or metagenomics.
- Understanding of protein classification approaches and the tools that underpin them.
- Research experience dealing with large datasets.
- An eagerness to work in a highly collaborative atmosphere while still being able to work independently and to meet deadlines in a timely manner.
- Strong Python skills, with demonstrable ability to produce well documented and tested software.
- Experience with UNIX/Linux, and HPC or cloud environments
- Motivation to work in an international team on interdisciplinary projects.
- Strong communication and interpersonal skills, with the ability to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).