Software Engineer; HPC Sustainability
Listed on 2026-09-20
-
Software Development
Backend Developer, Python
About the team
EMBL-EBI's HPC infrastructure underpins world-leading bioinformatics research, operating large-scale compute clusters on behalf of scientific users across the institute. The HPC sub team sits within IT and Technical Services (ITS) and is responsible for the reliability, operation, and ongoing development of these clusters and their supporting tooling.
Your role Your role is embedded within that HPC sub team, reporting to the HPC Lead. You will combine hands-on backend development with a meaningful sustainability mission: improving how EBI measures, reports, and reduces the environmental impact of its scientific computing. A significant part of your work involves a close collaboration with the Green Algorithms team at the University of Cambridge - an internationally recognised open source project for quantifying the carbon footprint of research computing and to align EBI's tooling with the global GA ecosystem and contribute to the E-SCOUT international trial.
Beyond this collaboration, you will contribute to the broader HPC sustainability agenda at EBI, including improving energy and carbon reporting, supporting user facing tooling, and helping establish best practices for sustainable HPC operations. This is a technically demanding role for someone who is equally comfortable working independently in a production codebase and collaborating across organisational boundaries.
- Develop and maintain Python based backend services that process HPC job telemetry and compute per job energy consumption and carbon footprint metrics.
- Contribute to the Green Algorithms open source project, working in close collaboration with the Cambridge GA team to integrate EBI's real time job monitoring capabilities into the shared GA codebase.
- Work with time series databases to ensure efficient storage, aggregation, and retrieval of large scale HPC telemetry data.
- Extend and improve carbon and energy reporting tooling, including alignment with established methodologies (e.g., regional carbon intensity, memory energy, GPU utilisation weighting).
- Support the development of Grafana based dashboards that give HPC users and stakeholders visibility into the environmental impact of their workloads.
- Contribute to wider HPC sustainability initiatives within EBI, including documentation, user guidance, and participation in the E-SCOUT international trial.
- Produce clear technical documentation and actively share knowledge within the team.
You have Strong experience writing production-grade Python, with a focus on data pipelines, concurrent processing, and clean, maintainable code. This is a backend data engineering role, not web API development. Solid experience with PostgreSQL and time-series data management. Familiarity with Timescale DB including continuous aggregates is an advantage, as is comfort with both psycopg2 and psycopg
3. Working knowledge of Apache Kafka as a message broker: consumer groups, offset management, and event-driven pipeline design. Kafka is a core component of EBI's real time job monitoring system. Confidence working with pandas Data Frames for data transformation, aggregation, and analysis used extensively in the Green Algorithms codebase. Proficient with Git, pytest, ruff, and Poetry. Able to contribute effectively to a collaborative open-source environment with CI/CD pipelines.
Proficiency in agentic coding and AI-assisted development tools (e.g., Claude Code, Git Hub Copilot, Cursor, or similar) to accelerate development within a well-structured codebase. Strong written and verbal communication skills. Able to work across two independent teams (EBI and Cambridge), translate technical trade-offs into clear proposals, and build consensus with stakeholders from…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).