Senior Researcher, Interpretability and AI Safety
Listed on 2026-09-06
-
Research/Development
Research Scientist
Senior Researcher in Interpretability and AI Safety
central Oxford
We are seeking a full time Senior Researcher to join Technical Safety and Governance (TSG) labat the Department of Engineering Science in central Oxford. The post is funded by the Oxford Martin AI Governance Initiative and is fixed-term to 1 year, with the possibility of an extension for an additional year.
The Senior Researcher will work on Interpretability, evaluations and AI safety for continuously learning systems. They will evaluate systems whose behaviour keeps moving, and track, from the inside, whether the structures that carry their capabilities and safety properties hold up under the change.
You will be responsible for the follow : (full details of duties available from the Job Description)
Research Collaboration and engagementYou will have completed a PhD in machine learning, computer science, or a closely related field and with several years of experience and have research experience in at least one of: mechanistic or representational interpretability; model evaluation and benchmarking; continual or lifelong learning; or AI safety (alignment, scalable oversight, adversarial robustness). A strong publication record at relevant venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or established AI safety venues and workshops) together with the ability to design, run, and make sense of large-scale experiments on foundation models.
Informal enquiries may be addressed to Fazl Barez at
For more information about working at the Department, see
The Department holds an Athena Swan Bronze award, highlighting its commitment to promoting women in Science, Engineering and Technology.
AI, interpretability, continuous learning, evaluations chain-of-thought
#J-18808-LjbffrTo Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: