Machine Learning Evaluation Specialist
Greater London, London, Greater London, W1B, England, UK
Listed on 2026-10-08
-
Research/Development
AI Evaluation, Data Scientist, AI Business & Operations -
IT/Tech
AI Evaluation, Machine Learning/ ML Engineer, Data Scientist, AI Business & Operations
About
The Role
The quality of AI depends entirely on the quality of the problems used to test it. We're looking for researchers and domain experts with deep machine learning knowledge to design the evaluation challenges that define — and push — the limits of today's most capable AI systems.
About
The Role
The quality of AI depends entirely on the quality of the problems used to test it. We're looking for researchers and domain experts with deep machine learning knowledge to design the evaluation challenges that define — and push — the limits of today's most capable AI systems.
This isn't routine review work. You'll apply your hard-earned research expertise to craft problems that state-of-the-art models genuinely struggle to solve. Your contributions directly shape how the next generation of AI is measured, benchmarked, and improved.
- Organization:
Alignerr - Type:
Hourly Contract - Location:
Fully Remote - Commitment: 10–40 hours/week
- Design complex, original machine learning problems rooted in your specific domain of expertise
- Craft evaluation tasks that require advanced domain knowledge well beyond standard ML pipelines
- Draw from your own research experience to create problems that genuinely challenge state-of-the-art AI
- Define rigorous problem statements, evaluation criteria, and gold-standard solutions
- Assess AI-generated ML solutions for correctness, creativity, and methodological soundness
- Document problem difficulty levels, required domain knowledge, and expected AI failure modes
- Collaborate asynchronously with a global team of researchers and engineers
Who You Are
- Graduate-level expertise (MS or PhD preferred) in a scientific or technical domain that intersects with machine learning
- Strong working knowledge of ML methods — model selection, feature engineering, evaluation metrics, and pipeline design
- Deep familiarity with active research problems in your field
- Able to identify precisely where general ML knowledge falls short and specialized domain insight becomes critical
- Experience publishing or conducting original research is highly valued
- Excellent written communication — you can articulate complex problems clearly and precisely
- Self-motivated and comfortable working independently on intellectually demanding tasks
- Computational biology, genomics, or bioinformatics
- Climate science and environmental modeling
- Medical imaging and healthcare ML
- Materials science and computational chemistry
- Astrophysics and signal processing
- Natural language processing for low-resource or specialized corpora
- Robotics, control theory, or reinforcement learning in complex environments
- Financial modeling and quantitative analysis
- Work at the true frontier of AI evaluation and safety research
- Collaborate with top research labs pushing the boundaries of what AI can do
- Finally put your specialized domain expertise to use in a high-impact, meaningful way
- Full autonomy over your schedule — work when and how you do your best thinking
- Flexible, fully remote contract with potential for ongoing work and deeper research involvement
- Build your profile as a recognized contributor to cutting-edge AI development
- Join a global community of researchers and engineers who take this work seriously
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).