Head of Research
Listed on 2026-09-12
-
Research/Development
AI Evaluation, Research Scientist
Directly focused on building benchmarks and evaluation frameworks for LLMs and AI models — closely tied to vibe-coding and benchmarking work.
About the RoleLead Vals AI's research efforts to design and validate new evaluation methodologies and benchmarks for LLMs and other AI systems, set research direction, publish influential work, and build and manage a research team while working closely with enterprise customers and lab partners in San Francisco.
Job Description RoleHead of Research responsible for defining and advancing evaluation methodologies and benchmarks for large language models and other AI systems. The role sets research direction across Vals’ portfolio, publishes work that moves the field forward, recruits and grows a research team, and partners directly with enterprise customers and AI labs on real‑world evaluation problems.
Key Responsibilities- Develop new paradigms for evaluating long-horizon, real‑world tasks that current benchmarks and judge‑model approaches fail to capture.
- Oversee and set direction across Vals’ research portfolio and ongoing projects.
- Publish research and present results to the community, customers, and partners.
- Recruit, mentor, and grow a high-quality research team alongside the founders.
- Collaborate closely with enterprise customers and lab partners to solve practical evaluation problems.
- PhD in ML/NLP (in progress or completed) or equivalent industry research track record.
- Deep familiarity with the LLM evaluation landscape, including existing benchmarks, failure modes, judge‑model approaches, and human‑in‑the‑loop methodologies.
- Preference for research that influences real‑world deployments rather than easily gamed benchmarks.
- Strong written and verbal communication skills for publishing, presenting, and customer engagement.
- Ability to work in‑person in San Francisco.
- A widely cited benchmark or evaluation framework you built or co‑built.
- Prior experience at a frontier lab (e.g., Anthropic, OpenAI, Google Deep Mind, Meta FAIR) or a research‑led startup.
- Domain depth in verticals such as legal, finance, insurance, or healthcare.
- Experience leading or mentoring other researchers and maintaining a public research presence (papers, blog posts, talks, OSS contributions).
- Highly competitive salary and equity.
- Relocation and transportation support.
- Health and dental insurance coverage.
- Lunch and dinner provided, plus free snacks, coffee, and drinks.
- 401(k) plan.
- Unlimited PTO.
- $1,500 housing stipend (within one‑mile radius).
- Collaboration with leading AI labs and opportunity to work on foundational evaluation research.
- Python
- Django
- React
- AWS
- AWS CDK
- Compensation Range: $225,000 - $275,000 (USD)
Research Leadership Experimental Design LLM Evaluation Publication & Communication Team Building Mentorship Customer‑facing Collaboration Domain Expertise Learning Agility Ownership Problem Solving
Experience LevelUSD 225,/year
Employment TypeFull-time
- Relocation and transportation support
- Dental insurance
- Lunch and dinner provided
- Free snacks/coffee/drinks
- $1,500 housing stipend (within one‑mile radius)
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).