×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

AI Advocate - Open-Source & Research

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: aijoblist
Full Time position
Listed on 2026-07-30
Job specializations:
  • Research/Development
    AI Evaluation
Salary/Wage Range or Industry Benchmark: 155000 - 240000 USD Yearly USD 155000.00 240000.00 YEAR
Job Description & How to Apply Below

About Snorkel

At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.

We’re on a mission to help enterprises transform expert knowledge into specialized AI  AI landscape has gone through incredible changes between 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production‑ready systems.

We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!

About the Role

You'll be Snorkel's primary technical voice in the open‑source and research communities. The work spans three audiences: frontier AI research teams (post‑training, RL environments, evals and benchmarks), enterprise ML and applied AI teams building specialized models on proprietary expertise, and the broader data‑centric AI community.

You’ll partner closely with our research, forward deployed research, and product teams to translate the methodology behind Snorkel's work into world‑class technical content, open‑source contributions, conference presence, and a thriving community of data‑centric AI practitioners.

Success looks like: a strong Snorkel open‑source presence, a steady cadence of high‑signal technical writing and research artifacts, marquee presence at the conferences that matter (NeurIPS, ICML, ICLR, AI Engineer World's Fair), and an engaged community of researchers and practitioners who view Snorkel as the trusted authority on data development for modern AI.

Responsibilities
  • Own Snorkel's external technical voice. Write methodology posts, technical deep‑dives, and research‑grade content on data development for frontier models.
  • Lead Snorkel's open‑source presence. Define the GTM approach, ship code, review PRs, recruit contributors, and keep the libraries credible and current. Build OSS that demonstrates Snorkel's methodology in practice, including reproducible evals and benchmark artifacts.
  • Advance the conversation on AI evaluation and benchmarking. Publish original work on how to measure agentic AI systems. Domain‑specific evals, agent evals, LLM‑as‑judge calibration, contamination and saturation, and the connection between evals and post‑training data.
  • Drive conference and research community presence. Land talks, papers, and workshops at NeurIPS, ICML, ICLR, AI Engineer World's Fair, and the right practitioner venues. Build relationships with academic labs and AI research teams.
  • Partner with the research team. Translate what's learned in research collaborations into externally shareable methodology, case studies, and tooling.
  • Set the bar for technical credibility. Design evals and benchmarks, prototype RL environments, and write code worth using. Your authority comes from doing the work, not just talking about it.
Preferred Qualifications
  • Experience. 6+ years in applied ML research, AI engineering, developer/research advocacy, or a research‑intensive technical role with significant public output. Prior Dev Rel/advocate experience welcome but not required.
  • Deep technical fluency in modern AI. Post‑training techniques (RLHF, DPO, RLAIF), evaluation methodologies, RL environment design, training data pipelines, synthetic data generation, and at least one applied domain (coding agents, reasoning, multimodal, agents).
  • Hands‑on experience with AI evaluation and benchmarks. You’ve built and run real evals: public benchmarks (MMLU, GPQA, SWE‑bench, HELM, BIG‑bench, Arena‑style head‑to‑head(s)), domain‑specific custom evals, and LLM‑as‑judge pipelines with proper calibration.
  • You build with AI, not just about AI. A power user of frontier coding agents (Claude Code, Cursor, Codex, and the like) in your day‑to‑day workflow, and you’ve built non‑trivial agentic systems yourself – multi‑step, tool‑using, with real evals and an opinion on what breaks.
  • A real public body of work. Talks, papers, blog…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary