×
Register Here to Apply for Jobs or Post Jobs. X

Founding AI Research Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-09-10
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer, Python, Software Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 250000 USD Yearly USD 150000.00 250000.00 YEAR
Job Description & How to Apply Below

About Klavis AI

Klavis AI is building high-quality agentic and coding data for frontier AI post-training.

Frontier models are increasingly bottlenecked not just by compute, but by the quality of the coding environments, trajectories, rubrics, rewards, and verification data used to train them. We build that data layer: long-horizon coding tasks, terminal-based software engineering environments, hidden-test verification, expert rubrics, gold trajectories, dockerized environments, and agentic tool-use workflows ready for RL and SFT. We already work with multiple frontier AI labs on production coding and agentic data.

The Role

We’re hiring a founding engineer who is genuinely exceptional at using LLMs and coding agents to build, test, debug, and ship real software.

You should be the kind of engineer who can make Claude Code, Codex, MCP tools, shell environments, custom evals, Docker, and agent pipelines feel like an extension of your hands. We are not looking for someone who has only tried basic ChatGPT prompts or simple API wrappers. We are looking for someone who already uses AI agents to build, debug, refactor, test, and ship faster than traditional engineering teams.

You’ll work directly with the founders to build the systems and datasets that help frontier labs train better coding and tool-use agents. In your first 30 days, you’ll onboard into our internal systems, ship improvements to them, and personally produce high-quality tasks end-to-end. By the end of the first month, you should understand what makes a task valuable for frontier post-training.

In 90 days, you’ll own a core product or infrastructure area from design to production. You’ll help define what “excellent” agentic and coding data means, build systems that scale production, and directly influence how frontier labs train their next-generation agents.

What you’ll work on

  • Build high-quality long-horizon coding datasets for frontier AI post-training
  • Create coding tasks, hidden tests, gold solutions, rubrics, dockerized environments, and agent trajectories
  • Design realistic software engineering workflows across terminals, repos, APIs, databases, and developer tools
  • Use LLMs and coding agents aggressively to accelerate engineering and data production
  • Build infrastructure for data generation, environment orchestration, verification, evaluation, and QA
  • Work with human experts to define difficult, realistic, and verifiable coding tasks
  • Design agentic tool-use workflows across SaaS apps, APIs, MCP servers, and external tools where needed
  • Turn ambiguous customer needs into reliable, scalable data products

What we're looking for

  • Have 3+ years of professional software engineering experience
  • Are extremely strong with LLMs, coding agents, and AI-assisted engineering workflows
  • Can build across Python, Type Script, shell, Docker, Git, APIs, databases, and modern dev tools
  • Have strong taste for realistic long-horizon coding tasks, test design, agent workflows, evals, and edge cases
  • Care deeply about correctness, verification, reproducibility, and data quality
  • Move fast without accepting sloppy work
  • Want the ownership, ambiguity, and intensity of joining at the founding stage

Strong signals

  • You have created long-horizon coding benchmarks, hidden tests, task environments, coding-agent evals, or data pipelines
  • You have built custom coding-agent workflows, MCP servers, eval systems, or LLM orchestration tools
  • You use Claude Code, Codex, Cursor, or similar tools daily and deeply understand their failure modes
  • You have strong open-source, infra, systems, ML engineering, devtools, or competitive programming experience
  • You can show examples of agents helping you ship real software, not just demos

Why join

  • Work on a core bottleneck for frontier AI: post-training data quality
  • Build products already used by frontier AI labs
  • Join a YC-backed company at the founding stage
  • Work directly with technical founders with deep agentic AI, infra, and ML systems experience
  • Own important engineering and product decisions from day one
  • Help define how future AI coding and tool-use agents are trained

Compensation

$150K - $250K + Equity 0.50% - 1.00%

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary