×
Register Here to Apply for Jobs or Post Jobs. X

Mountain View, Founding Engineer - Reinforcement Learning

Job in San Mateo, San Mateo County, California, 94409, USA
Listing for: S27a
Full Time position
Listed on 2026-07-22
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Mountain View, USA Founding Engineer - Reinforcement Learning

Founding Engineer - Reinforcement Learning

Location: Bay Area / San Mateo, CA
Employment Type: Full-Time
Department: Engineering

About Us

Deccan AI is a model training and evaluation startup headquartered in the Bay Area, with a delivery office in Hyderabad. We are founded by IIT Bombay, IIM Ahmedabad, and ex-Google alumni, and we work with leading AI labs and enterprise teams on high-quality human data, evaluations, and AI-first scaled operations.

We believe frontier AI systems are only as good as the data, evaluations, and human judgment behind them. Our work sits at the intersection of customer needs, research ambition, and operational execution. That makes the work high-touch, technical, ambiguous, and execution-heavy.

We are looking for people who can operate with urgency, write clearly, build trust, and turn ambiguity into outcomes.

About the Role

We are hiring a Founding Engineer - Reinforcement Learning to serve as the SME for Deccan’s RL practice.

The market is moving beyond generic training data and simple evaluations. Frontier AI teams increasingly need high-quality environments, tasks, verifiers, feedback loops, and human-in-the-loop systems that can support RL, RLHF, agentic workflows, and model improvement. This work is technically hard, operationally messy, and only valuable when it reaches a bar that serious AI teams will actually use.

This is a hands‑on founding engineering role. You will design and build RL environments, evaluation harnesses, reward / verifier systems, task pipelines, and tooling that help Deccan serve frontier AI labs and technically sophisticated customers. You should be comfortable moving from research ambiguity to working systems, and from customer or researcher requirements to concrete environments that can be tested, improved, and scaled.

What

You’ll Do
  • Design and build RL environments, task suites, evaluation harnesses, and feedback systems for frontier AI workflows.
  • Translate ambiguous research or customer needs into concrete technical specs: task definitions, environment behavior, reward signals, verifier logic, grading criteria, data requirements, and failure modes.
  • Prototype quickly, then harden the pieces that need to become reusable infrastructure.
  • Design human‑in‑the‑loop workflows for RLHF, preference data, expert grading, model behavior evaluation, and quality improvement.
  • Build tools to measure environment quality, task difficulty, model failure patterns, grader consistency, and production readiness.
  • Work across engineering, ML, research, delivery, quality, and GTM so RL opportunities are technically credible before Deccan commits to them.
  • Help define what frontier level means for Deccan’s RL environments, then raise the bar through real implementation.
  • Write clear technical docs, playbooks, and customer‑facing explanations so the team can reuse what you build.
  • Stay close to the market: agentic systems, coding environments, tool‑use tasks, verifiers, reward modeling, RLHF, and emerging post‑training workflows.
  • Create enough structure that Deccan can move from one‑off RL experiments to a repeatable practice.
What You Will Own
  • Technical architecture for Deccan’s RL environments and supporting systems.
  • Environment and task design for RL / RLHF workflows.
  • Verifier, reward, grading, and quality‑measurement logic.
  • Prototype‑to‑production path for RL tooling and environments.
  • Technical feasibility assessments for RL opportunities before Deccan makes commitments.
  • Internal engineering standards for reliability, instrumentation, reproducibility, and documentation.
  • A reusable foundation that helps Deccan compete for serious RL and post‑training work.
What We’re Looking For

We are looking for a strong engineer with deep ML judgment and a builder’s bias.

  • technical expertise to design and implement RL environments, task pipelines, evaluation systems, and feedback loops;
  • systems‑minded to build reusable tools instead of one‑off scripts;
  • comfortable with ambiguity, incomplete specs, and fast‑changing customer requirements;
  • rigorous about measurement, failure modes, quality bars, and reproducibility;
  • high‑agency to create the first version without waiting for a mature team around you;
  • clear…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary