Mountain View, Founding Engineer - Reinforcement Learning
Listed on 2026-07-22
-
Software Development
Founding Engineer - Reinforcement Learning
Location: Bay Area / San Mateo, CA
Employment Type: Full-Time
Department: Engineering
Deccan AI is a model training and evaluation startup headquartered in the Bay Area, with a delivery office in Hyderabad. We are founded by IIT Bombay, IIM Ahmedabad, and ex-Google alumni, and we work with leading AI labs and enterprise teams on high-quality human data, evaluations, and AI-first scaled operations.
We believe frontier AI systems are only as good as the data, evaluations, and human judgment behind them. Our work sits at the intersection of customer needs, research ambition, and operational execution. That makes the work high-touch, technical, ambiguous, and execution-heavy.
We are looking for people who can operate with urgency, write clearly, build trust, and turn ambiguity into outcomes.
About the RoleWe are hiring a Founding Engineer - Reinforcement Learning to serve as the SME for Deccan’s RL practice.
The market is moving beyond generic training data and simple evaluations. Frontier AI teams increasingly need high-quality environments, tasks, verifiers, feedback loops, and human-in-the-loop systems that can support RL, RLHF, agentic workflows, and model improvement. This work is technically hard, operationally messy, and only valuable when it reaches a bar that serious AI teams will actually use.
This is a hands‑on founding engineering role. You will design and build RL environments, evaluation harnesses, reward / verifier systems, task pipelines, and tooling that help Deccan serve frontier AI labs and technically sophisticated customers. You should be comfortable moving from research ambiguity to working systems, and from customer or researcher requirements to concrete environments that can be tested, improved, and scaled.
WhatYou’ll Do
- Design and build RL environments, task suites, evaluation harnesses, and feedback systems for frontier AI workflows.
- Translate ambiguous research or customer needs into concrete technical specs: task definitions, environment behavior, reward signals, verifier logic, grading criteria, data requirements, and failure modes.
- Prototype quickly, then harden the pieces that need to become reusable infrastructure.
- Design human‑in‑the‑loop workflows for RLHF, preference data, expert grading, model behavior evaluation, and quality improvement.
- Build tools to measure environment quality, task difficulty, model failure patterns, grader consistency, and production readiness.
- Work across engineering, ML, research, delivery, quality, and GTM so RL opportunities are technically credible before Deccan commits to them.
- Help define what frontier level means for Deccan’s RL environments, then raise the bar through real implementation.
- Write clear technical docs, playbooks, and customer‑facing explanations so the team can reuse what you build.
- Stay close to the market: agentic systems, coding environments, tool‑use tasks, verifiers, reward modeling, RLHF, and emerging post‑training workflows.
- Create enough structure that Deccan can move from one‑off RL experiments to a repeatable practice.
- Technical architecture for Deccan’s RL environments and supporting systems.
- Environment and task design for RL / RLHF workflows.
- Verifier, reward, grading, and quality‑measurement logic.
- Prototype‑to‑production path for RL tooling and environments.
- Technical feasibility assessments for RL opportunities before Deccan makes commitments.
- Internal engineering standards for reliability, instrumentation, reproducibility, and documentation.
- A reusable foundation that helps Deccan compete for serious RL and post‑training work.
We are looking for a strong engineer with deep ML judgment and a builder’s bias.
- technical expertise to design and implement RL environments, task pipelines, evaluation systems, and feedback loops;
- systems‑minded to build reusable tools instead of one‑off scripts;
- comfortable with ambiguity, incomplete specs, and fast‑changing customer requirements;
- rigorous about measurement, failure modes, quality bars, and reproducibility;
- high‑agency to create the first version without waiting for a mature team around you;
- clear…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).