More jobs:
RL Infrastructure Engineer — Frontier AI Research
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-06-18
Listing for:
Aionia Group
Full Time
position Listed on 2026-06-18
Job specializations:
-
Software Development
AI Reliability/ Performance Engineer, AI Engineer (Applied/Software), Cloud Engineer - Software, DevOps
Job Description & How to Apply Below
A rare infrastructure role in a frontier RL research operation.
Compensation: $300K–$500K base + equity. San Francisco, on-site. Hiring urgent.
OpportunityIn a seed‑stage, well‑funded AI company, a small engineering team works with top researchers to automate task objectives and scale learning curriculums across thousands of GPUs.
Own the Systems LayerThis role builds the systems layer that enables researchers and applied ML engineers to run, debug, and reproduce large‑scale RL experiments—including distributed rollouts, training orchestration, inference, evaluation, data pipelines, observability, and reliability.
You will own infrastructure projects end‑to‑end, translating experimental workflows into durable, scalable solutions alongside leading RL minds.
Responsibilities- Design and deploy infrastructure for distributed RL training and inference across thousands of GPUs.
- Improve reliability, debuggability, and throughput for large‑scale RL experiments.
- Build interfaces for researchers and ML engineers to launch, inspect, compare, and reproduce experiments.
- Eliminate bottlenecks in training, rollout generation, evaluation, data movement, and cluster utilization.
- Establish engineering standards for RL infrastructure: testing, observability, versioning, and reproducibility.
- 2+ years building infrastructure for LLM or RL systems.
- Experience at a high‑engineering‑bar organization—top AI startups, frontier labs, or Big Tech RL research teams.
- Hands‑on experience with GPU clusters, distributed training, model serving, or high‑throughput inference systems.
- Familiarity with vLLM, SGLang, and modern LLM‑RL training frameworks.
- Degree in CS, EECS, Mathematics, or a related field.
- Very high level of curiosity and hypothesis‑driven thinking.
- Experience collaborating with ML researchers on infrastructure for messy experimental workflows.
- Evidence of strong independent technical work—open‑source projects, competitions, or notable infrastructure contributions.
- Familiarity with veRL, SkyRL, Slime, FSDP, Deep Speed, or similar distributed training frameworks.
- No hands‑on LLM or RL infrastructure experience.
- No meaningful GPU or distributed systems exposure.
- No evidence of strong technical ownership or independent work.
- $300,000 – $500,000 base (average offer $400K–$425K)
- Competitive equity at seed stage.
- Visa: H1B transfer, OPT, and O‑1 supported.
- San Francisco, on‑site.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×