×
Register Here to Apply for Jobs or Post Jobs. X

RL Infrastructure Engineer — Frontier AI Research

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Aionia Group
Full Time position
Listed on 2026-06-18
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer, AI Engineer (Applied/Software), Cloud Engineer - Software, DevOps
Salary/Wage Range or Industry Benchmark: 300000 USD Yearly USD 300000.00 YEAR
Job Description & How to Apply Below

A rare infrastructure role in a frontier RL research operation.

Compensation: $300K–$500K base + equity. San Francisco, on-site. Hiring urgent.

Opportunity

In a seed‑stage, well‑funded AI company, a small engineering team works with top researchers to automate task objectives and scale learning curriculums across thousands of GPUs.

Own the Systems Layer

This role builds the systems layer that enables researchers and applied ML engineers to run, debug, and reproduce large‑scale RL experiments—including distributed rollouts, training orchestration, inference, evaluation, data pipelines, observability, and reliability.

You will own infrastructure projects end‑to‑end, translating experimental workflows into durable, scalable solutions alongside leading RL minds.

Responsibilities
  • Design and deploy infrastructure for distributed RL training and inference across thousands of GPUs.
  • Improve reliability, debuggability, and throughput for large‑scale RL experiments.
  • Build interfaces for researchers and ML engineers to launch, inspect, compare, and reproduce experiments.
  • Eliminate bottlenecks in training, rollout generation, evaluation, data movement, and cluster utilization.
  • Establish engineering standards for RL infrastructure: testing, observability, versioning, and reproducibility.
Qualifications
  • 2+ years building infrastructure for LLM or RL systems.
  • Experience at a high‑engineering‑bar organization—top AI startups, frontier labs, or Big Tech RL research teams.
  • Hands‑on experience with GPU clusters, distributed training, model serving, or high‑throughput inference systems.
  • Familiarity with vLLM, SGLang, and modern LLM‑RL training frameworks.
  • Degree in CS, EECS, Mathematics, or a related field.
  • Very high level of curiosity and hypothesis‑driven thinking.
Nice to Have
  • Experience collaborating with ML researchers on infrastructure for messy experimental workflows.
  • Evidence of strong independent technical work—open‑source projects, competitions, or notable infrastructure contributions.
  • Familiarity with veRL, SkyRL, Slime, FSDP, Deep Speed, or similar distributed training frameworks.
Exclusions
  • No hands‑on LLM or RL infrastructure experience.
  • No meaningful GPU or distributed systems exposure.
  • No evidence of strong technical ownership or independent work.
Compensation & Logistics
  • $300,000 – $500,000 base (average offer $400K–$425K)
  • Competitive equity at seed stage.
  • Visa: H1B transfer, OPT, and O‑1 supported.
  • San Francisco, on‑site.
Interview Process
  • Background, fit, and logistics.
  • Technical interview—deep dive on infrastructure experience, systems design, and RL/LLM stack.
  • On‑site loop—1–2 days with researchers and engineering team in San Francisco.
  • #J-18808-Ljbffr
    To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
    (If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary