×
Register Here to Apply for Jobs or Post Jobs. X

Founding Mid-Training​/RL Infrastructure Engineer

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Model AI, Inc.
Full Time, Apprenticeship/Internship position
Listed on 2026-10-10
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 190000 - 230000 USD Yearly USD 190000.00 230000.00 YEAR
Job Description & How to Apply Below

Founding Large-Scale Mid-Training/RL Infrastructure Engineer

Location: Onsite in Palo Alto

Compensation: Competitive Salary + Equity

About Peano AI

Peano AI is building the infrastructure and application stack for the next generation of agentic AI systems.

We believe token usage will grow exponentially over the coming years, but routing all inference and training through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving and training infrastructure paired with an application layer designed for long-running, agentic workloads.

Peano AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, large-scale open-source model deployment, and pretraining of our own foundation model. By combining infrastructure and application design, we aim to make open-source models and custom foundation models significantly more performant, practical, and competitive.

About This Role

We are looking for a Large-Scale Mid-Training/RL Infrastructure Engineer to help build, optimize, and scale out our in-house foundation model training stack, with an emphasis on both core pretraining and RL-enhanced methods. This role is deeply technical and directly impacts Peano AI's core model product.

You will have a rare opportunity to design and run distributed infrastructure at massive scale using thousands of the latest NVIDIA GB300, VR200 GPUs, and next‑generation TPU v7x accelerators. You'll tackle ambitious optimization and scaling challenges with state-of-the-art compute.

You will architect, scale, and tune the distributed and accelerator-aware training infrastructure for massive language models, including data pipelines, model sharding, checkpointing, rollout and reward-based RL, and overall system throughput. Experience with large-scale model pretraining and mid-training interventions is required.

Preferred experience with frameworks such as Megatron, Transformer-Engine, verl, slime, and related large-scale training and RL toolkits. Familiarity with training pipeline scaling, mixed-precision, memory optimization, and deep understanding of both supervised and RL-based training cycles is essential.

What You’ll Do
  • Build, optimize, and scale the distributed training infrastructure for Peano AI’s foundation models, leveraging thousands of NVIDIA GB300/VR200 GPUs and TPU v7x accelerators for high throughput and efficiency.

  • Own and improve throughput, efficiency, scalability, cost, and reliability of end-to-end pretraining and RL pipelines.

  • Architect distributed data loading, sharding, checkpointing, batching, and accelerator utilization for multi-node LLM training at unmatched scale.

  • Integrate and optimize RL components: rollout, reward modeling, environment orchestration, and mid-training signal injections.

  • Work with and extend frameworks such as Megatron, Transformer-Engine, verl, slime, and other high-performance training libraries.

  • Tune memory usage, mixed-precision, and runtime performance for massive models and long-context workloads.

  • Debug and profile training performance bottlenecks across model code, distributed compute, networking, and infrastructure.

  • Collaborate with application, infrastructure, and research teams to ensure our foundation models deliver on performance and functionality.

  • Translate research prototypes and experimental features into production-ready, scalable training systems that operate across state-of-the-art compute clusters.

Qualifications
  • Significant experience with large-scale deep learning model training and distributed system design, including large GPU/TPU clusters (GB300/VR200,…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary