×
Register Here to Apply for Jobs or Post Jobs. X

Backend Engineer

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: Bespoke Labs
Full Time position
Listed on 2026-09-10
Job specializations:
  • Software Development
    Backend Developer, DevOps, Software Engineer, Unix/Linux
Salary/Wage Range or Industry Benchmark: 100000 - 140000 USD Yearly USD 100000.00 140000.00 YEAR
Job Description & How to Apply Below

About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-Mini Chart-7B and Bespoke-Mini Check, and built the environment infrastructure that frontier labs and enterprises use to make their agents reliable.

Bespoke is uniquely positioned to capture a large share of data and RL environment curation.

About the Role

We're looking for an Infrastructure Engineer to own the execution layer beneath our RL environments: the systems that let an agent operate inside a realistic, multi-tool world coherently for hours or days.

This is a hard systems problem disguised as an AI job. As the tasks agents can complete keep lengthening, the environments that train them have to stay coherent across far longer horizons than anything that exists today. That means sandboxing and isolation you can trust, execution that's fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot.

You'll build the platform that makes all of this possible.

You'll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

What You'll Do
  • Environment Execution & Sandboxing:

    • Design and own the sandboxing and execution layer that environments run inside. Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.

    • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known‑good state, patch, and continue.

    • Extend execution to long‑horizon and multi‑node environments, where an agent operates across many tools and services over hours or days.

  • Performance & Scale

    • Own the performance characteristics of the platform: throughput, latency, and cost‑per‑rollout at scale.

    • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.

    • Profile and remove bottlenecks across the stack, from container startup to environment teardown.

    • Build the observability that lets us understand what's happening inside thousands of concurrent, long‑running rollouts.

  • Environment Platform

    • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.

    • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.

  • Collaboration & Production Excellence

    • Scale prototypes into production systems with reproducible workflows and high engineering standards.

    • Write the documentation and tools that let internal teams and external users build on the platform.

  • What We're Looking For
  • Systems & Infrastructure

    • Strong track record building production systems or research infrastructure at scale: distributed systems, execution engines, container/sandboxing infrastructure, or similar.

    • Deep comfort with the systems layer: containers and isolation (e.g. name spaces, cgroups, VMs, gVisor/Firecracker‑style sandboxing), file systems, process and state management.

    • Experience making systems fast and cheap — profiling, scheduling, resource utilization, and cost optimization at scale.

    • Proficiency with cloud platforms (GCP, AWS) and distributed computing.

    • Strong engineering fundamentals and a systematic…

  • To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
    (If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary