Software Engineer, Systems Generalist
Listed on 2026-07-22
-
Software Development
Python, DevOps, Software Engineer, Cloud Engineer - Software
About the Role
Thinking Machines Lab’s mission is to empower humanity through the advancement of collaborative general intelligence. We build a future where everyone has access to the tools and knowledge to make AI work for their unique needs. We are scientists, engineers, and builders who have created widely used AI products such as ChatGPT and Character.ai, open‑weight models like Mistral, and open‑source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.
We are looking for generalist infrastructure and systems engineers to help build the systems that power our foundation models and enable internal research and product teams to create and ship models and powered products. You will join a small, high‑impact team responsible for architecting and scaling the core infrastructure. You will work across the full technical stack, solve complex distributed systems problems, and build robust, scalable platforms.
Infrastructure is the bedrock that enables every breakthrough. You will work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights across our models, products, and data assets.
Note:
This is an "evergreen role" that we keep open on an ongoing basis. We receive many applications and may not always have an immediate role that aligns perfectly with your experience. We encourage you to apply, and we review applications continuously to reach out as new opportunities open. You may re‑apply after gaining more experience, but please avoid applying more than once every six months.
We may also post singular roles for project or team‑specific needs; you are welcome to apply directly for those in addition to the evergreen role.
We generally interview, but during project selection we take into account your interests and experience alongside organizational needs. This flexible approach matches talented engineers with infrastructure teams where they have the greatest impact and growth potential. Depending on your expertise and interest, you may contribute to one or more of the following areas:
- Core Infrastructure – Support teams that train, research, and ultimately serve AI models. Build the infrastructure for clusters to reliably and safely train frontier models, including systems for large Kubernetes clusters with GPU workloads or infrastructure for Tinker.
- Data Infrastructure – Build and maintain data systems for research and products. Design and optimize data pipelines using tools such as Spark and other modern data infrastructure technologies, while embedding governance best practices.
- Developer Productivity – Extend research and engineering productivity by building tooling, frameworks, and systems that ensure well‑configured, optimized developer environments.
- Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large‑scale clusters and container orchestration systems such as Kubernetes or Slurm.
- Comfort operating across the stack and owning projects end‑to‑end.
- Ability to thrive in a highly collaborative environment with many cross‑functional partners and subject matter experts.
- Bias for action with a mindset to take initiative and work across different stacks and teams to ensure shipping.
- Strong debugging across application, OS, and network layers.
- Proficiency in Python or Rust (or similar), containers, and modern CI.
- Experience with Kubernetes, controllers/operators, or performance profiling.
- Familiarity with GPU/ML workflows or large‑scale data/evaluation pipelines.
This role is based in San Francisco, California.
CompensationDepending on background, skills, and experience, the expected annual salary range is $350,000–$475,000 USD.
Visa SponsorshipWe sponsor visas and are committed to working through the visa process together.
BenefitsThinking Machines offers generous health, dental, and vision benefits, unlimited paid time off, paid parental leave, and relocation support as needed.
Equal Employment OpportunityAs set forth in Thinking Machines’ Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).