×
Register Here to Apply for Jobs or Post Jobs. X

Research Engineer, Infrastructure, Numerics

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Doist
Full Time position
Listed on 2026-07-22
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 350000 - 475000 USD Yearly USD 350000.00 475000.00 YEAR
Job Description & How to Apply Below

Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence. We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals. We are scientists, engineers, and builders who’ve created some of the most widely used AI products, including ChatGPT and Character.ai, open‑weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.

About

the Role

We’re looking for an infrastructure research engineer to design and build the core systems that enable efficient large‑scale model training with a focus on numerics. This role is ideal for someone who thrives at the intersection of research and systems engineering: a builder who understands both the maths of optimisation and the realities of distributed compute.

What You’ll Do
  • Design and optimise distributed training infrastructure for large‑scale LLMs, focusing on performance, stability, and reproducibility across multi‑GPU and multi‑node setups.
  • Implement and evaluate low‑precision numerics (e.g., BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.
  • Develop kernels and communication primitives that use hardware‑level support for mixed and low‑precision arithmetic.
  • Collaborate with research teams to co‑design model architectures and training recipes that align with emerging numeric formats and stability constraints.
  • Prototype and benchmark scaling strategies such as data, tensor, and pipeline parallelism that integrate precision‑adaptive computation and quantised communication.
  • Contribute to the design of internal orchestration and monitoring systems to ensure that thousands of distributed experiments can run efficiently and reproducibly.
  • Publish and share learnings through internal documentation, open‑source libraries, or technical reports that advance the field of scalable AI infrastructure.
Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Thrives in a highly collaborative environment involving many, different cross‑functional partners and subject‑matter experts.
  • Bias for action and initiative to work across different stacks and teams to ship solutions.
  • Strong engineering skills, ability to contribute performant, maintainable code and debug complex codebases in areas such as floating‑point numerics, low‑precision arithmetic, and distributed systems.

Preferred qualifications – we encourage you to apply if you meet some but not all of these:

  • Familiarity with distributed frameworks such as PyTorch/XLA, Deep Speed, Megatron‑LM.
  • Experience implementing FP8, INT8, or block‑floating point (MX) formats and understanding their numerical trade‑offs.
  • Prior contributions to open‑source deep learning infrastructure such as PyTorch, Deep Speed, or XLA.
  • Publications, patents, or projects related to numerical optimisation, communication‑efficient training, or systems for large models.
  • Experience training and supporting large‑scale AI models.
  • Track record of improving research productivity through infrastructure design or process improvements.
Location & Compensation

Location:

San Francisco, California.
Compensation: $350,000 – $475,000 USD per year (depending on background, skills and experience).
Visa sponsorship:
We sponsor visas.

Benefits

Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Equal Employment Opportunity

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary