×
Register Here to Apply for Jobs or Post Jobs. X

AI Research Scientist

Job in Somerville, Middlesex County, Massachusetts, 02145, USA
Listing for: Thespian Labs
Full Time position
Listed on 2026-10-08
Job specializations:
  • Research/Development
    Data Scientist, Research Scientist, AI Business & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Research the generative models at the core of one foundation model for embodied behavior. One model, any body, from a single forward pass.

Thespian Labs is an embodied intelligence research lab with roots at MIT and deep expertise in AI. We are building a foundation model for behavior, the layer between reasoning and execution where a body has to do something coherent while the world keeps moving. Reasoning can plan and decide. Execution can render pixels and drive motors. The layer in between is the one nobody has built, and it is the whole of our work.

We believe intelligence is something a body does in the world. A screen can answer you, but a body can be with you. That is why we are building one model that drives any body from a single forward pass, with perception, action, expression, and voice in a single stream. If language was the last decade of this work, embodiment is the next.

About

the Role

As an AI Research Scientist at Thespian Labs, you will research, design, and implement state-of-the-art generative models that sit at the core of our platform, the models that synthesize an embodied performance from a single forward pass. Working with novel, often diffusion-based and transformer architectures, you will generate the complex, high-fidelity 3D content, along with the motion, visual, and audio components, that let one model drive a humanoid on a real floor or a digital character on a screen.

This is core technology, not a wrapper around someone else’s. The team is small and the loop is short. You will own research problems end to end, from architecture to a model running on a real body, and work closely with the data team to turn unique, large-scale datasets into behavior the world can see. We move fast, we stay close to the latest research, and we hold the work to a high bar.

Not a benchmark, but the world.

Responsibilities
  • Develop and train advanced generative models, particularly diffusion-based architectures, for the synthesis of dynamic, high-fidelity 3D content that drives a body in real time.
  • Explore and implement cutting-edge techniques across generative modeling, including the visual, motion, and audio generation models that come together as a single embodied performance.
  • Design the system so these components work as one, composing multiple model types into a single forward pass rather than a pipeline of disconnected parts.
  • Collaborate closely with the data engineering team to define data requirements and leverage unique, large-scale datasets for model training.
  • Optimize models for quality, performance, and scalability, so research-grade results run fast enough to act in a live room and on a real robot.
  • Stay current with the latest research in generative AI and 3D computer vision, and bring new findings into our work quickly.
  • Take research from idea to deployment, owning problems end to end and closing the loop onto real humanoids and live digital characters.
What We’re Looking For
  • Proven experience developing and training deep learning models, particularly transformers and/or diffusion models.
  • Strong programming skills in Python and proficiency with modern ML frameworks like PyTorch.
  • A solid understanding of 3D computer vision and geometry, including how shape, motion, and space are represented and generated.
  • A strong portfolio of relevant projects or publications at top-tier conferences such as CVPR, ICCV, or NeurIPS.
  • The ability to take a model from research to a system that runs, with care for quality, performance, and scale rather than just a result in a notebook.
  • A bias for ownership and the appetite to drive open, ambiguous problems end to end in a small team.
Preferred Qualifications
  • Familiarity with audio, motion, or other generative domains beyond your core…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary