Research Scientist, Multi-Modal Understanding & Synthesis
Job in
Redmond, King County, Washington, 98052, USA
Listed on 2026-09-20
Listing for:
Meta
Full Time
position Listed on 2026-09-20
Job specializations:
-
Research/Development
AI Business & Operations, Research Scientist, Data Scientist, AI Evaluation
Job Description & How to Apply Below
Meta is seeking a Research Scientist to drive foundational research in multi-modal understanding, synthesis, and world models. In this role, you will advance the state of the art in building AI systems that perceive, reason across, and generate content spanning vision, language, audio, and other modalities. You will develop world models that learn rich internal representations of human behavior, enabling prediction, planning, and simulation.
Collaborating with world-class researchers and engineers, you will define research directions, publish influential work, and translate breakthroughs into technologies that power Meta's next-generation AI products.
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- PhD in Machine Learning, Computer Vision, Natural Language Processing, or a closely related field
- 6+ years of experience conducting AI research in multi-modal learning, generative models, or world models, including experience leading major research initiatives from conception through publication or production deployment
- Experience implementing and evaluating multi-modal systems using deep learning frameworks such as PyTorch or Tensor Flow, with proficiency in Python
- Experience publishing original research in peer-reviewed machine learning or AI venues
- Experience driving cross-functional technical decisions and communicating research findings and trade-offs to both research and engineering audiences through written documents and presentations
- Experience with techniques spanning multiple modalities such as vision-language models, multi-modal transformers, or cross-modal representation learning
- Experience developing large-scale multi-modal foundation models or vision-language models
- Experience with world models, predictive learning, or model-based reinforcement learning for planning and reasoning
- First-author publications at top-tier venues such as NeurIPS, ICLR, or CVPR demonstrating contributions to multi-modal learning, generative models, or world models
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×