×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist, Multi-Modal Human Understanding

Job in Burlingame, San Mateo County, California, 94012, USA
Listing for: Meta
Full Time position
Listed on 2026-09-07
Job specializations:
  • Research/Development
    AI Evaluation, AI Business & Operations
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta

  • 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
  • 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
  • Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
  • Experience writing production-quality or research-quality code in Python for multi-modal AI applications
  • Experience developing or fine-tuning Vision-Language Models for human understanding tasks
  • Experience with video foundation models, temporal transformers, or large-scale video pretraining
  • Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
  • Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary