×
Register Here to Apply for Jobs or Post Jobs. X

Applied Scientist III - VLM R&D

Job in Kirkland, King County, Washington, 98034, USA
Listing for: Wyze
Full Time position
Listed on 2026-07-25
Job specializations:
  • Research/Development
    AI Business & Operations, Data Scientist
  • IT/Tech
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), AI Business & Operations, Data Scientist
Salary/Wage Range or Industry Benchmark: 134000 - 181000 USD Yearly USD 134000.00 181000.00 YEAR
Job Description & How to Apply Below

Overview

At Wyze, we make smart home technology accessible to everyone. We are known for disrupting markets with high-quality, affordable products—from cameras to lighting to sensors and more. We believe technology should simplify life, not complicate it. We’re a fast-moving, customer-obsessed team driven by curiosity and powered by data.

The Opportunity

We are looking for an Applied Scientist to drive the development of our in-house vision-language models (VLMs) for smart home video understanding. In this role, you will track breakthrough research from academia and the broader AI community, and rapidly translate it into our production VLM development. You will help build the next generation of smart home physical AI — foundation models that understand the physical world of the home.

You will work with tens of millions of authorized videos to deeply investigate user event patterns and build a physical smart home foundation model that impacts over 10 million Wyze households. We believe in advancing the field, not just our product: we encourage publishing your research and releasing open-source models and datasets to benefit the broader community. This is a rare opportunity to shape a category-defining product at the intersection of frontier multimodal research and real-world deployment at massive scale.

What

You’ll Do
  • Follow the latest breakthroughs in multimodal and vision-language research from academia and industry, evaluate their relevance, and apply them to our in-house VLM development for smart home video understanding
  • Train, fine-tune, and evaluate multimodal vision-language models on large-scale, real-world home video data
  • Design and run rigorous evaluation pipelines to measure model quality on video understanding tasks such as event detection, activity recognition, and temporal reasoning
  • Investigate user event patterns across tens of millions of authorized videos to inform model design and product direction
  • Contribute to the architecture and training of a physical smart home foundation model, drawing on advances in visual transformers, physical world foundation models, and embodied AI
  • Build rapid proofs of concept using AI-assisted research and development workflows, and carry promising directions from idea to validated prototype
  • Publish research at top venues and contribute open-source models and datasets that help advance the community
  • Help define research problems, set technical direction, and anticipate where academic research and industry solutions are heading
What We’re Looking For
  • PhD in Computer Vision, Machine Learning, or a related field; or a Master’s degree with a strong track record of research or applied impact (publications, open-source contributions, or shipped ML systems)
  • Hands-on experience training and evaluating multimodal vision-language models
  • Experience in one or more of: visual transformer algorithm innovation, physical world foundation models, or embodied AI
  • Strong research sense: the ability to define the right problems, choose promising directions, and predict how research trends will translate into industry solutions
  • Proficiency with AI-assisted research and fast POC development — you use modern AI tools to multiply your own research velocity
  • Solid engineering skills in Python and deep learning frameworks (e.g., PyTorch), with the ability to work with large-scale video data pipelines
Nice to Have
  • Publications at top venues (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, or similar)
  • Experience with video understanding, long-context temporal modeling, or efficient inference for edge/cloud deployment
  • Experience deploying ML models in consumer products at scale
Compensation

The base pay range for this role is $134,000 – $181,000 per year.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary