×
Register Here to Apply for Jobs or Post Jobs. X

Applied Researcher: -Device Multimodal Reasoning

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Apple Inc.
Full Time position
Listed on 2026-08-29
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Data Scientist, AI Business & Operations, Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 150400 - 277600 USD Yearly USD 150400.00 277600.00 YEAR
Job Description & How to Apply Below
Position: Applied Researcher: On-Device Multimodal Reasoning

Applied Researcher:
On-Device Multimodal Reasoning

Sunnyvale, California, United States Machine Learning and AI

The Video Computer Vision (VCV) organization is an applied research and engineering team developing real-time, on-device Computer Vision and Machine Perception technologies across Apple products. Within VCV, our team builds next-generation multimodal AI systems that combine on-device multimodal encoders, large language models, and foundation models to create intelligent systems capable of understanding, reasoning, and acting across language, vision, audio, and tools.

Our work is deeply integrated into the Apple ecosystem, partnering across hardware, software, and ML teams to deliver real-time, scalable, and privacy-preserving experiences reaching millions of users.

Description

We are seeking an Applied Researcher with deep expertise in multimodal reasoning at small model scale — making vision-language models in the smaller regime (under ~10B parameters, down to sub-1B) think, plan, and act reliably under strict compute, memory, and latency constraints. In this role you will own the reasoning side of the on-device multimodal stack: designing compact VLMs that reason over images, video, and 3D scene content;

compressing and distilling the reasoning process itself; and engineering the decoding and inference path that makes multi-step reasoning affordable on an Apple device. This role offers the unique opportunity to define what on-device intelligence looks like for hundreds of millions of users. You'll push the boundaries of what small models can achieve — enabling real-time multimodal understanding and multi-step reasoning without reliance on cloud connectivity.

You'll collaborate with hardware teams, compiler engineers, and ML researchers to unlock capabilities that few organizations can deliver at Apple's scale and quality bar. This role spans multiple dimensions of efficient on-device reasoning — including VLM architecture and connector design, reasoning post-training (SFT/RL), chain-of-thought compression, speculative and structured decoding, visual token reduction, quantization and distillation, and hardware-aware inference optimization. A core focus of this role is efficient reasoning: compressed and latent chain-of-thought, reasoning distillation from frontier teachers, adaptive test-time compute (knowing when — and how long — to think), speculative and structured decoding, KV-cache compression, and visual-token efficiency.

A second focus is reasoning over real-time visual perception experts. Rather than consuming pixels alone, the VLM should be able to invoke and reason over the outputs of specialist on-device vision models — feed-forward 3D scene and geometry estimators (VGGT-style reconstruction, depth, camera pose), human body and hand mesh/pose recovery, object detectors, localize rs, and trackers — and fuse those structured, metric outputs into its reasoning about the scene.

This raises real research questions: how to represent geometry, body parameters, and detections compactly in a token-budgeted context; how to schedule which experts run at which frame rate within a real-time budget; and how to train a small model to invoke, trust, and cross-check them. Efficiency is treated as a first-class metric here: reasoning quality is measured at a fixed latency, memory, and power budget.

Responsibilities
  • Design, train, and post-train compact vision-language models (under ~10B, including sub-1B) that perform multi-step visual reasoning, grounded visual understanding, and language generation within on-device resource budgets
  • Research and implement efficient reasoning techniques — compressed and latent chain-of-thought, reasoning-trace distillation, early-exit and budget-aware reasoning, adaptive compute allocation, and test-time scaling that maximizes reasoning quality per FLOP
  • Own the decoding stack for on-device inference: speculative and self-speculative decoding, draft models and multi-token prediction, structured/constrained generation, KV-cache compression and quantization, prefill/decode scheduling, and streaming latency (TTFT, tokens/sec)
  • Build reasoning over real-time…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary