AI Research Scientist, Video Understanding & Physical AI
Listed on 2026-08-30
-
Research/Development
Robotics
The Problem
As robots powered by learned policies enter real manufacturing environments, making them trustworthy becomes as important as making them capable. Learned policies can be confidently wrong — executing smoothly while doing something other than what was intended — and understanding what a robot is actually doing on a factory floor, in real time and from observation alone, is an open research frontier.
Our lab researches the AI systems that make robot fleets trustworthy in production: real-time perception and reasoning over robot behavior — and, at its core, detecting when a robot is doing something wrong — built on multi-camera video understanding and multimodal signals from the operating environment and the robot itself.
Samsung SDS builds and operates the systems behind Samsung's global manufacturing — thousands of production lines worldwide. This research is being developed with that environment as its destination.
This is not a monitoring dashboard project. It is a frontier problem in video understanding: not just recognizing what a robot is doing, but reliably detecting when it is doing it wrong — on continuous, real-world behavior, in real time, under production latency and cost constraints.
The TeamYou would join a small, hands-on lab of PhD-level researchers — no layers between you and the research. Your research happens alongside real robots that our lab operates end-to-end: we collect our own data through teleoperation, train open-source robot foundation models on our GPUs, and deploy them to humanoid robots to test their behavior. Ongoing work extends to dexterous manipulation. The work that proves out in the lab has a path to pilot deployment in real manufacturing settings.
We run on a simple contract:
the mission is fixed; the method is yours. The lab's direction is clear and executive-sponsored, and everyone's work compounds toward it — but how you get there (which architectures, which formulations, which experiments) is your call to make and defend.
At this size, a new researcher is not headcount. Your technical judgment shapes how we get there from your first week.
What You Will Work On- Detecting anomalous robot behavior from real-time streaming video of humanoid robots at work — fusing external multi-camera views and, potentially, the robot's own egocentric video
- Efficient VLM research
: adapting vision-language models to achieve low-latency, low-cost on-prem deployment without sacrificing reasoning quality - Video-language grounding
: connecting continuous visual observations of robot behavior with language and structured task knowledge - Multimodal fusion beyond vision
: combining camera streams with robot state signals and manufacturing context data into a unified representation for judgment - World models and video prediction
: moving beyond detection — reasoning about the causes and dynamics of robot behavior, and anticipating what happens next
- Own the technical agenda. This is an executive-sponsored research effort at an early, formative stage. You will shape the architecture, the research questions, and the evaluation standards — not inherit them.
- Ship into the physical world at scale. Samsung SDS operates the systems behind Samsung's global manufacturing. When this research succeeds, it does not end as a paper or a demo — the deployment path runs onto real production lines, at a scale almost no research organization anywhere can offer. Papers are a milestone here, not the finish line.
- A data setting few others have. Synchronized multi-camera video of robots at work, paired with rich operational context from real manufacturing environments — a combination that few academic labs or frontier AI labs can match.
- Real robots, every day. The lab trains and runs learned policies on its own physical platforms — from teleoperation-based learning to dexterous manipulation — so your research has a living testbed: real robots executing real learned behaviors, generating the kind of behavioral data most video researchers never get to touch. Publishing and patenting are part of how we work.
- PhD in Computer Vision, Machine Learning, Robotics, or a related field,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).