Senior AI Engineer – Depth Estimation and Dense Scene Understanding
Listed on 2026-07-13
-
Engineering
AI Engineer (Applied/Software)
Senior AI Engineer
We are seeking a highly motivated Senior AI Engineer to develop next-generation AI models for depth estimation and dense scene understanding from images and video. The primary focus of this role is advancing state-of-the-art solutions for monocular depth estimation, stereo depth estimation, metric depth prediction, video-based depth estimation, and dense per-pixel vision models.
The ideal candidate will have deep expertise in computer vision and deep learning, with substantial experience developing depth-related AI models using modern neural network architectures. Knowledge of optical flow, motion estimation, temporal modeling, semantic segmentation, geometric vision, and 3D perception is highly desirable as complementary technologies that improve depth quality, temporal consistency, and scene understanding.
This role combines cutting-edge AI research with real-world deployment and will be instrumental in developing future vision systems capable of accurate, temporally consistent spatial understanding on edge devices. The candidate will drive innovations in depth estimation, video-based scene understanding, and dense perception while collaborating closely with hardware and system teams to design efficient AI solutions optimized for performance, power, memory, and latency constraints.
Experience with edge AI deployment, hardware-aware model design, and system-level optimization is highly valued.
Key Responsibilities
Depth Estimation Research and Development
- Design, develop, and optimize state-of-the-art AI models for monocular depth estimation, stereo depth estimation, video/multi-frame depth estimation, dense per-pixel depth prediction, and depth fusion, completion, and refinement.
- Develop novel AI architectures to improve depth accuracy, geometric consistency, edge preservation, temporal consistency, generalization across diverse environments, and robustness in challenging visual conditions.
Video Understanding and Motion-Based Learning
- Develop AI models that leverage temporal information from video sequences.
- Research and implement techniques involving optical flow estimation, temporal feature fusion, temporal attention mechanisms, Video Transformers, recurrent and memory-based architectures, and motion-depth-segmentation joint learning.
- Design approaches that utilize temporal and motion cues to improve depth prediction accuracy, enhance temporal stability, reduce frame-to-frame depth/segmentation flickering, improve scene understanding over time, and handle dynamic scenes and object motion.
Preferred Qualifications
- Experience with state-of-the-art depth and vision models, including Depth Anything, Metric3D, Uni Depth, Zoe Depth, MiDaS, RAFT, Flow Former, DINOv2, and Segment Anything (SAM).
- Experience with video foundation models, self-supervised depth learning, motion-depth joint learning, multi-modal vision models, neural rendering, NeRF, and Gaussian Splatting.
- Experience deploying AI models on edge devices.
- Ph.D. or M.S. with relevant job experience.
Minimum Qualifications:
• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR Master's degree in Computer Science, Engineering, Information Systems, or related field and 1+ year of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR PhD in Computer Science, Engineering, Information Systems, or related field.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).