×
Register Here to Apply for Jobs or Post Jobs. X

Multimodal Perception Engineer (Vision + Tactile + Audio

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Proception Inc.
Full Time position
Listed on 2026-07-30
Job specializations:
  • Software Development
    Robotics, Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
Position: Multimodal Perception Engineer (Vision + Tactile + Audio)

Multimodal Perception Engineer (Vision + Tactile + Audio)

Build and deploy multimodal perception systems that fuse vision, tactile, force, and audio sensing. You'll work at the frontier of robotics perception and machine learning to create systems that enable robots to understand and interact with the world more like humans do. From foundation models to real-time deployment, you will shape the future of sensor-driven intelligence.

Requirements
  • MS or PhD in Computer Science, Robotics, Machine Learning, or a related field—or equivalent industry experience
  • Experience integrating vision, tactile, force, or audio sensors in robotic systems
  • Deep understanding of sensor fusion, time sync, calibration, and failure modes
  • Strong grasp of probability, optimization, signal processing, and linear algebra
  • Proficient in Python and PyTorch for model development
  • Experience with C++ for high-performance or embedded deployment
  • Comfortable working in Linux/Unix environments and with embedded hardware
  • (+) Experience training or fine-tuning ViTs, MAE, or transformer-based models
  • (+) Experience with diffusion models for generative or representation learning
  • (+) Experience with camera calibration, stereo/depth sensors, or event-based vision systems
  • (+) Experience deploying perception models to real-time or resource-constrained platforms
  • (+) Experience designing or managing sensor data collection pipelines for ML model training
Responsibilities
  • Design and implement multimodal sensor fusion systems combining vision, touch, force, and audio
  • Develop and fine-tune transformer and diffusion models for robotic perception tasks
  • Adapt foundation models for real-time inference on physical robots
  • Fuse high-frequency sensor data to estimate object pose, contact events, and material properties
  • Deploy models to real-world robots with performance and latency constraints
  • Collaborate with reinforcement learning, control, and hardware teams to close the perception-action loop
  • Take technical ownership of perception systems and help define future direction
Benefits
  • Competitive salary and meaningful equity
  • Comprehensive health, dental, and vision coverage
  • Work with world-class researchers and engineers in AI and robotics
  • Backed by top investors (YC, leading VCs)
  • High-ownership role with opportunity to lead multimodal perception efforts
  • Help define and build next-generation embodied intelligence systems

Interested in this role?

Applying takes a few minutes — you'll need an account to submit.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary