Founding AI Engineer (Multimodal AI / Computer Vision
Listed on 2026-08-04
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
We’re seeking Founding AI Engineer to join Full-time onsite in San Francisco, CA 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems. Location :
San Francisco, CA On-site work policy:
On-site 5 days per week in San Francisco. Full-time position Salary: $180K - $240K Hiring Count:
Looking to hire 2 candidate for this role Tech stack:
Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, Lang Chain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker Visa sponsorship available:
For H-1B transfers and TN visas supported. No new H-1B sponsorship. Candidates must currently reside in the USA or Canada. About This Role We're looking for an AI Engineer with 1-5 years of experience in applied AI/ML who has shipped agentic multimodal systems to real users - not just RAG wrappers or single-turn chatbots. You should be comfortable building production VLM pipelines with real hardware constraints (latency, power, connectivity) and have a track record of owning AI systems end-to-end in fast-paced environments.
Bonus points if you have wearable AI, autonomous driving, or industrial domain experience. What you'll do:
- Build and ship the production agentic-VLM pipeline running on industrial smart glasses - multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service)
- Own model orchestration and runtime optimization for edge inference, trading off model quality against latency with graceful fallback across connectivity conditions
- Design and build the eval harness and data flywheel from scratch - failure-mode capture, customer-data fine-tune loops, and measurable model quality improvements each release
- Ship real-time voice-video AI interfaces adapted to different end-user profiles: video-heavy, conversational speech, and proactive alerts
- Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data
- Drive multimodal model training when needed for on-premise deployments: open-source model SFT, RL post-training, and quantization
- Seniority - 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems.
- Work experience
- Shipped multimodal and computer vision systems in the VLM era. (production, not demos, not pure research). - Owned the model layer end to end
- Experience at a startup OR AI team building production AI products, preferably as a founder/CTO/founding engineer. - Big- company tenure (Meta Reality Labs, Snap, Apple Vision, Google, big streaming infra) is a BONUS only when on a directly- relevant team (AR/smart- glasses, real- time video/streaming, on- device/edge ML) OR paired with a builder signal (founder / early- startup / side projects / OSS).
- A long single big- tech tenure with no builder signal and no relevant- team work is a NEGATIVE.
- Production AR / wearable AI experience (Meta Reality Labs, Snap, Apple Vision) or autonomous driving Computer Vision.
- Industrial domain exposure (data centers, energy grid, aerospace, manufacturing).
- Education
- Strong CS/ML/Eng background OR demonstrated equivalent shipping record. - Master's with vision or multimodal research component.
- Hard skills Applied VLM / multimodal engineering: makes VLMs reliably do visual reasoning in production (image/video understanding, detection/segmentation as needed) in the modern VLM/VLA era. Depth is in shipping, hardening, and applied fine- tuning, not pretraining from scratch - this means VISION- language / video- language multimodality (images, video, VLM/VLA), NOT sensor- fusion, materials, audio- only, or time- series 'multimodal'. Applied agentic AI / model orchestration experience (vs.
pure research) Evals discipline: rigorous eval harnesses (ground- truth, trajectory and tool- call accuracy, regression) to compare models/orchestrations and drive iteration On- prem / self- hosted model deployment: serving and optimizing open- weight ML models on customer hardware (for high- IP environments). Hands- on fine- tuning and deployment of a vLLM is a proxy In- context…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).