×
Register Here to Apply for Jobs or Post Jobs. X

Founding AI Engineer (Multimodal AI ​/ Computer Vision

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: MeeBoss
Full Time position
Listed on 2026-08-04
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
Position: Founding AI Engineer (Multimodal AI / Computer Vision)

We’re seeking Founding AI Engineer to join Full-time onsite in San Francisco, CA 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems. Location :
San Francisco, CA On-site work policy:
On-site 5 days per week in San Francisco. Full-time position Salary: $180K - $240K Hiring Count:
Looking to hire 2 candidate for this role Tech stack:
Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, Lang Chain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker Visa sponsorship available:
For H-1B transfers and TN visas supported. No new H-1B sponsorship. Candidates must currently reside in the USA or Canada. About This Role We're looking for an AI Engineer with 1-5 years of experience in applied AI/ML who has shipped agentic multimodal systems to real users - not just RAG wrappers or single-turn chatbots. You should be comfortable building production VLM pipelines with real hardware constraints (latency, power, connectivity) and have a track record of owning AI systems end-to-end in fast-paced environments.

Bonus points if you have wearable AI, autonomous driving, or industrial domain experience. What you'll do:

  • Build and ship the production agentic-VLM pipeline running on industrial smart glasses - multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service)
  • Own model orchestration and runtime optimization for edge inference, trading off model quality against latency with graceful fallback across connectivity conditions
  • Design and build the eval harness and data flywheel from scratch - failure-mode capture, customer-data fine-tune loops, and measurable model quality improvements each release
  • Ship real-time voice-video AI interfaces adapted to different end-user profiles: video-heavy, conversational speech, and proactive alerts
  • Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data
  • Drive multimodal model training when needed for on-premise deployments: open-source model SFT, RL post-training, and quantization
Role requirements
  • Seniority - 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems.
  • Work experience
    - Shipped multimodal and computer vision systems in the VLM era. (production, not demos, not pure research).
  • Owned the model layer end to end
    - Experience at a startup OR AI team building production AI products, preferably as a founder/CTO/founding engineer.
  • Big- company tenure (Meta Reality Labs, Snap, Apple Vision, Google, big streaming infra) is a BONUS only when on a directly- relevant team (AR/smart- glasses, real- time video/streaming, on- device/edge ML) OR paired with a builder signal (founder / early- startup / side projects / OSS).
  • A long single big- tech tenure with no builder signal and no relevant- team work is a NEGATIVE.
  • Production AR / wearable AI experience (Meta Reality Labs, Snap, Apple Vision) or autonomous driving Computer Vision.
  • Industrial domain exposure (data centers, energy grid, aerospace, manufacturing).
  • Education
    - Strong CS/ML/Eng background OR demonstrated equivalent shipping record.
  • Master's with vision or multimodal research component.
  • Hard skills Applied VLM / multimodal engineering: makes VLMs reliably do visual reasoning in production (image/video understanding, detection/segmentation as needed) in the modern VLM/VLA era. Depth is in shipping, hardening, and applied fine- tuning, not pretraining from scratch - this means VISION- language / video- language multimodality (images, video, VLM/VLA), NOT sensor- fusion, materials, audio- only, or time- series 'multimodal'. Applied agentic AI / model orchestration experience (vs.

    pure research) Evals discipline: rigorous eval harnesses (ground- truth, trajectory and tool- call accuracy, regression) to compare models/orchestrations and drive iteration On- prem / self- hosted model deployment: serving and optimizing open- weight ML models on customer hardware (for high- IP environments). Hands- on fine- tuning and deployment of a vLLM is a proxy In- context…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary