Senior Vision-Language Model (VLM) Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-08-03
Listing for:
Girder AI
Full Time
position Listed on 2026-08-03
Job specializations:
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Employment Type: Full-Time
Salary Range: $170,000 – $240,000 USD + Equity
About UsWe are building the next generation of physical and multimodal AI systems. Our team develops state-of-the
- Vision-Language Models (VLMs) and Vision-Language-Action (VLA) models that power complex visual understanding, reasoning, and real-world execution across image, video, and embodied AI environments.
We are looking for an experienced VLM Engineer / Scientist to lead the architecture, pre-training, fine-tuning, and deployment of large-scale visual-text foundation models.
What you will do:- Model Architecture & Training: Design, train, and optimize state-of-the-art vision-language architectures (e.g., contrastive learning, cross-attention fusion, autoregressive multimodal transformers, or VLA paradigms).
- Multimodal Data Engineering: Architect pipelines for harvesting, curating, and filtering large-scale image-text, video-text, and sensor data. Implement automated, agentic labeling and synthetic data generation workflows.
- Alignment & Fine-Tuning: Execute post-training, instruction tuning, and preference alignment techniques (e.g., SFT, DPO, RLHF) tailored for multimodal visual reasoning and grounding.
- Optimization & Inference: Quantize, distill, and optimize VLMs for low-latency edge deployment or high-throughput cloud inference pipelines (e.g., vLLM, TensorRT-LLM, ONNX Runtime).
- Evaluation & Benchmarking: Develop rigorous evaluation suites targeting visual QA, document understanding, spatial reasoning, grounding, and object interaction.
- Education: Master’s or PhD in Computer Science, Machine Learning, Computer Vision, or a related field (or equivalent hands-on industry experience).
- Experience: 4+ years of industry or applied research experience building and scaling deep learning models.
- Multimodal Expertise: Proven track record working with Vision-Language Models (e.g., CLIP, LLaVA, BLIP, Qwen-VL, Pali Gemma, InternVL) or multimodal transformer architectures.
- Core Tech Stack: Advanced proficiency in Py Torch , Python, distributed training frameworks (Deep Speed, Megatron-LM, FSDP), and CUDA acceleration.
- Scale
Experience:
Direct experience training models on large-scale GPU clusters (hundreds to thousands of GPUs) and handling massive multimodal datasets.
- Publications in top-tier machine learning/computer vision conferences (CVPR, NeurIPS, ICCV, ECCV, ICML).
- Experience with Vision-Language-Action (VLA) models, robotics, 3D vision, or autonomous systems.
- Experience with video understanding models, temporal reasoning, or agentic visual workflows.
- Base Salary: $170,000 – $240,000 / year (commensurate with level and experience) + Equity
- Benefits: 100% company-paid medical, dental, and vision coverage; flexible PTO;
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×