×
Register Here to Apply for Jobs or Post Jobs. X

Senior Vision-Language Model (VLM) Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Girder AI
Full Time position
Listed on 2026-08-03
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 170000 - 240000 USD Yearly USD 170000.00 240000.00 YEAR
Job Description & How to Apply Below

Employment Type: Full-Time

Salary Range: $170,000 – $240,000 USD + Equity

About Us

We are building the next generation of physical and multimodal AI systems. Our team develops state-of-the
- Vision-Language Models (VLMs) and Vision-Language-Action (VLA) models that power complex visual understanding, reasoning, and real-world execution across image, video, and embodied AI environments.

We are looking for an experienced VLM Engineer / Scientist to lead the architecture, pre-training, fine-tuning, and deployment of large-scale visual-text foundation models.

What you will do:
  • Model Architecture & Training: Design, train, and optimize state-of-the-art vision-language architectures (e.g., contrastive learning, cross-attention fusion, autoregressive multimodal transformers, or VLA paradigms).
  • Multimodal Data Engineering: Architect pipelines for harvesting, curating, and filtering large-scale image-text, video-text, and sensor data. Implement automated, agentic labeling and synthetic data generation workflows.
  • Alignment & Fine-Tuning: Execute post-training, instruction tuning, and preference alignment techniques (e.g., SFT, DPO, RLHF) tailored for multimodal visual reasoning and grounding.
  • Optimization & Inference: Quantize, distill, and optimize VLMs for low-latency edge deployment or high-throughput cloud inference pipelines (e.g., vLLM, TensorRT-LLM, ONNX Runtime).
  • Evaluation & Benchmarking: Develop rigorous evaluation suites targeting visual QA, document understanding, spatial reasoning, grounding, and object interaction.
What we are Looking For:
  • Education: Master’s or PhD in Computer Science, Machine Learning, Computer Vision, or a related field (or equivalent hands-on industry experience).
  • Experience: 4+ years of industry or applied research experience building and scaling deep learning models.
  • Multimodal Expertise: Proven track record working with Vision-Language Models (e.g., CLIP, LLaVA, BLIP, Qwen-VL, Pali Gemma, InternVL) or multimodal transformer architectures.
  • Core Tech Stack: Advanced proficiency in Py Torch , Python, distributed training frameworks (Deep Speed, Megatron-LM, FSDP), and CUDA acceleration.
  • Scale

    Experience:

    Direct experience training models on large-scale GPU clusters (hundreds to thousands of GPUs) and handling massive multimodal datasets.
Nice-to-Haves
  • Publications in top-tier machine learning/computer vision conferences (CVPR, NeurIPS, ICCV, ECCV, ICML).
  • Experience with Vision-Language-Action (VLA) models, robotics, 3D vision, or autonomous systems.
  • Experience with video understanding models, temporal reasoning, or agentic visual workflows.
Compensations & Benefits
  • Base Salary: $170,000 – $240,000 / year (commensurate with level and experience) + Equity
  • Benefits: 100% company-paid medical, dental, and vision coverage; flexible PTO;
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary