Demo

Senior Vision-Language Model (VLM) Engineer

Girder AI
San Francisco, CA Full Time
POSTED ON 8/1/2026
AVAILABLE BEFORE 1/27/2027

Location: San Francisco, CA / Hybrid / Remote

Team: Multimodal AI / Foundation Models

Employment Type: Full-Time

Salary Range: $170,000 – $240,000 USD Equity


About Us

We are building the next generation of physical and multimodal AI systems. Our team develops state-of-the-art Vision-Language Models (VLMs) and Vision-Language-Action (VLA) models that power complex visual understanding, reasoning, and real-world execution across image, video, and embodied AI environments.

We are looking for an experienced VLM Engineer / Scientist to lead the architecture, pre-training, fine-tuning, and deployment of large-scale visual-text foundation models.


What you will do:


  • Model Architecture & Training: Design, train, and optimize state-of-the-art vision-language architectures (e.g., contrastive learning, cross-attention fusion, autoregressive multimodal transformers, or VLA paradigms).
  • Multimodal Data Engineering: Architect pipelines for harvesting, curating, and filtering large-scale image-text, video-text, and sensor data. Implement automated, agentic labeling and synthetic data generation workflows.
  • Alignment & Fine-Tuning: Execute post-training, instruction tuning, and preference alignment techniques (e.g., SFT, DPO, RLHF) tailored for multimodal visual reasoning and grounding.
  • Optimization & Inference: Quantize, distill, and optimize VLMs for low-latency edge deployment or high-throughput cloud inference pipelines (e.g., vLLM, TensorRT-LLM, ONNX Runtime).
  • Evaluation & Benchmarking: Develop rigorous evaluation suites targeting visual QA, document understanding, spatial reasoning, grounding, and object interaction.


What we are Looking For:


  • Education: Master’s or PhD in Computer Science, Machine Learning, Computer Vision, or a related field (or equivalent hands-on industry experience).
  • Experience: 4 years of industry or applied research experience building and scaling deep learning models.
  • Multimodal Expertise: Proven track record working with Vision-Language Models (e.g., CLIP, LLaVA, BLIP, Qwen-VL, PaliGemma, InternVL) or multimodal transformer architectures.
  • Core Tech Stack: Advanced proficiency in PyTorch, Python, distributed training frameworks (DeepSpeed, Megatron-LM, FSDP), and CUDA acceleration.
  • Scale Experience: Direct experience training models on large-scale GPU clusters (hundreds to thousands of GPUs) and handling massive multimodal datasets.


Nice-to-Haves


  • Publications in top-tier machine learning/computer vision conferences (CVPR, NeurIPS, ICCV, ECCV, ICML).
  • Experience with Vision-Language-Action (VLA) models, robotics, 3D vision, or autonomous systems.
  • Experience with video understanding models, temporal reasoning, or agentic visual workflows.


Compensations & Benefits


  • Base Salary: $170,000 – $240,000 / year (commensurate with level and experience) Equity
  • Benefits: 100% company-paid medical, dental, and vision coverage; flexible PTO;


How to Apply

Submit your resume, GitHub profile, and links to any relevant papers, projects, or open-source contributions.

Salary : $170,000 - $240,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Not the job you're looking for? Here are some other Senior Vision-Language Model (VLM) Engineer jobs in the San Francisco, CA area that may be a better fit.

  • EchoTwin AI San Francisco, CA
  • Company Overview EchoTwin AI is pioneering AI-driven infrastructure intelligence, redefining how cities are managed. Powered by a proprietary visual intell... more
  • 5 Days Ago

  • Preference Model San Francisco, CA
  • About Us Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real-world ML tasks. The pres... more
  • 26 Days Ago

AI Assistant is available now!

Feel free to start your new journey!