What are the responsibilities and job description for the Senior Vision-Language Model (VLM) Engineer position at Girder AI?
Location: San Francisco, CA / Hybrid / Remote
Team: Multimodal AI / Foundation Models
Employment Type: Full-Time
Salary Range: $170,000 – $240,000 USD Equity
About Us
We are building the next generation of physical and multimodal AI systems. Our team develops state-of-the-art Vision-Language Models (VLMs) and Vision-Language-Action (VLA) models that power complex visual understanding, reasoning, and real-world execution across image, video, and embodied AI environments.
We are looking for an experienced VLM Engineer / Scientist to lead the architecture, pre-training, fine-tuning, and deployment of large-scale visual-text foundation models.
What you will do:
- Model Architecture & Training: Design, train, and optimize state-of-the-art vision-language architectures (e.g., contrastive learning, cross-attention fusion, autoregressive multimodal transformers, or VLA paradigms).
- Multimodal Data Engineering: Architect pipelines for harvesting, curating, and filtering large-scale image-text, video-text, and sensor data. Implement automated, agentic labeling and synthetic data generation workflows.
- Alignment & Fine-Tuning: Execute post-training, instruction tuning, and preference alignment techniques (e.g., SFT, DPO, RLHF) tailored for multimodal visual reasoning and grounding.
- Optimization & Inference: Quantize, distill, and optimize VLMs for low-latency edge deployment or high-throughput cloud inference pipelines (e.g., vLLM, TensorRT-LLM, ONNX Runtime).
- Evaluation & Benchmarking: Develop rigorous evaluation suites targeting visual QA, document understanding, spatial reasoning, grounding, and object interaction.
What we are Looking For:
- Education: Master’s or PhD in Computer Science, Machine Learning, Computer Vision, or a related field (or equivalent hands-on industry experience).
- Experience: 4 years of industry or applied research experience building and scaling deep learning models.
- Multimodal Expertise: Proven track record working with Vision-Language Models (e.g., CLIP, LLaVA, BLIP, Qwen-VL, PaliGemma, InternVL) or multimodal transformer architectures.
- Core Tech Stack: Advanced proficiency in PyTorch, Python, distributed training frameworks (DeepSpeed, Megatron-LM, FSDP), and CUDA acceleration.
- Scale Experience: Direct experience training models on large-scale GPU clusters (hundreds to thousands of GPUs) and handling massive multimodal datasets.
Nice-to-Haves
- Publications in top-tier machine learning/computer vision conferences (CVPR, NeurIPS, ICCV, ECCV, ICML).
- Experience with Vision-Language-Action (VLA) models, robotics, 3D vision, or autonomous systems.
- Experience with video understanding models, temporal reasoning, or agentic visual workflows.
Compensations & Benefits
- Base Salary: $170,000 – $240,000 / year (commensurate with level and experience) Equity
- Benefits: 100% company-paid medical, dental, and vision coverage; flexible PTO;
How to Apply
Submit your resume, GitHub profile, and links to any relevant papers, projects, or open-source contributions.
Salary : $170,000 - $240,000