Senior/Staff Research Engineer — Vision-Language-Action Models (Autonomous Driving)
- $170,000–$260,000 per year
- Hybrid
- All ai & machine learning jobs
- Full-time
- Research and Development
About the role
You will join our core AI team at the frontier of autonomous decision-making, building the Vision-Language-Action (VLA) models that form SuperDrive's reasoning layer. You'll train VLA models that generate high-level driving decisions and trajectory guidance for on-board strategic decision-making, and design the knowledge distillation and compression techniques that transition large models onto on-board compute.
Responsibilities
- Design, train, and evaluate Vision-Language-Action models that generate high-level driving decisions and trajectory guidance in support of Plus's reasoning layer.
- Own a VLA workstream end to end — data, architecture, large-scale training, and on-vehicle validation.
- Build training and evaluation pipelines and rigorous metrics for VLA performance in driving contexts.
- Develop distillation and compression recipes to deploy large reasoning models on on-board compute.
- Apply SFT and RL post-training to improve reasoning, robustness, and long-tail behavior.
- Collaborate with perception, planning, and platform teams to bring models from research to production
Required qualifications
- M.S. minimum, Ph.D. preferred in CS, EE, Mathematics, Statistics, or a related field.
- 3+ years implementing and training models in a deep learning framework (PyTorch, TensorFlow, or JAX).
- Direct, hands-on experience training vision-language / vision-language-action models.
- Hands-on experience with model training, evaluation, and deployment in production.
- Thorough understanding of state-of-the-art vision-language / VLA models, diffusion, flow matching, and transformers.
- Experience with large-scale / distributed model training.
Preferred Qualifications
- Model distillation, quantization, and inference optimization (ONNX/TensorRT, mixed precision, custom kernels).
- SFT and RL post-training of large multimodal models.
- Hands-on experience with multi-modal sensor data (camera, LiDAR, radar).
- Publications at top venues (CVPR, NeurIPS, ICML, ICLR, CoRL, RSS, ICRA).
- Autonomous driving / ADAS experience.
Description as published by PlusAI.