Senior Machine Learning Engineer - AI Foundation
About the role
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.
We are looking for a full-time Machine Learning Engineer - AI Foundation, with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training very large foundation model and accelerating model training/inference.
Our mission is to solve the autonomous driving problem. You will work with a team of talented software engineers, machine learning engineers and research scientists to push the boundary of state-of-art machine learning models which will enable the next-generation E2E solution of autonomous driving.
Job Responsibilities:
- Design and implement training data pipeline that streams data from hundreds of petabytes of labeled and unlabeled data from a fleet of over a million vehicles.
- Implement training framework for all physical AI foundation models in XPeng, including VLA 2.0, XWorld, Robotics.
- Accelerate training with state of the art parallelisms, e.g., FSDP, Expert Parallel, Context Parallel, and data types.
- Accelerate model inference on the cloud for closed-loop simulation, reinforcement learning, and enterprise LLM/VLM applications.
- Master's Degree in CS/CE/EE, or equivalent, in industry experience.
- Deep knowledge of PyTorch.
- Knowledge of model inference framework (e.g. vLLM, SGLang)
- In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models.
- Experience of performing large scale distributed training of models.
- A track record of profiling model and doing detective work to improve model training and inference speed.
- Previous experience in the autonomous driving industry.
- Experience with CUDA language for writing custom ops.
- Experience with edge computing systems.
- Knowledge of disributed computing frameworks, such as Ray.
- A track record of efficiently solving complex problems collaboratively on larger teams
- A fun, supportive and engaging environment.
- Infrastructures and computational resources to support your work.
- Opportunity to work on cutting edge technologies with the top talents in the field.
- Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
- Competitive compensation package.
- Snacks, lunches, dinners, and fun activities.
Description as published by XPENG.