Software Engineer, ML Platform (ML Training)

zoox
Foster City, CA
On-site
Full-time
Posted 13 days ago
Software

Job Description

Zoox is on a mission to reimagine transportation and ground-up build autonomous robotaxis that are safe, reliable, clean, and enjoyable for everyone. We are still in the early stages of deploying our robotaxis on public roads, and it is a great time to join Zoox and have a significant impact in executing this mission. The ML Platform team at Zoox plays a crucial role in enabling innovations in ML and CV to make autonomous driving as seamless as possible. 
 
The Opportunity
Would you like to be part of the ML Training platform team that enables autonomous driving, driving scenario understanding, learned planning and trajectories, large scale driving simulations and several other ML use cases at Zoox? You will get to work across all ML teams within Zoox -  Foundation Models, Perception, Prediction, Planner, Simulation, Data Science, Collision Avoidance, etc. and have the opportunity to significantly push the boundaries of how ML is practiced within Zoox.
 
The ML Training platform team builds and operates the core part of the ML platform that powers model training at scale. We are responsible for developing and operating ML tools, deep learning frameworks, and distributed model training infrastructure to support foundational models and reinforcement learning. This team also owns the model repository and model lifecycle management tools used by our applied research teams for in- and off-vehicle ML use cases. You will play a crucial role in reducing the time it takes from ideation to productionization of cutting-edge AI innovation. This team has a lot of growth opportunities as we expand our robotaxi deployments and venture into new ML domains.

Responsibilities

In this role, you will:

  • Build the Zoox Training framework leveraged by all ML teams within Zoox. This framework needs to be highly scalable, reliable, and efficient.
  • Design, implement, and operate a robust and efficient ML platform to enable the training, validation, serving, and monitoring of ML models.
  • Collaborate closely with cross-functional teams, including ML researchers, software engineers, and data engineers, to define requirements and align on architectural decisions.
  • Requirements

    Qualifications

  • 2+ years of ML infrastructure experience.
  • Experience with training frameworks like PyTorch, DeepSpeed, JAX, Ray, etc.
  • Experience working with cloud providers like AWS.
  • Bonus Qualifications

  • Experience with building large-scale, cost-efficient distributed model training and ML compute infrastructure.
  • Experience with building model lifecycle management tools and experimentation
  • Ready to Apply?

    Take the next step in your career journey

    Apply Now

    Explore more

    Browse more jobs like this

    Work arrangement

    Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.