Senior/Staff Software Engineer, ML Performance Optimization

zoox
On-site
Posted almost 3 years ago
Software

Job Description

Zoox is on a mission to reimagine transportation and ground-up build autonomous robotaxis that are safe, reliable, clean, and enjoyable for everyone. We are still in the early stages of deploying our robotaxis on public roads, and it is a great time to join Zoox and have a significant impact in executing this mission. The ML Platform team at Zoox plays a crucial role in enabling innovations in large-scale Foundation models, VLMs, and VLAs to make autonomous driving as seamless as possible. 
The Opportunity
Are you excited to drive our ML Performance Optimization initiatives and make our ML models that enable autonomous driving as fast and efficient as possible? You will get to work with SOTA accelerators, cutting-edge techniques in distributed training, quantization, distillation, and pruning, among other things, working closely with all the Autonomy teams within Zoox - Perception, Prediction, Planner, Simulation, Collision Avoidance, and have the opportunity to significantly push the boundaries of how ML is practiced within Zoox.
We build and operate the base layer of ML tools, model development, and serving systems that our applied research teams use for in- and off-vehicle ML use cases. You will work alongside a team of strong software engineers and act as a force multiplier for our internal customers. This team has many growth opportunities as we expand our robotaxi deployments and venture into new ML domains. If you want to learn more about our stack behind autonomous driving, please look here. If you want to learn more about our ML Infrastructure, here is one of our past talks at re:Invent.

Responsibilities

In this role, you will:

  • Develop and execute a strategic vision for the ML Performance Optimization team to unlock ML innovation in autonomous driving and rider experience. 
  • Lead the design, implementation, and operation of cutting-edge ML Training OR Inference performance optimization techniques to scale our VLM, VLA, and Foundational models and deploy them efficiently in our robotaxi.
  • Collaborate closely with x-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions.
  • Enable the engineers in the team to grow their careers by providing technical guidance and mentorship.
  • Requirements

    Qualifications

    Note: You do not have to meet all the requirements below to be considered for this position:

  • Strong experience with training frameworks like PyTorch, leveraging GPUs efficiently for distributed model training.
  • Experience with GPU-accelerated inference using TensorRT or similar frameworks.
  • Experience using profiling tools like NVIDIA's Nsight or PyTorch's Profiler for identifying model training and serving bottlenecks.
  • Proficient in Python and C++
  • Experience with model compression techniques to reduce model size and improve performance.
  • Bonus Qualifications

  • 10+ years of total experience, including 4+ years of working on large-scale model training or inference platforms.
  • Excellent leadership skills with a demonstrated ability to lead high-performing engineering teams.
  • Ready to Apply?

    Take the next step in your career journey

    Apply Now

    About the job

    Posted on
    Oct 28, 2023
    Job type
    Full-time
    Salary range
    USD 242000-389000 per-year-salary
    Location
    Foster City, CAOn-site

    Explore more

    Browse more jobs like this

    Work arrangement

    Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.