Engineering Manager AI GPU Cloud

Scaleway
Hybrid
Posted about 1 month ago
GPU Cloud

Job Description

WHY WE NEED YOU ? 

As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI infrastructure.

Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.

You will combine engineering leadership and technical ownership, helping the team automate critical infrastructure workflows, improve observability, and operate production-grade GPU platforms powering our sovereign cloud.

YOUR FUTURE TEAM 

We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.

You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.

The team works on some of our most critical AI and HPC infrastructure challenges, including GPU cluster automation, server lifecycle management, observability, reliability, and the integration of new GPU technologies.

You will collaborate closely with other GPU Cloud engineering teams, Product, Operations, and stakeholders across Scaleway.

YOUR DAILY ROUTINE 

Tasks

  • Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
  • Provide technical leadership on the architecture, reliability, and operation of large-scale GPU infrastructure
  • Design and drive automation for server provisioning and lifecycle management across GPU clusters
  • Build and improve observability, monitoring, logging, and alerting capabilities for production infrastructure
  • Own the SRE team's technical roadmap, priorities, and delivery
  • Ensure the reliability, scalability, performance, and resilience of production GPU clusters
  • Drive continuous improvements in automation, incident management, and operational processes
  • Collaborate closely with Engineering, Product, Operations, and other GPU Cloud teams
  • Recruit, onboard, coach, and develop engineers within the team
  • Lead incident response and post-incident improvements when critical production issues occur


ABOUT YOU 

HARDSKILLS:

  • Strong experience managing engineering teams in high-constraint production environments
  • Proven expertise with Kubernetes container orchestration
  • Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
  • Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
  • Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
  • Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)

SOFT SKILLS:

  • Strong engineering leadership and team management capabilities
  • Technical rigor and high attention to detail in production-critical environments
  • Ability to handle high-pressure operational situations and manage incident stress pragmatically
  • Excellent communication skills with the ability to convey challenging messages effectively
  • Collaborative mindset with a focus on empowering engineers rather than micromanaging


WHAT YOU WILL FIND AT SCALEWAY ++++ 

  • Hybrid work: We offer up to 3 days of remote work per week.
  • Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
  • Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
  • Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
  • International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
  • Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.

🚀 Why join the Scaleway adventure?

✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.

✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.

✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.

🔜 THE NEXT STEPS …

  • Discovery call with HR
  • Technical interview with the HPC team to understand your technical skills and approach to the role
  • Manager interview to validate your expertise
  • Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team 
  • HR interview and office visit to tour our offices and meet your future colleagues

Ready to Apply?

Take the next step in your career journey

Apply Now

About the job

Posted on
Jul 16, 2026
Job type
Full-time
Location
ParisHybrid

Explore more

Browse more jobs like this

Work arrangement

Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.