Job Description
WHY WE NEED YOU?
Our growth is driving us to strengthen our SRE team to support and scale our production environments.
Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world.
#HPC #AI #GPU #CLUSTERS
YOUR FUTURE TEAM
We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.
You will join a newly formed team dedicated to building and operating Scaleway’s future AI infrastructure. As part of this group, you will design, maintain, and scale core systems and observability tools, partner with product teams, and ensure the reliability and performance of AI services across Scaleway.
YOUR DAILY ROUTINE
- Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents- Troubleshoot high-impact production issues in collaboration with other engineering teams
- Participate in an on-call rotation to handle incidents and ensure service continuity
- Implement and maintain observability solutions to monitor AI infrastructure and application health
- Contribute to AI infrastructure lifecycle management across different environments and countries
- Promote and apply best practices in terms of stability, resiliency, scalability, and security
- Maintain clear technical documentation for tools and procedures
- Contribute to system and tool evolution based on production feedback
- Collaborate closely with development teams to ensure infrastructure readiness- Participate in team rituals and knowledge-sharing initiatives
ABOUT YOU
🎯 SOFTSKILLS :
- Proactive and solution-oriented mindset
- Passion for automation and continuous improvement
- Strong collaboration and communication skills
- Ability to work independently and in a team
- Willingness to mentor and share knowledge
💻 HARDSKILLS :
- Experience with Python, Go or C++
- Strong scripting skills (Bash, Python)
- Hands-on experience with Linux systems (Ubuntu/Debian)
- Preferred hands-on experience with GPU & HPC infrastructure
- Knowledge of networking (TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)
- Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)
- Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)
- Experience managing relational databases (MariaDB)
- Understanding of CI/CD pipelines (GitLab)
- Comfortable with English (written and spoken)
WHAT YOU WILL FIND AT SCALEWAY ++++
Hybrid work: We offer up to 3 days of remote work per week.
Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
🚀 Why join the Scaleway adventure ?
✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
🔜 THE NEXT STEPS …
- Discovery call with a recruiter (30 min)
- Technical interview to validate your expertise (1h)
- Interview with the manager to understand your approach to the role (45 min)
- Interview with the manager to understand your approach to the role (45 min)
- Interview with the Head of the Tribe to deepen your discussions and assess your fit with the team (45 min)
- HR interview to tour our offices and meet your future colleagues
Keep looking
Related roles you might like
Event Operations Intern
Scaleway
On-siteInternship7 days ago
Field Marketing Manager
Scaleway
On-siteFull-time7 days ago
Field Marketing Specialist
Scaleway
On-siteFull-time7 days ago
Product Manager - Storage
Scaleway
On-siteFull-timeabout 1 month ago
DevOps Cybersecurity Engineer
Scaleway
On-siteFull-timeabout 1 month ago
Field Marketing Manager Italy
Scaleway
On-siteFull-timeabout 2 months ago
Pre-Sales Architect
Scaleway
On-siteFull-timeabout 2 months ago
Software Engineer IAM - Internship
Scaleway
On-siteInternshipabout 2 months ago
Backend Software Engineer (Python / DevOps)
Scaleway
On-siteFull-timeabout 2 months ago
Pre-Sales Solutions Architect - AI & GPU Infrastructure
Scaleway
On-siteFull-time2 months ago
Product Manager MLops
Scaleway
On-siteFull-time2 months ago
Object Storage Senior Product Manager
Scaleway
On-siteFull-time3 months ago
Pre-Sales Architect - Sweden
Scaleway
On-siteFull-time3 months ago
Digital Learning & Enablement Specialist
Scaleway
On-siteFull-time3 months ago
Software Engineer - Kubernetes Specialist
Scaleway
On-siteFull-time3 months ago
Explore more
Browse more jobs like this
Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.
