AI summary
The Site Reliability Engineer will architect and maintain DeviantArt's high-throughput AWS infrastructure, ensuring uptime, optimizing database performance, and automating resource provisioning while participating in a 24/7 on-call rotation.
Job Description
We are seeking a Site Reliability Engineer to architect, scale, and maintain DeviantArt’s high-throughput AWS infrastructure, supporting over 1.5 billion monthly page views. In this role, you will drive infrastructure automation, optimize high-scale database performance, harden system security, and ensure maximum uptime through robust deployment pipelines and proactive site reliability management.
- High-Scale AWS Infrastructure: Architect and maintain auto-scaling AWS systems designed to reliably serve DeviantArt's 1.5B+ monthly page views.
- High Availability & Incident Response: Maximize platform uptime and rapidly resolve performance degradation or service outages.
- Database & Search Optimization: Scale, maintain, and troubleshoot sharded MySQL databases and search infrastructure (e.g., Vespa.ai) for low-latency queries.
- CI/CD & Dev Parity: Build zero-downtime pipelines (Terraform, Kubernetes, GitHub Actions, CodeBuild) and maintain production-parity dev environments.
- Infrastructure Automation: Automate resource provisioning and configuration management, backed by clear documentation and automated testing.
- Security & DDoS Mitigation: Enforce robust security protocols, defend against targeted DDoS attacks, and execute routine patch management.
- Cloud Cost Optimization: Monitor and tune AWS resource utilization to minimize infrastructure spend without compromising site performance.
- 24/7 Operations & On-Call: Ensure continuous reliability as part of a 3-engineer 24/7 on-call rotation.
Requirements
- Minimum of 10 years of experience working with systems at scale in either a Dev Ops, Platform Engineer, or Site Reliability Engineer type role.
- Excellent analytical skills with the ability to troubleshoot complex problems, analyze system bottlenecks, and implement effective solutions, from frontend through backend systems, sometimes during production degradation or outage.
- Exceptional command line Linux skills.
- In-depth knowledge of AWS services, infrastructure as code using Terraform, and container orchestration with Kubernetes.
- Experience building Docker images, and composing Helm charts.
- Proficiency in Python and Bash for scripting and automation.
- Experience with sharded MySQL databases.
- Proven track record in security compliance standards on large-scale web infrastructures and in-depth knowledge of DDOS mitigation strategies.
- A proactive mindset in identifying potential issues and taking pre-emptive actions to prevent downtime or performance degradation.
- Excellent communication skills, open minded, and capable of effectively collaborating with cross-functional teams and articulating technical concepts to non-technical stakeholders.
- Bonus points: if you have experience with GitOps tools and methodologies.
About the job
- Posted on
- Oct 9, 2026
- Job type
- Full-time
- Location
- Remote, OR, usRemote
Keep looking
Related roles you might like
Frontend Software Engineer (100% Remote - worldwide)
Tether Operations Limited
Bilingual Client Intake Specialist (Mandarin & English)
Law Offices of Sabrina Li
Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.
