AI summary
The SRE Tech Lead is responsible for system design, ensuring product reliability, and mentoring team members while driving automation and improving system performance.
Job Description
Location: Remote
Salary: £80,000 - £90,000
About us
At Arbor, we’re on a mission to transform the way schools work for the better.
We believe in a future of work in schools where being challenged doesn’t mean being burnt out and overworked. Where data guides progress without overwhelming staff. And where everyone working in a school is reminded why they got into education every day.
Our MIS and school management tools are already making a difference in over 7,000 schools and trusts. Giving time and power back to staff, turning data into clear, actionable insights, and supporting happier working days.
At the heart of our brand is a recognition that the challenges schools face today aren’t just about efficiency, outputs and productivity - but about creating happier working lives for the people who drive education everyday: the staff. We want to make schools more joyful places to work, as well as learn.
About the role
We are looking for an experienced and collaborative Site Reliability Technical Lead to join our Site Reliability team and take ownership of system and solution design to ensure our products are robust, scalable, and secure. The remit and focus of the role is to blend deep technical expertise with leadership, requiring you to mentor and coach engineers, embed a culture of quality and reliability, and guide the team in making sound technical decisions. It’s a broad and exciting role, so we’re looking for someone up for a challenge - if you’re highly technical and a good communicator, this is the role for you.
Core responsibilities
- Architectural Leadership: Define and guide system architecture, balancing trade-offs between speed, scalability, maintainability, and security to meet business goals.
- Reliability and Performance: Champion accountability from design through to production by ensuring systems are observable and meet agreed Service Level Objectives (SLOs). Drive continuous improvement in platform reliability, performance, and efficiency.
- Incident Management: Lead Root Cause Analysis (RCA) when issues occur and contribute to optimizing the incident response process and framework.
- Automation: Drive automation initiatives across the team to reduce operational toil and improve system efficiency.
- Technical Standards: Uphold coding standards, promote automated testing, and work with the architecture community to drive technology adoption and share best practices across teams. Ensure production readiness standards for all services.
- Planning and Delivery: Lead technical estimation and feasibility assessments, ensuring plans are realistic and aligned with team capacity. Contribute to structured release planning and support post-release reviews.
- Mentorship and Coaching: Mentor and coach engineers through constructive feedback, knowledge sharing, and motivation. Foster alignment and help the team galvanise around technical solutions and goals.
- Collaboration: Work closely with Product Managers, Engineering Managers, and other engineers to align technical direction with product strategy. Communicate complex technical concepts clearly to both technical and non-technical stakeholders.
Requirements
About you
- Experience: Extensive professional experience in SRE, DevOps, or Platform Engineering on complex, scalable systems.
- Cloud Systems: Extensive expertise with AWS and distributed cloud architectures.
- Platform Scale: Proven experience operating platforms serving a high volume of requests (~1000 req/sec).
- Infrastructure as Code: Advanced proficiency with Terraform and configuration management tools.
- Programming: Strong skills in Python, Go, or a similar language for automation and tooling.
- Observability: Deep experience with monitoring and observability platforms (e.g., DataDog, Prometheus, or equivalent), plus incident/problem management.
- System Design: Expert understanding of distributed systems, microservices, and resilience patterns.
- Containerisation: Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes, ECS).
- CI/CD: Practical experience with building and maintaining CI/CD pipelines for automated deployments.
- Leadership: Demonstrated ability in mentoring and supporting the growth of fellow engineers.
Bonus Skills
- Experience with chaos engineering and reliability testing.
- Knowledge of security best practices and compliance frameworks.
- Background in agile and lean methodologies (Scrum/Kanban).
- Contributions to open-source projects or the SRE community.
About the job
- Posted on
- May 26, 2026
- Job type
- Full-time
- Location
- United KingdomRemote
Keep looking
Related roles you might like
Technical Business Analyst (Payroll) - 12 month FTC - SAMpeople
Arbor Education
Applied AI Engineer (JavaScript, Intermediate to Senior, Remote in Canada)
SimpliCity Digital Inc ("SimpliCity CMS")
Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.
