Manager, Software Engineering

sophos
Remote
Posted 6 days ago
Product Development

AI summary

The Manager of Software Engineering will lead a distributed team focusing on the reliability and scalability of cloud services, overseeing incident response, automation, and continuous improvement.

Eligible from: US and Canada

Job Description

Role Summary

We are seeking an experienced Manager, Software Engineering (SRE) to lead a team focused on improving the reliability, scalability, and operational maturity of Sophos cloud services and engineering delivery systems. In this role, you will lead a team distributed across the U.S. and Canada. The team is responsible for production reliability, operational readiness, incident response, escalation management, observability, automation, release reliability, and continuous improvement across cloud-based platforms.

This role requires a leader with a strong background spanning site reliability engineering, platform engineering, DevOps practices, release engineering, infrastructure management, and cloud operations. You should have enough technical depth to guide discussions around AWS, Kubernetes/EKS, CI/CD, Linux, infrastructure as code, observability, cost optimization, security practices, compliance needs, and automation, while primarily focusing on team leadership, execution, stakeholder alignment, prioritization, and improving how engineering teams deliver and operate services at scale.

Responsibilities

What You Will Do

Leadership and Team Management

  • Lead, coach, and develop a team of engineers focused on site reliability, platform operations, release engineering, and automation. Set clear priorities, expectations, and delivery goals aligned to business and engineering needs.
  • Support hiring, onboarding, performance management, and career development for team members.
  • Foster a culture of ownership, collaboration, continuous improvement, and operational excellence. 
  • SRE Strategy and Operational Maturity

    • Partner with senior leadership to define and execute the roadmap for SRE and operational improvements.
    • Drive improvements in reliability, availability, scalability, deployment confidence, and production readiness.
    • Drive operational efficiency and cost optimization practices across cloud platforms, balancing reliability, performance, and responsible cloud spend.
    • Help establish best practices for incident response, escalation management, observability, runbooks, post-incident reviews, and toil reduction.
    • Promote the use of reliability metrics, operational health indicators, and continuous improvement practices.
    • Platform, Release, and Automation Enablement

      • Partner with engineering, architecture, security, release, and operations teams to improve shared tooling and delivery workflows.
      • Support improvements across CI/CD pipelines, release automation, deployment reliability, and environment stability. Guide team efforts involving cloud infrastructure, Kubernetes/EKS, Linux systems, Terraform/IaC, Ansible, and automation tooling.
      • Encourage responsible use of approved AI tools to improve troubleshooting, documentation, automation, and engineering productivity.
      • Incident Management and Reliability

        • Ensure effective incident and escalation management practices are in place for timely response, communication, resolution, and follow-up.
        • Assess operational situations quickly, prioritize response activities, and balance customer impact, business risk, reliability, security, and delivery needs.
        • Drive root cause analysis and long-term corrective actions to reduce repeat incidents.
        • Improve observability, alerting, monitoring, and operational visibility across services and platforms.
        • Support capacity planning and operational readiness for scalable, high-availability systems.
        • Security, Compliance, and Governance

          • Partner with security, AppSec, compliance, and engineering teams to ensure operational practices support secure and reliable service delivery.
          • Help manage CI/CD security gates, vulnerability remediation workflows, audit readiness, and compliance-related operational controls.
          • Support vendor management for SRE, platform, observability, security, and cloud tooling, including renewals, evaluations, usage reviews, and service escalations.
          • Ensure team documentation, runbooks, operational evidence, and process controls are maintained to support audits and internal reviews.
          • Planning and Delivery

            • Lead sprint planning, backlog management, prioritization, and delivery tracking for SRE and platform initiatives.
            • Balance competing priorities across reliability work, operational escalations, security needs, technical debt, roadmap delivery, and team capacity.
            • Participate in quarterly planning, roadmap development, and cross-team dependency management.
            • Work with engineering, product, architecture, and security stakeholders to align reliability and operational priorities with business needs.
            • Use Agile/Scrum practices to improve execution, transparency, predictability, and continuous improvement across the team.
            • Collaboration and Communication

              • Work closely with engineering teams to understand service needs and ensure SRE priorities align with product and platform goals.
              • Lead technical and operational discussions with engineering leaders, product stakeholders, and executive audiences, including Directors, VPs, and C-level staff.
              • Present reliability updates, operational risks, roadmap progress, incident trends, and investment needs in a clear, concise, and audience-appropriate way.
              • Facilitate meetings, drive decisions, manage follow-ups, and ensure cross-functional conversations result in clear ownership and action.
              • Communicate status, risks, operational trends, and improvement opportunities to technical and non-technical stakeholders.

What You Will Bring

  • Extensive experience in site reliability engineering, DevOps practices, platform engineering, infrastructure management, operational efficiency, cloud infrastructure, or production operations.
  • Proven experience leading and managing engineers, fostering a culture of collaboration, accountability, ownership, and continuous improvement.
  • Strong understanding of reliability engineering, operational excellence, incident management, escalation management, observability, production support, and scalable system operations.
  • Strong working knowledge of cloud technologies, especially AWS; Azure experience is a plus.
  • Familiarity with operating and supporting containerized platform environments, including Kubernetes/EKS, Docker, Helm, and Linux-based systems.
  • Familiarity with infrastructure as code, configuration management, and automation practices using tools such as Terraform and Ansible.
  • Understanding of CI/CD pipelines, release automation, deployment practices, and developer tooling such as Jenkins, GitHub Actions, Artifactory, or similar tools.
  • Proficiency with scripting and automation concepts using languages such as Python, Bash, or Groovy.
  • Experience with monitoring, logging, and observability tools such as ELK stack, Prometheus, Grafana, CloudWatch, or similar platforms.
  • Working knowledge of security, application security, vulnerability management, compliance controls, audit readiness, and vendor management in a cloud or SaaS environment.
  • Familiarity with Agile/Scrum practices, sprint planning, backlog management, quarterly planning, roadmap development, and cross-team dependency management.
  • Excellent problem-solving and analytical skills, with the ability to assess situations quickly, prioritize work, manage escalations, and balance competing technical and business priorities.
  • Strong presentation and facilitation skills, with the ability to communicate clearly and confidently with technical teams, senior leaders, executives, vendors, and cross-functional stakeholders.
  • Ability to guide technical discussions, evaluate trade-offs, balance reliability, security, compliance, delivery speed, and operational efficiency, and help teams make sound engineering decisions.
  • Experience improving operational processes, reducing manual toil, mentoring engineers, managing delivery, and building healthy, accountable teams.
  • Experience with Java-based services, build systems, or code review workflows is a plus.
  • Comfort using or encouraging AI-assisted engineering tools responsibly within company guidelines.
  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. A Master’s degree is a plus.

Ready to Apply?

Take the next step in your career journey

Apply Now

About the job

Posted on
Sep 4, 2026
Job type
Full-time
Location
United StatesRemote

Explore more

Browse more jobs like this

Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.