Legion ⢠Remote
Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance. This is a handsâon leadership role where you'll spend ~20â30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and crossâteam execution.
Hire and build a globallyâdistributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement. Own the reliability and infrastructure roadmap for our AWSâbased production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency. Lead the organizationâs security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation. Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes. Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce meanâtimeâtoâresolution. Apply agentic AI infrastructure concepts to SDLC and DevOps processes. Drive InfrastructureâasâCode, CI/CD, and automation practices. Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements. Ensure the platform meets the highest standards of security, compliance, and data protection. Lead and participate in the Incident Management onâcall rotation. Stay current on cloud, DevOps, and security best practices and provide technical guidance and thought leadership.
8â12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience. Deep handsâon experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3). Demonstrated experience running security operations (SecOps) â vulnerability management, incident response, and remediation of production security issues. 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana). Strong experience with InfrastructureâasâCode (e.g., Terraform, CloudFormation) and CI/CD automation. Proficiency in at least one of Go, Python, or Bash, with dayâtoâday use of Git and test automation pipelines. Handsâon experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS). Proven track record partnering crossâfunctionally with engineering and IT teams. Demonstrated experience leading incident management and onâcall practices for highâavailability production systems.
8â12 years of relevant experience, seniorâlevel expertise, proven leadership and peopleâmanagement skills, deep AWS and security operations knowledge, strong observability and IaC background, programming proficiency (Go/Python/Bash), Linux/Unix operations experience, crossâfunctional collaboration, incident management leadership.
Skills
Source & Verification
Source: We Work Remotely (RSS)
Discovered 7 Sept 2026 ⢠Last checked 7 Sept 2026
Similar Jobs