Unknown Company • Iraq-wide
Looking for a Senior Site Reliability Engineer to join an Infrastructure Squad. This is a deeply hands-on role at the core of a high-traffic system, responsible for maintaining reliability, performance, and stability in a fast-paced environment. The engineer will handle real-time production challenges, manage alerts, and participate in a critical 24/7 on-call rotation.
Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time; Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment; Investigate incidents, perform root cause analysis, and implement long-term fixes; Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem; Deploy, manage, and optimize infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD); Maintain and evolve CI/CD pipelines and infrastructure-as-code practices.
Strong hands-on experience with Kubernetes in high-load environments; Experience with GitOps tools such as FluxCD or ArgoCD; Proven experience in incident response, root cause analysis, and postmortems in production systems; Solid experience with AWS, Terraform, Docker, and CI/CD pipelines; Experience with monitoring and observability tools like Datadog, Prometheus, Grafana, and ELK or CloudWatch; Strong understanding of networking concepts and protocols; Proficiency in at least one scripting language (e.g. Python, Go, Node.js).
Competitive Salary, Quarterly Bonuses, Unlimited Paid Time Off, Unlimited Paid Sick Leave, Remote & Flexible Working, Private Medical Insurance, Financial Support for Life Events, Professional Development Budget, International Exposure, Regular Company Events
Skills
Source & Verification
Source: Jobicy (RSS)
Discovered 14 Sept 2026 • Last checked 14 Sept 2026
Similar Jobs