All roles

Senior Site Reliability Engineer, Platform

Seattle, USA·Engineering

KubernetesTerraformAWSPrometheusGoLinux

About the role

We are hiring a Senior SRE to keep Acme's core platform fast, resilient, and observable. You will own reliability for services that thousands of engineers and millions of users depend on every day.

What you'll do

  • Build and operate Kubernetes infrastructure across multiple regions
  • Define SLOs, error budgets, and incident response practices
  • Automate provisioning with Terraform and reduce operational toil
  • Lead blameless postmortems and drive long-term reliability fixes

Requirements

  • 5+ years in SRE, DevOps, or production engineering roles
  • Hands-on Kubernetes, Terraform, and a major cloud provider
  • Strong scripting or programming in Go, Python, or similar

Nice to have

  • Experience with observability stacks like Prometheus and Grafana
  • On-call leadership for large-scale systems

Benefits

  • Competitive salary and equity package
  • Fully covered health benefits for you and dependents
  • Flexible hybrid schedule and home-office stipend