Summary
What you’ll impact
Senior Site Reliability Engineer (DevOps) at the organization will own the reliability of the production infrastructure on managed Kubernetes and AWS, building tools and automation to reduce manual toil. The role collaborates with Platform Engineering, Product, and QA to design resilient architectures, improve CI/CD pipelines, and enhance observability while ensuring cost efficiency, security, and developer enablement.
Responsibilities
What you'll do
- Own the reliability of production infrastructure running on managed Kubernetes and AWS, keeping a high-traffic system up and catching issues before customers do
- Work agent-first day to day: build, debug, and automate using agentic coding workflows as your default mode, not a fallback tool
- Build internal tools and automation that eliminate recurring manual work (toil) for the team, rather than just documenting runbooks around it
- Develop and operate monitoring, alerting, and observability systems (metrics, tracing, logging) across the stack
- Partner with engineering teams to design for reliability and performance from the start, not bolt it on after incidents
- Automate infrastructure management through infrastructure-as-code, and improve CI/CD pipelines and local developer workflows
- Lead and evolve incident response practices, including postmortems and blameless learning
- Optimize infrastructure for cost efficiency while maintaining high availability and security standards
- Contribute to security, compliance, and disaster recovery efforts as the platform scales
- Support developer enablement: build and improve the in-house tooling, local dev workflows, and internal platforms other engineers rely on
Requirements
What you’ll bring
- 5-8 years in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, with real ownership of a production system at meaningful scale — someone who drives reliability and tooling initiatives rather than waiting to be assigned them
- Fluent working agent-first day to day — directing coding agents to do real engineering work, not just occasional autocomplete
- Strong foundation in AWS-hosted data and networking services (RDS/Postgres, ElastiCache/Redis, VPC/networking) and experience running workloads on managed Kubernetes
- A track record of building tools, not just running playbooks: scripts, services, or internal platforms that removed manual work for a team
- Hands-on experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, CloudFormation)
- Expertise in observability stacks (metrics, tracing, logging) and modern monitoring practices
- Familiarity with security and compliance in cloud environments (SOC 2, GDPR, etc. a plus)
- A collaborative mindset with a passion for empowering developers to move fast safely
- Experience with (or strong interest in) developer enablement — building the internal tooling and platforms that make other engineers more productive
- Nice to have: experience operating across multiple cloud providers, as our infrastructure footprint expands