Summary
What you’ll impact
Our company is seeking a Senior Site Reliability Engineer to monitor and improve the reliability, performance, and scalability of its core workflow automation platform. The role involves on‑call duties, infrastructure management, orchestration work, and collaboration with engineering and leadership to define SLOs/SLAs and enhance developer experience.
Responsibilities
What you'll do
- Monitoring our core business-logic software, both via on-call and in non-urgent situations: describing its existing behavior and defining SLOs or SLAs that get us (you and we) to respond.
- Extending and monitoring our infrastructure stack. Because we are B2B, we have fewer raw users, but their usage is often business-critical and we need the infrastructure for that.
- Maintaining both a mental and a reified model of our systems: from this model, estimating risk, planning projects, and debugging efficiently.
- Collaborating with our engineering team and company leadership from your unique lens on site reliability — being vocal and clear in your expertise in this specialization
- Working on our core orchestration logic that determines how to efficiently run tens of thousands of workflows at the same time - we recently rebuilt our orchestrator on top of Temporal and we love it!
- Advising many parallel major backend engineering projects, both early in planning and through release. Join such projects for building from time-to-time.
- Optimizing our services for scalability, stability, and observability as our customer base grows and our product becomes more sophisticated.
- Improving our developer experience in tactical ways, and improving our overall engineering processes and practices more broadly, especially given the opportunities & pressures of LLM coding agents.
Requirements
What you’ll bring
- 5+ years of SRE, DevOps, or Platform engineering experience
- A proven record of building efficient, performant, and easy to extend systems.
- Has maintained quantitative metrics of site reliability, while also demonstrating judgment about appropriate strictness for SLOs & SLAs. Given our team size, the expectations are somewhat less formal and mature than many SRE teams have; you will both strengthen our approach, but also navigate tradeoffs.
- Familiarity with containerization and orchestration tools (e.g., Docker, Kubernetes) and how they interact with backend services, and with Linux
- Experience implementing and managing AWS infrastructure
- You’re not afraid to ask for help, and you’re happy to give it, too.
- You’re an enthusiastic communicator and you like working with a team that provides both mutual support and thoughtful critique.
- You're excited to join a hybrid team and work out of our NYC or SF office ~3 days a week.