Summary
What you’ll impact
Our organization is seeking a Site Reliability Engineer to build and shape an AI-native platform that powers contact-center conversation intelligence. The SRE will design and implement platform engineering patterns, improve CI/CD pipelines, extend the autonomous agent harness, and participate in on-call rotation to ensure high availability.
Responsibilities
What you'll do
- Contribute to patterns, design, and implementation of our domains; help shape the future of platform engineering.
- Build and improve systems that help reduce toil and enable the company's production infrastructure to remain available and operable under large-scale, real-time conversational AI traffic.
- Extend and iterate our agent harness: Unsupervised AI agents are currently used by about 10% of the dev team - help us grow that number. The agent harness includes CI, sandboxes, guardrails, and validation (e.g. agent-first eval loops).
- Own and improve our CI/CD pipelines and surrounding developer tooling: build and test performance, deployment ergonomics, and paved paths for new services.
- Participate in on-call rotation and incident management to ensure platform uptime and quality. (SRE owns the base infrastructure, not the applications; non-business-hours pages are rare)
Requirements
What you’ll bring
- 6+ years’ experience in software development enablement roles.
- Solid experience owning CI/CD platforms end to end - including domains like caching, architecture, and developer self-service.
- Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning. You pair these skills with a defensible opinion on where to avoid using AI tools.
- Familiarity with Node/TypeScript including making code changes (e.g. exposing new metrics), Python and Terraform for automation, and developing in a Kubernetes/Helm ecosystem.
- Practical experience with observability: logs/metrics/tracing, monitoring/alerting, incident management process, and tooling.
- Experience working in fully remote teams - tell us how you’ve made one work better.
- Bonus: Harness engineering experience - building platforms for autonomous agents.
- Bonus: Production-at-scale experience with GCP.
- Bonus: Telephony and SIP architectures, FreeSWITCH in particular.