Summary
What you’ll impact
The role is the first Machine Learning Engineer focused on the Benchmarks and Evaluations vertical, partnering with the general manager, researchers, and enterprise customers to build production benchmarks and supporting infrastructure. It involves hands‑on ML evaluation, backend system development, and customer‑facing delivery in a high‑ownership, remote environment.
Responsibilities
What you'll do
- Define, design, and build production benchmarks and evaluations with customers and internal researchers.
- Build and own backend infrastructure, including data pipelines, execution environments, storage, and orchestration.
- Create sandboxed environments for agentic evaluations involving tools, code execution, and multi-step tasks.
- Own the engineering portion of customer engagements from technical scoping through production delivery.
- Identify repeatable evaluation patterns and infrastructure gaps that can become scalable products.
- Move quickly through ambiguity while maintaining strong technical judgment and clear written communication.
Requirements
What you’ll bring
- 4 or more years of engineering experience, including hands-on machine learning model evaluation work.
- Experience deploying end-to-end ML evaluation or benchmark systems to production against demanding customer timelines.
- Strong proficiency with ML evaluation frameworks and benchmark design, including approaches such as LLM-as-judge.
- Ownership of backend and infrastructure systems such as large-scale data pipelines, execution environments, storage, and orchestration.
- Customer-facing engineering experience managing enterprise stakeholders.
- A degree in computer science, physics, or a related technical field.
- Current unrestricted U.S. work authorization without visa sponsorship.