Summary
What you’ll impact
Our company is seeking an AI Benchmark Engineer to build and operate the inference performance side of our measurement platform. The role involves creating automated benchmarking harnesses, defining metrics and methodology, and ensuring results are reproducible and defensible for a range of models, engines, and hardware. The engineer will work remotely with required overlap in Eastern Time and collaborate with data, GPU, and pricing teams.
Responsibilities
What you'll do
- You will write the harnesses, run the workloads, chase down the anomalies, and own the methodology that explains why the numbers are what they are.
- Build and maintain the inference benchmarking harness. Automated, containerized, reproducible runs across inference engines (vLLM, TensorRT-LLM, SGLang, TGI), model families, and hardware targets.
- Own the inference metrics that matter. Time-to-first-token, inter-token latency, sustained and peak throughput, latency percentiles under concurrency, goodput under SLO constraints, and their expression per dollar and per watt.
- Design workload profiles that reflect real usage. Input and output length distributions, concurrency patterns, streaming versus batch, prefill-heavy versus decode-heavy — rather than synthetic best-case runs.
- Capture and validate the configuration fingerprint. Engine version, model revision and quantization, driver and firmware, kernel and runtime versions, tensor and pipeline parallelism, scheduler settings. A result without its full configuration is not a result.
- Establish run repeatability and statistical rigor. Define warmup and steady-state criteria, variance thresholds, minimum run counts, and outlier handling. Know when a run is invalid and be willing to throw it out.
- Investigate performance anomalies. When a cohort underperforms, determine whether the cause is thermal, network, contention, misconfiguration, or a genuine silicon difference — and produce the evidence.
- Partner with our data and GPU benchmarking engineers on cohort definition, cross-suite consistency, and how inference results align with hardware-level benchmarks.
- Write the methodology. Public-facing write-ups, versioned methodology documents, and technical documentation that a skeptical customer can audit and reproduce.
- Work with our pricing products so measured inference performance can be expressed in economic terms — cost of capability per measured token per second.
- Track the field. New engines, new serving techniques, new model architectures, and emerging public benchmarks worth incorporating or explicitly rejecting.
Requirements
What you’ll bring
- 4+ years of engineering experience, with substantial hands‑on work in LLM inference, model serving, or performance engineering.
- Strong Python. You write automation and tooling that other people can run and trust, not one‑off scripts.
- Direct experience with at least one production inference engine (vLLM, TensorRT‑LLM, SGLang, TGI, or equivalent), including configuring and tuning it — not just calling its API.
- Working knowledge of what actually drives inference performance: batching and scheduling, KV cache behavior, quantization, tensor and pipeline parallelism, memory bandwidth limits, and prefill versus decode characteristics.
- Comfort with Linux, containers, and running workloads on GPU infrastructure.
- Genuine measurement discipline. You are skeptical of your own results, you control variables, you distinguish signal from noise, and you can explain the uncertainty in a number.
- Clear technical writing. You can document a methodology decision well enough that an outside engineer can reproduce it and a hostile reviewer cannot dismiss it.
- Self‑direction. The methodology here is not written yet. You will be defining the standard, not executing someone else’s spec.