Featured Job

Research Engineer

San Mateo, CA Full-time On-site $280k — $280k per year 09/12/2026 Job ID: 000192
Apply Now
Experiment design Ablations Benchmarking Error analysis NL-to-SQL generation

Summary

What you’ll impact

Our company is seeking a Research Engineer to design experiments, develop and productionize advanced NLP and LLM techniques, and improve conversational AI systems for enterprise customers. The role involves rigorous evaluation, prototype development, and contributing to the AI community, based in San Mateo, CA.

Responsibilities

What you'll do

  • Design and run rigorous experiments, including ablations, benchmarking, and error analysis, to improve NL-to-SQL generation, retrieval, ranking, and agentic reasoning.
  • Advance multi-turn conversational understanding through improved context retention, entity resolution, conversational memory, and intelligent clarification strategies.
  • Build robust evaluation frameworks, including automated benchmarks, regression suites, and LLM-as-a-judge methodologies to measure and improve model quality.
  • Prototype, validate, and productionize ML techniques, including fine-tuning, distillation, retrieval optimization, and agent architectures.
  • Evaluate emerging foundation models and AI research, rapidly translating promising advances into production experiments.
  • Contribute to the broader AI community through technical reports, or conference talks.

Requirements

What you’ll bring

  • 5+ years of experience in applied machine learning, NLP, large language models, or AI research, with a proven track record of shipping production AI systems.
  • Master's degree in Computer Science, Machine Learning, Artificial Intelligence, NLP, or a related field required; PhD is a strong plus.
  • Deep experience with modern LLMs and agent architectures.
  • Strong software engineering skills with production-quality Python and experience using modern AI frameworks.
  • Demonstrated ability to design rigorous evaluation frameworks, benchmark models, and make data-driven decisions balancing quality, latency, cost, and reliability.
  • Demonstrated ability to improve model cost/quality through experimentation—including retrieval optimization, fine-tuning, distillation, inference optimization, and agent design.

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.