Featured Job

Senior Applied Scientist

Palo Alto Full-time On-site 09/12/2026 Job ID: 000194
Apply Now
Reinforcement Learning On-policy Distillation RLHF RLVR LLM-as-judge

Summary

What you’ll impact

The LLM Post-Training Applied Scientist at the organization will own the reinforcement learning and on‑policy distillation pipeline that turns raw model capability into safe, reliable clinical behavior. The role involves designing and implementing post‑training methods, building reward models and evaluation frameworks, and collaborating with multidisciplinary teams to improve model safety and reasoning in healthcare settings.

Responsibilities

What you'll do

  • Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges
  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)
  • Develop conversational AI environments and simulations for healthcare RL training—creating synthetic clinical scenarios and datasets that enable safe, scalable post-training without relying solely on human feedback
  • Automate post-training research loops using agents and tooling to systematically explore hyperparameters, methods, and data strategies—turning post-training into a scientific, reproducible process
  • Run rigorous experiments and analysis to understand what drives post-training gains, isolate the contributions of different components, and build intuition about what works in healthcare contexts
  • Collaborate with research, engineering, and clinical teams to translate clinical requirements into post-training objectives, validate improvements against real-world metrics, and scale successful methods to production

Requirements

What you’ll bring

  • Master's degree in Computer Science, Machine Learning, or a related field
  • 5+ years of professional experience in NLP, LLM training, or reinforcement learning
  • 2+ years of hands-on experience with RL for LLM post-training
  • Proficiency in Python and PyTorch for large-scale training
  • Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods
  • Experience training or fine-tuning models at scale (50B+ parameters)
  • Publications at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
  • Healthcare or regulated domain experience
  • Experience with distributed training frameworks (FSDP, DeepSpeed, vLLM)
  • Familiarity with safety alignment and interpretability research

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.