Startups · AI

Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic AI · Menlo Park, CA · On-site

← All jobs
About Hippocratic AI

Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. Over 180 million clinical interactions across 1,000+ use cases with 60+ partners worldwide. Backed by General Catalyst, Kleiner Perkins and a16z.

About the role

Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)

What they're looking for

  • Master's degree in Computer Science, Machine Learning, or a related field
  • 5+ years of professional experience in NLP, LLM training, or reinforcement learning
  • 2+ years of hands-on experience with RL for LLM post-training
  • Proficiency in Python and PyTorch for large-scale training
  • Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods
  • Experience training or fine-tuning models at scale (50B+ parameters)
More about this role

As HAI's LLM Post-Training Applied Scientist, you will own the reinforcement learning and on-policy distillation pipeline that transforms raw model capability into reliable, safe clinical behavior. Your post-training methods will directly determine how our AI agents reason through complex clinical scenarios, handle safety-critical decisions, and ultimately impact millions of patient interactions. This role exists because post-training is where capability becomes trustworthiness—and in healthcare, that's everything.

Own your first major outcome: By day 90, you will have shipped a post-training improvement that meaningfully advances model performance on a critical clinical capability (clinical reasoning, safety alignment, or task completion), evaluated the gains rigorously, and contributed that learning to our post-training roadmap.

Drive lasting impact: At 12 months, you will have designed and shipped multiple post-training methods that measurably improve our models' clinical safety and reasoning, built reusable infrastructure (reward models, verifiers, evaluation frameworks) that accelerate future post-training work, published your research or contributed to HAI's intellectual...

Read the full posting on Hippocratic AI's site ↗

Research & Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.