Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. Over 180 million clinical interactions across 1,000+ use cases with 60+ partners worldwide. Backed by General Catalyst, Kleiner Perkins and a16z.
About the role
Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)
What they're looking for
- Master's degree in Computer Science, Machine Learning, or a related field
- 5+ years of professional experience in NLP, LLM training, or reinforcement learning
- 2+ years of hands-on experience with RL for LLM post-training
- Proficiency in Python and PyTorch for large-scale training
- Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods
- Experience training or fine-tuning models at scale (50B+ parameters)
More about this role
As HAI's LLM Post-Training Applied Scientist, you will own the reinforcement learning and on-policy distillation pipeline that transforms raw model capability into reliable, safe clinical behavior. Your post-training methods will directly determine how our AI agents reason through complex clinical scenarios, handle safety-critical decisions, and ultimately impact millions of patient interactions. This role exists because post-training is where capability becomes trustworthiness—and in healthcare, that's everything.
Own your first major outcome: By day 90, you will have shipped a post-training improvement that meaningfully advances model performance on a critical clinical capability (clinical reasoning, safety alignment, or task completion), evaluated the gains rigorously, and contributed that learning to our post-training roadmap.
Drive lasting impact: At 12 months, you will have designed and shipped multiple post-training methods that measurably improve our models' clinical safety and reasoning, built reusable infrastructure (reward models, verifiers, evaluation frameworks) that accelerate future post-training work, published your research or contributed to HAI's intellectual...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area