Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. Over 180 million clinical interactions across 1,000+ use cases with 60+ partners worldwide. Backed by General Catalyst, Kleiner Perkins and a16z.
About the role
We are building a recursive self-improvement system — a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives.
What they're looking for
- Built or owned part of a feedback loop — a reward model, an evaluation harness, or the data pipeline for an RLHF/RLAIF or active-learning system
- Ran a retraining or continual-learning pipeline where a model consumed its own predictions or production data (e.g. ranking, recommendations, fraud, spam)
- Fine-tuned LLMs with human or AI feedback, or built agentic evaluation harnesses
- The ideal candidate has built or shipped a full system that improved from its own outputs or feedback end to end. This is rare at this level, so treat it as a standout differentiator rather than a filter. Examples:
- RLHF / RLAIF pipelines
- Self-play systems
More about this role
We are building a recursive self-improvement system — a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives.
This is an engineering-first role with deep reinforcement learning requirements. You should be equally comfortable writing robust production ML code and reasoning about reward design, credit assignment, and why feedback-driven systems become unstable.
Build and maintain the training, evaluation, and deployment loops at the core of the self-improvement system, with a strong emphasis on reproducibility and reliability.
Design and implement reward and feedback signals; investigate and mitigate reward hacking, specification gaming, and distribution drift.
Build evaluation harnesses and metrics before models — because a self-improving system is only as safe as its measurement of “better.”
Own data pipelines and automated data flywheels that...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area