Backed by Accel, Greylock and Norwest.
About the role
Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations). In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe's machine learning and AI systems reliable, scalable, and safe in production.
What they're looking for
- Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure
- Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems
- Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies
- Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks
- Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents
- Drive operational strategy for GoFundMe's generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability
More about this role
GoFundMe is the world’s most powerful community for good, dedicated to helping people help each other. By uniting individuals and nonprofits in one place, GoFundMe makes it easy and safe for people to ask for help and support causes – for themselves and each other. Together, our community has raised more than $40 billion since 2010.
Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations). In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe's machine learning and AI systems reliable, scalable, and safe in production. This role requires strong technical judgment across the ML lifecycle (data → training → online inference → monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure.
Candidates considered for this role will be located in the San Francisco Bay Area. There will be an in-office requirement of 3x a week.
- Own the reliability, scalability, and operational health of ML/AI production systems...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area