Startups · AI

MLOps Engineer

SpreeAI · Hybrid (San Francisco, California, US) · Hybrid

← All jobs
About SpreeAI

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences.

About the role

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

What they're looking for

  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform
More about this role

AI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU...

Read the full posting on SpreeAI's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.