SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences.
About the role
This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability. You will partner closely with Applied Science, AI Platform, Product, and Partner Engineering to enable rapid research iteration and reliable model delivery at scale.
What they're looking for
- Build and operate SPREEAI’s end-to-end ML platform spanning training, evaluation, deployment, and monitoring
- Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems
- Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation
- Enable reliable and scalable inference deployments through standardized serving, orchestration, and monitoring frameworks
- Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability
- Establish production SLOs for latency, availability, error rate, GPU saturation, cold-start time, cost per inference, and model quality drift
More about this role
About the Role
SPREEAI is building the future of AI-powered commerce through photorealistic virtual try-on and multimodal intelligence. We bring together cutting-edge AI and real-world retail to deliver production systems that redefine how people shop online.
We are looking for a Principal Engineer to build the infrastructure, deployment pipelines, and observability systems that enable multimodal AI models to move from research prototypes to reliable, production-grade deployments powering real-time virtual try-on experiences for global retail partners.
This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability. You will partner closely with Applied Science, AI Platform, Product, and Partner Engineering to enable rapid research iteration and reliable model delivery at scale.
What You'll Own
ML Platform & Training Enablement
- Build and operate SPREEAI’s end-to-end ML platform spanning training, evaluation, deployment, and monitoring.
- Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems.
- Define platform standards for model packaging, model registry, dataset lineage,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area