Startups · AI

Staff Engineer, Machine Learning Systems & Reliability - Moveworks

ServiceNow · Mountain View, CALIFORNIA · On-site

← All jobs
About ServiceNow

Moveworks : the Agentic AI Assistant platform that empowers the entire workforce. Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Backed by Greylock and Sequoia.

About the role

We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems.

What they're looking for

  • A track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure
  • Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust
  • Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization
  • Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability
  • Experience distinguishing service-health problems from data-quality or model-quality problems
  • Familiarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems
More about this role

We are building AI-enabled product capabilities that improve through data, feedback, and real-world use. We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale.

We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems.

This role sits at the intersection of ML systems, platform engineering, and site reliability engineering. You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production—and take ownership of how those systems perform and evolve once deployed.

  • Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining.
  • Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety,...

Read the full posting on ServiceNow's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.