Trusted AI for critical infrastructure to deliver secure, explainable agents across cloud, on-prem, and edge with rapid time to impact.
About the role
We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production.
What they're looking for
- Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics
- Define and operate SLIs/SLOs for AI APIs and critical services, covering availability, latency, errors, throughput, model quality, output safety, user-facing correctness, and cost
- Develop actionable observability through metrics, logs, traces, dashboards, and SLO-based alerts
- Participate in a sustainable on-call rotation, lead incident response, improve runbooks, and facilitate blameless postmortems
- Reduce operational toil and improve resilience through infrastructure as code, automation, and disaster-recovery planning
- Design automated load, stress, spike, soak, and scalability tests that model realistic AI production traffic
More about this role
We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production. You will partner closely with AI/ML, Platform, Security, and Product teams to build dependable systems that are performant, resilient, and ready to scale.
The Impact
You will help define how Seekr operates mission-critical AI systems in production. Your work will make releases safer, incidents less frequent and easier to resolve, and service health more visible and measurable. By establishing meaningful SLOs, strengthening observability, improving resilience, and reducing operational toil, you will directly improve the experience of our customers and engineering teams.
Duties and Responsibilities
- Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics.
- Define and operate SLIs/SLOs for AI APIs and...
Browse similar: AI jobs · AI startup jobs · Startup jobs