# AI Site Reliability Engineer at Seekr

- Company: Seekr
- What the company does: Trusted AI for critical infrastructure to deliver secure, explainable agents across cloud, on-prem, and edge with rapid time to impact.
- Company website: https://www.seekr.com
- Type: Startups (AI role)
- Level: Mid level
- Location: Austin, Texas, United States; Reston, Virginia, United States
- Work setup: On-site
- Posted: 2026-08-21
- Apply by: 2026-10-08
- Apply: https://www.seekr.com/careers/?gh_jid=4726415005
- Page: https://www.1752.vc/careers/jobs/seekr-ai-site-reliability-engineer/

## About the role

We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production.

## What they're looking for

- Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics
- Define and operate SLIs/SLOs for AI APIs and critical services, covering availability, latency, errors, throughput, model quality, output safety, user-facing correctness, and cost
- Develop actionable observability through metrics, logs, traces, dashboards, and SLO-based alerts
- Participate in a sustainable on-call rotation, lead incident response, improve runbooks, and facilitate blameless postmortems
- Reduce operational toil and improve resilience through infrastructure as code, automation, and disaster-recovery planning
- Design automated load, stress, spike, soak, and scalability tests that model realistic AI production traffic

Tags: DevOps
