Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. Backed by Battery.
About the role
This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space.
What they're looking for
- Active Top Secret (TS/SCI) security clearance
- 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems
- Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes
- Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD)
- Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
- Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
More about this role
This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space.
This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you.
Note: this role includes on-call responsibilities.
- Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed...
Browse similar: Startup jobs · Remote jobs