⚡️ Why Altium? Altium is transforming the way electronics are designed and built. From startups to world’s technology giants, our digital platforms give more power to PCB designers, supply chain, and manufacturing, letting them collaborate as never before.
About the role
Senior Site Reliability Engineer ensures the reliability, availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks, improving observability, and contributing to incident management, while also collaborating with development and technology teams to build more reliable and scalable applications.
What they're looking for
- 6+ years in SRE, DevOps or related role in a large-scale environment
- Software development experience (ideally working with and as a .NET developer)
- Strong understanding of SDLC, microservice and HA architecture
- Observability - NewRelic, ELK, Grafana, PagerDuty, OTEL or similar
- Experience with Kubernetes clusters in production setting, AWS, IOC
- Experience with operational tasks
More about this role
Senior Site Reliability Engineer ensures the reliability, availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks, improving observability, and contributing to incident management, while also collaborating with development and technology teams to build more reliable and scalable applications.
Join Altium as a Senior Site Reliability Engineer to ensure the reliability and performance of the Altium Cloud Platforms.
- Understanding how an Altium Cloud Platform works
- Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection.
- Develop and implement reliability frameworks and patterns that standardize and elevate the resilience of our SaaS products across multiple regions and environments.
- Cultivate a shared responsibility model where the SRE team collaborates with and educates engineering teams on reliability best practices.
- Contribute to incident response and management, ensuring rapid resolution, clear stakeholder communication, and...
Browse similar: Startup jobs