Startups

Senior Site Reliability Engineer

NeuroFlow · Remote · Remote

← All jobs
About NeuroFlow

Who We Are NeuroFlow CEO and West Point graduate Christopher Molaro served in the army for five years, including a tour in Iraq as a platoon leader. Backed by MedTech Innovator.

About the role

As our Senior Site Reliability Engineer, you'll own how reliable our production systems are and how we measure it. You'll set our SLOs, lead incident response, build the tooling that lets teams ship and experiment safely, and keep our infrastructure secure and compliant across AWS, Azure, and our data centers. You'll work with engineering, security, and product partners to set the standard for how we build and ship. As a HIPAA compliant organization, NeuroFlow expects all team members to: Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • Define SLOs and SLIs with stakeholders inside and outside engineering, back them with error budgets, and use them to decide when to ship and when to slow down
  • Run and improve our SaaS monitoring, alerting, and logging stack (Dynatrace, New Relic, Datadog, or similar) so we meet compliance requirements and catch problems before customers do
  • Take part in on-call for production incidents and help engineers work through customer issues
  • Lead blameless postmortems. When a failure keeps showing up, find the root cause and fix it for good
  • Find the manual, repetitive work that eats engineering time, measure what it costs, and automate it away. Track and report the reduction
  • Leave behind automation and documentation that survives a change of owner
More about this role

About the Role

As our Senior Site Reliability Engineer, you'll own how reliable our production systems are and how we measure it. You'll set our SLOs, lead incident response, build the tooling that lets teams ship and experiment safely, and keep our infrastructure secure and compliant across AWS, Azure, and our data centers. You'll work with engineering, security, and product partners to set the standard for how we build and ship.

What You'll Do

Reliability & Incident Response

  • Define SLOs and SLIs with stakeholders inside and outside engineering, back them with error budgets, and use them to decide when to ship and when to slow down.
  • Run and improve our SaaS monitoring, alerting, and logging stack (Dynatrace, New Relic, Datadog, or similar) so we meet compliance requirements and catch problems before customers do.
  • Take part in on-call for production incidents and help engineers work through customer issues.
  • Lead blameless postmortems. When a failure keeps showing up, find the root cause and fix it for good.
  • Find the manual, repetitive work that eats engineering time, measure what it costs, and automate it away. Track and report the reduction.
  • Leave behind automation...

Read the full posting on NeuroFlow's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.