Startups · AI

Staff Software Engineer, Site Reliability Engineering, Vertex 1P GenAI SRE

Google · United States · On-site

← All jobs
About Google

Learn more about Google. Explore our innovative AI products and services, and how we. Backed by Kleiner Perkins.

About the role

Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success.

What they're looking for

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience
  • 8 years of experience with software development in one or more programming languages
  • 2 years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering roles managing large - scale infrastructure
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems
  • 3 years of experience in machine learning infrastructure
More about this role
  • Develop scalable and sustainable system architecture and designs for products, services and enhancements.
  • Defend performance of critical services in alignment with customer expectations and SLOs.
  • Own and define strategy and set direction and establish roadmaps for Vertex AI services to increase reliability, efficiency and ultimately feature velocity.
  • Resolve outages or service disruptions and help design solutions to ensure systems are protected from similar classes of problems in the future.
  • Collaborate with development counterparts to incorporate and deliver enhancements to systems resulting in improved reliability, scalability and or performance.
  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 2 years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering roles managing large - scale infrastructure.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.
  • 3 years of experience in machine learning infrastructure.
  • Master's degree in Computer Science or...

Read the full posting on Google's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.