Startups · AI

Site Reliability Engineer (SRE) Intern — AI Infrastructure

Tencent · US-California-Palo Alto · On-site

← All jobs

About the role

We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure.

What they're looking for

  • Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field
  • Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards
  • Exposure to scripting languages such as Bash or Python
  • Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK)
  • Strong problem-solving and analytical skills
  • Ability to work both independently and as part of a team
More about this role

Role Summary

We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure.

  • Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments.
  • Assist with hardware diagnostics, system functionality checks, and firmware updates as required.
  • Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.).
  • Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.
  • Document incident details, resolutions, and lessons learned to improve future problem-solving.
  • Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team.

-...

Read the full posting on Tencent's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.