Tencent logo

    Site Reliability Engineer Intern — AI Infrastructure

    Tencent

    Singapore, SingaporeInternship6 Oct 2026

    About this internship

    About TencentTencent is an Internet-based platform company founded in Shenzhen, China, in 1998. We use technology to enrich the lives of Internet users and assist the digital upgrade of enterprises. Our mission is “Value for Users, Tech for Good.” We embrace a culture of teamwork and creativity and are driven by our values of Integrity, Proactivity, Collaboration, and Creativity. We are rapidly expanding our international operations and are looking for top talent to propel us forward. Combining the results-oriented nature of a start-up with the resources of a profitable and leading Internet company, Tencent offers a unique opportunity for aspiring individuals to thrive. Role SummaryWe are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations and maintenance of AI-accelerated high-performance computing (HPC) infrastructure. In this role, you will work closely with internal business team and vendor's engineering team to build and operate GPU clusters. This is a hands-on opportunity to gain direct exposure to cutting-edge AI infrastructure. Team IntroductionThe AI Compute Centre sits within Tencent's Overseas IT department, acting as the bridge between internal AI/GPU infrastructure demand and the external resources that fulfil it. We play key roles in the full lifecycle of GPU clusters — from requirement gathering and capacity planning, through architectural design and development, to delivery, operations, and DevOps — across regions worldwide. Key ResponsibilitiesSupport the deployment, configuration, and maintenance of high-end GPU servers, storage servers, networking equipment, and software components in secure environments.Assist with hardware diagnostics, system functionality checks, and firmware updates as required.Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, HPC clusters, Kubernetes, Slurm, etc.).Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance.Document incident details, resolutions, and lessons learned to improve future problem-solving.Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team.Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. Skills & ExperienceCurrently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field.Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards.Exposure to scripting languages such as Bash or Python.Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes, NVIDIA BCM), and observability tools (e.g., Prometheus, Grafana, ELK).Assist senior engineers in checking GPU cluster network connectivity, including InfiniBand or RoCE links. Run predefined tests, collect results, and flag link or performance anomalies.Monitor dashboards for compute, storage, and management networks. Help record bandwidth, latency, and RDMA test results, and compare them with agreed baselines.Support AI infrastructure incident investigations by collecting switch, NIC, and host logs, updating incident records, and tracking vendor follow-up actions under guidance.Ability to work both independently and as part of a team.