Demo

Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance
Seattle, WA Other
POSTED ON 9/22/2026
AVAILABLE BEFORE 11/22/2026
Our Infrastructure Engineering team supports the company's fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable. We have three subgroups for this role: - Cloud Host Delivery, Delivery & Standardization - Cloud Host Operation, Operation Efficiency & Reliability - Cloud Management & Security Responsibilities - What You'll Do - Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds. - Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure. - Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards. - Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability. - Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

Qualifications


Minimal Qualifications - Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field. - 5 years of experience in Linux operations, SRE, or DevOps; - Proficient in at least one programming language such as Go, Python, or C , with solid engineering capabilities in platform development, system tooling, and automation. - Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills. - Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes. - Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset. Preferred Qualifications - Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms. - Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU. - Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces. - Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines. - Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization. - Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred. - Experience operating large-scale production environments is a strong plus.

Hourly Wage Estimation for Tech Lead Cloud Site Reliability Engineer - DCS Cloud in Seattle, WA
$59.00 to $70.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Tech Lead Cloud Site Reliability Engineer - DCS Cloud?

Sign up to receive alerts about other jobs on the Tech Lead Cloud Site Reliability Engineer - DCS Cloud career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at ByteDance

  • ByteDance York, NY
  • Responsibilities The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in ... more
  • 13 Days Ago

  • ByteDance York, NY
  • Responsibilities The DCS platform team's goal is to build all the necessary and advanced tooling to support full-lifecycle operation and management of dome... more
  • 13 Days Ago

  • ByteDance Seattle, WA
  • About the team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogen... more
  • 14 Days Ago

  • ByteDance Seattle, WA
  • About the Team Join ByteDance’s database R&D team, where you’ll build and own cutting-edge database products supporting ByteDance’s global infrastructure. ... more
  • 14 Days Ago


Not the job you're looking for? Here are some other Tech Lead Cloud Site Reliability Engineer - DCS Cloud jobs in the Seattle, WA area that may be a better fit.

  • Alibaba Cloud Bellevue, WA
  • Alibaba Cloud Computing Platform includes a proprietary big data platform, ODPS (MaxCompute, Hologres, DataWorks, etc.), open-source big data platforms (E-... more
  • 14 Days Ago

  • ByteDance Seattle, WA
  • Our Infrastructure Engineering team supports the company's fast growth by building and operating hyper-scale datacenters, managing the life cycle of server... more
  • 22 Days Ago

AI Assistant is available now!

Feel free to start your new journey!