Production System Engineer
Listed on 2026-09-24
-
IT/Tech
IT Infrastructure, Systems Engineer, Unix/Linux, SRE/Site Reliability
Responsibilities
About the Team
The Server Management Dev Ops team is responsible for the end-to-end lifecycle management of servers across Byte Dance’s self-built data centers in the United States and Europe. Our scope covers new hardware introduction, data center delivery, production operations, hardware maintenance, configuration and firmware changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse. The team serves as a central coordination point between multiple functions, including:
- Hardware New Product Introduction (NPI)
- Server and data center operations
- Field maintenance and infrastructure management
- Hardware vendors and service providers
- Supply chain and asset management
- Infrastructure platform and automation engineering teams
Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.
Role OverviewWe are looking for a hands‑on Production Systems Engineer with a strong foundation in Linux systems, server infrastructure, automation, and production operations. This role is open to engineers across a range of experience levels, from early-career engineers with strong technical fundamentals to experienced infrastructure engineers who can take ownership of complex systems and large-scale initiatives.
The scope and level of ownership will grow with experience, ranging from hands‑on infrastructure engineering and automation development to leading complex global infrastructure initiatives across organizational boundaries.
Responsibilities- Server Infrastructure Operations:
Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers. - Automation Development:
Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency. - Linux Systems:
Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues. - GPU and AI Infrastructure:
Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements. Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making. - Monitoring and Data Analysis:
Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement. - Technical Documentation:
Create and improve technical documentation, standard operating procedures, troubleshooting guides, and internal knowledge bases. - Cross-functional Collaboration:
Work with infrastructure engineers, hardware teams, data center operations, platform developers, supply chain teams, and other stakeholders on global infrastructure projects.
- Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
- 2 years of experience in systems engineering, infrastructure operations, Dev Ops, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
- Strong foundation in Linux system administration and troubleshooting, with an understanding of basic server architecture, operating systems, storage, networking, and hardware management concepts.
- Programming or scripting experience in Python, Bash, Go, or another modern programming language, with the ability to develop tools or automation for infrastructure or operational tasks.
- Hands‑on…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).