×
Register Here to Apply for Jobs or Post Jobs. X

Production System Engineer Graduate; Server Management

Job in New York, New York County, New York, 10261, USA
Listing for: ByteDance
Full Time position
Listed on 2026-09-09
Job specializations:
  • IT/Tech
    IT Infrastructure, Systems Engineer, Unix/Linux, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Production System Engineer Graduate (Server Management) - 2027 Start
Location: New York

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

New York

Team:

Infrastructure

Employment Type:

Regular

Job Code:

A55465

Share this listing:

Responsibilities

The Server Management Dev Ops team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in the United States and Europe. Our scope covers the complete server lifecycle, including new hardware introduction, data center delivery, production operations, hardware maintenance, configuration changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.

We are looking for a motivated Production Systems Engineer who is passionate about Linux systems, server hardware, automation, AI infrastructure, and large-scale data center operations. As a Production Systems Engineer, you will work alongside experienced infrastructure engineers on real production challenges involving large-scale server fleets, GPU infrastructure, automation platforms, and AI-assisted operational tools. You will contribute to building automation, improving operational efficiency, troubleshooting production issues, and supporting the lifecycle management of servers deployed across Byte Dance's global data centers.

Key Responsibilities
  • Server Infrastructure Operations:
    Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development:
    Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems:
    Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure:
    Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements.
  • Monitoring and Data Analysis:
    Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • AI for Infrastructure Operations:
    Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.
  • Strong analytical and troubleshooting skills with the ability to learn unfamiliar technologies quickly.
  • Good communication skills and the ability to collaborate effectively in cross-functional engineering teams.
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • Experience in systems engineering, infrastructure operations, Dev Ops, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
  • Solid understanding of Linux system (Debian or Ubuntu preferred) administration and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language.
  • Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.
Preferred Qualifications
  • Experience developing automation tools or infrastructure software using Python, Bash, Go, or similar languages.
  • Working with server hardware, PC building, homelabs, or data center infrastructure.
  • Expe…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary