More jobs:
Sr. System Development Engineer, AL/ML/Storage server team; AWS
Job in
Cupertino, Santa Clara County, California, 95014, USA
Listed on 2026-07-18
Listing for:
Amazon
Full Time
position Listed on 2026-07-18
Job specializations:
-
Software Development
Unix/Linux, Software Engineer
Job Description & How to Apply Below
Senior Systems Development Engineer – Services LLC
Key Job Responsibilities Fleet Health & Predictive Infrastructure- Build and own the automation infrastructure responsible for the health of the server fleet across storage and accelerator (AI/ML) compute platforms
- Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact
- Drive toward zero-touch operations – building automation that detects, diagnoses, triages, and remediates hardware and software faults without human intervention
- Develop monitoring tools, dashboards, and alerting systems to provide real‑time visibility into fleet health across lab and production environments
- Define and track fleet health metrics including failure rates, mean time to detect, mean time to repair, first‑time fix rate, and predictive accuracy
- Debug and resolve complex system‑level issues across storage, compute, GPU, and networking in production environments
- Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including PCIe, power, NIC, NVMe, and GPU subsystems
- Perform root cause analysis on hardware failures – correlating firmware, kernel, driver, and physical layer to isolate faults
- Build diagnostic tooling that automates root cause identification and reduces reliance on manual triage
- Improve manufacturing throughput and yield through test optimization
- Lead the definition and development of software, automation, and enabling tools for server hardware programs; track and report progress
- Design and build scalable system‑level software focused on durability, availability, security, and diagnostics
- Develop and maintain device drivers for Linux on ARM and x86 architectures
- Build automation solutions using modern programming languages such as Python, Ruby, Java, C/C++ and related tools
- Work with OS internals, storage subsystems, and accelerator/GPU software stacks in Linux‑based environments
- Build, manage, and deploy CI/CD pipelines for rapid deployment of code changes to both organization‑owned and customer‑owned systems
- Work across internal HWEng teams to ensure new server hardware addresses data path and control path functionality needed by dependent service teams
- Collaborate with internal customers to identify early potential problems when onboarding new servers into their ecosystem
- Engage with ODMs and design partners on testability, diagnostic, and automation requirements during hardware design and development
- Contribute to server design to improve robustness, testability, diagnosability, and reliability
- Partner with datacenter operations teams to close the loop between field failures and design improvements
Systems Development Engineers in AWS Hardware Engineering wear many hats. From orchestration tooling development to hardware integration to kernel driver debugging, we dive deep into problems across the breadth of AWS. Our teams are directly responsible for launching and maintaining server hardware in the fleet – including storage servers powering distributed storage platforms and AI/ML accelerator servers with GPUs.
Basic Qualifications- 6+ years of professional software development experience
- 6+ years of systems design, software development, operations, automation, and process improvement experience
- 6+ years of designing or architecting new and existing systems (patterns, reliability, scaling)
- 5+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, Power Shell, or Ruby
- Experience with Linux/Unix environments
- Experience leading the design, build, and deployment of complex and performant (reliable and scalable) software solutions in production
- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
- Experience taking a leading role in building complex software or computing…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×