Systems Development Engineer II, AWS Managed Operations; MO AWSOM Team
Listed on 2026-07-21
-
Software Development
DevOps, Cloud Engineer - Software, AI Engineer (Applied/Software), AWS
Systems Development Engineer II, AWS Managed Operations (MO) AWSOM Team
Job : | Amazon Development Center U.S., Inc.
Do you love decomposing problems to develop products that impact millions of people around the world? Would you enjoy identifying, defining, and building software solutions that revolutionize how businesses operate? Whether you thrive on diving deep into operating and improving some of the largest software systems, or the challenges of driving technical, business, and cultural change to improve reliability, performance, and efficiency excite you, we’re looking for you.
The AWS Managed Operations (MO) organization, founded in April 2023, aims to reduce operational load and toil through long‑term engineering projects. Managed Operations is building a best‑in‑class engineering and operations team that owns day‑to‑day operations for AWS regions, enhancing availability, reliability, latency, performance, and efficiency.
Amazon seeks highly motivated Systems Development Engineers who balance day‑to‑day operations of AWS software systems with long‑term engineering to reduce operational toil. You will learn continuously and dive deep into the systems and technologies that power one of the world’s largest cloud providers.
Key Job Responsibilities- Collaborate across diverse teams, projects, and environments, impacting our global customer base.
- Bring passion for innovation, data, search, analytics, and distributed systems.
- Build solutions to innovate agentic‑first interfaces.
- Solve challenging technical problems, often unprecedented, at every layer of the stack.
- Leverage Generative AI tools and AI‑assisted development environments to accelerate prototyping, development, validation, and testing workflows.
- Use AI‑powered code generation and review tools to rapidly iterate on infrastructure automation scripts, configuration templates, and system tooling; automate investigative, diagnostic, or operational tasks.
- Apply GenAI capabilities to accelerate test case generation, integration testing, and validation of complex infrastructure components—reducing cycle times while maintaining quality and security rigor.
- Rapidly synthesize technical documentation, compliance frameworks, and architectural patterns using GenAI tools to accelerate research and decision‑making in ambiguous problem spaces.
- Understand responsible use of GenAI in security‑sensitive environments, including data handling boundaries, model limitations, and appropriate human‑in‑the‑loop validation practices.
You will split your time approximately 50/50 between operating production systems and driving long‑term improvements to the reliability, availability, and performance of those systems.
Typical Week- Root cause analysis & remediation—investigate production issues such as deployment failures, identify underlying bugs, and implement fixes to restore system health.
- Systems‑level problem solving—identify patterns across incidents, design solutions addressing entire classes of problems, and collaborate to refine designs.
- SLO stewardship—evaluate the effectiveness of Service Level Objectives, partner with stakeholder teams to validate thresholds, and update infrastructure‑as‑code to keep monitoring meaningful and actionable.
- Systems automation for operational excellence—design and build systems that improve fleet management, such as safely migrating workloads to more optimal hardware types, delivering measurable gains in performance.
Our team is dedicated to supporting new members, fostering knowledge sharing and mentorship. Senior members provide mentorship and thorough yet kind code reviews. We enable career growth and assign projects that help members develop engineering expertise, empowering them to take on more complex tasks.
Basic Qualifications- 3+ years of designing or architecting systems (design patterns, reliability, scaling) for new and existing systems.
- 3+ years of administrative experience in networking, storage systems, operating systems, and hands‑on systems engineering.
- Experience programming with at least one modern language:
Python, Ruby, Golang, Java, C++, C#, Rust. - Ex…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).