Sr. Production Engineer
Listed on 2026-07-29
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
It takes powerful technology to connect our brands and partners with an audience of hundreds of millions of people. Whether you're looking to write mobile app code, engineer the servers behind our massive ad tech stacks, or develop algorithms to help us process trillions of data points a day, what you do here will have a huge impact on our business-and the world.
AboutYahoo Platform Experiences
Platform Experiences is a broad Yahoo organization that builds shared experiences and underlying platform capabilities across content, commerce, personalization, and site surfaces. The org includes squads such as Sites, Commerce AI, Catalog, Commerce Delivery System, Merchant Platforms & Data, and Recommendation Engine & Content Understanding. The team sits at the intersection of consumer experience and platform infrastructure. On the experience side, Sites focuses on organizing and presenting content in ways that maximize discoverability, SEO, and impact across properties.
On the platform side, the org owns core systems for catalog ingestion and enrichment, merchant/data platforms, and affiliate or SKU-based product promotion through CDS. Platform Experiences is also in an important expansion phase. The org is taking on former News content hubs, standardizing them on the Platform layout, evolving the Commerce application into the Platform application, unifying the design system on Yahoo OS, and scaling experience configuration and taxonomy to support News verticals.
The Production Engineering (PE/SRE/Dev Ops) team is the power behind engineering goodness for all above. By writing, designing and implementing software and infrastructure to drive velocity, operability, reliability and performance, the PE team ensures continuous quality and compliance on production systems. We also build and maintain tools and platforms for engineer productivity and are at the forefront of tech initiatives of infrastructure modernization with cloud platforms.
We cultivate an environment open to experimentation, where continuous learning and AI-assisted automation are embraced to scale our impact.
Yes, we are Dev Ops! We are all about:
Enabling a culture of ownership and excellence Engineering processes that are Automated and Agile Developing tools that are Self-Serve and (Re)
Usable
Efficiently bring products to market Proactively prevent defects from reaching customers Swiftly address and resolve any production issues Solve problems using automation and software
If you believe in the above, come join us.
A LotAbout You:
The Yahoo commerce team is seeking experienced Dev Ops/Cloud Infrastructure engineers with expertise in AWS and a strong knowledge of Web applications. If you have a strong focus on working in a dynamic environment and possess the required skills, we would love to have you on board with us.
We are looking for Dev Ops, Production Engineers (SRE's) who are problem solvers at heart, with solid ability to dig into code and own the reliability domain. As a senior member of the PE team, you will collaborate with developer partners to establish engineering and operational standards. You will drive reliability improvements, security structures, automation, CI/CD, and intelligent monitoring while helping mentor junior team members and accelerating developer velocity.
Responsibilities:- Analyze current frameworks and infrastructures to oversee high-capacity distributed systems, pinpoint potential enhancements, and establish architectural operational standards.
- Engage with development teams to shape the product trajectory, offering senior-level expertise in scaling, capacity management, security, and operational monitoring.
- Address intricate networking and system difficulties to bolster overall platform stability and performance.
- Leverage AI-assisted coding tools (e.g., Amazon Q) to accelerate the design, implementation, and refinement of Infrastructure as Code (Terraform) scripts and custom system automation tools.
- Execute frontend visibility strategies, including Real User Monitoring (RUM), performance tracking for Core Web Vitals, and managing error budgets using Datadog or Cloud Watch.
- Coordinate with…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).