Software Engineer, Distributed Systems
Listed on 2026-08-30
-
Software Development
Cloud Engineer - Software, Backend Developer, DevOps
At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.
Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.
Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.
About The Team AndThe Role
The Observability Platform team builds and operates the infrastructure that helps eBay teams monitor, fix, and improve the reliability of large-scale distributed systems. This platform supports the telemetry and reliability needs of thousands of microservices across eBay and operates at hyperscale, processing billions of time series and petabytes of log data using modern open-source technologies including Prometheus, Click House, Open Telemetry, and related tools.
As a Software Engineer on this team, you will design and build scalable distributed systems that power metrics, logs, traces, and related observability workflows across the stack—from ingestion and storage to query and visualization. You will partner closely with SREs, platform engineers, and service owners to solve complex reliability challenges, improve operational excellence, and help shape the future of observability at eBay.
This role offers the opportunity to work on critical systems, contribute to open-source technologies, and grow through direct exposure to some of eBay’s most complex infrastructure challenges. This role also includes participation in the team’s on‑call rotation in support of production reliability.
What You Will Accomplish- Design and deliver scalable, fault‑tolerant observability infrastructure that improves reliability while reducing operational overhead for platform and engineering teams
- Build and optimize high‑throughput services for ingesting, transforming, storing, and querying telemetry data across logs, metrics, and traces
- Strengthen the resilience of Kubernetes‑based production systems through self‑healing, autoscaling, and robust operational design
- Partner with SREs, platform teams, and service owners to translate observability needs into tools and platform capabilities that improve incident response and operational excellence
- Contribute to architecture reviews, production readiness discussions, and post‑incident findings to drive continuous improvement across the platform
- Expand your technical breadth by working across distributed systems, cloud‑native infrastructure, and optionally user‑facing observability experiences.
- 7+ years of experience in software engineering, infrastructure engineering, or a closely related field
- Strong programming skills in Golang or another systems‑level language, with experience building reliable backend or infrastructure services
- Deep understanding of distributed systems concepts such as fault tolerance, scalability, and system reliability
- Hands‑on experience deploying and operating containerized services in Kubernetes or similar cloud‑native environments
- Solid understanding of observability domains including metrics, logs, and traces
- Experience with tools such as Prometheus, Grafana, Open Telemetry, Click House, or similar technologies; familiarity with time‑series systems, high‑throughput data pipelines, open‑source infrastructure, or React/JavaScript is a plus
The base pay range for this position is expected in the range below:
$172,000 - $229,600
Base pay offered may vary depending on multiple individualized factors, including location, skills, and experience. The total compensation package for this position may also include other elements, including a target bonus and restricted stock units (as applicable) in addition to a full range of medical, financial, and/or other benefits (including 401(k) eligibility and various paid time off benefits, such as PTO and parental leave).
Details of participation in these
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).