Senior Cloud Engineer; Hybrid @ Bellevue, WA or Remote @ Florida
Bellevue, WA or Remote Florida
The Cloud Engineering team is dedicated to maintaining and evolving cloud infrastructure, CI/CD pipelines, Application Security, Change Management, and Backend application golden templates. We are the backbone of the technical organization, ensuring the foundation of our services is robust, secure, and efficient.
About the roleWe are looking for a Senior Cloud Engineer to join our Cloud Engineering team. This team is tasked with ensuring scalability, reliability, and security for all backend services and keeping the SDLC running smoothly for engineers across the company.
As a senior member of the team, you will own meaningful pieces of our cloud platform end-to-end — from design through deployment and operation — while helping shape the standards, tooling, and golden templates that the rest of the engineering organization builds on. You will balance deep, hands‑on technical ownership with mentoring teammates and influencing the technical direction of the systems that power a large‑scale mobile marketplace.
Whatwe love about this role
- Scale & Impact:
You will work on systems that drive the end‑user experience and support every technical team across the organization, operating across a multi‑account AWS and GCP footprint. - Ownership:
You will own critical infrastructure and platform capabilities end‑to‑end, making architectural decisions that affect reliability, security, and cost at scale. - Force Multiplier:
The templates, pipelines, and tooling you build raise the productivity and safety of every engineer at Offer Up. - Emerging Technology:
You will help bring new capabilities — including AI/ML developer tooling and infrastructure — into the hands of our engineers.
- Build cloud infrastructure as code — design, implement, and maintain AWS and GCP infrastructure with Terraform, optimizing for reliability, security, and cost.
- Own our Kubernetes & Git Ops platform — run and evolve EKS/GKE (cluster upgrades, Envoy/Gloo networking, Helm, ArgoCD delivery).
- Keep the SDLC fast and safe — build CI/CD pipelines and developer workflows (Git Hub Actions, monorepo tooling, JFrog Artifactory), and maintain the backend “golden templates” and shared libraries the wider org builds on.
- Lead platform migrations end‑to‑end — plan, execute, and de‑risk observability, data‑store, Kubernetes, and multi‑account migrations, including rollback strategy.
- Drive operational excellence & security — own observability in Datadog and on‑call, remediate CVEs, manage secrets and access controls (Cloudflare/ZTNA, Okta, SSO, Workload Identity Federation), and optimize cloud cost.
- Multiply the team — enable AI/ML developer tooling (Amazon Bedrock, Claude, AI‑assisted workflows), lead design and code reviews, set technical standards, mentor engineers, and partner with other teams.
- Experience &
Education:
5–8 years in cloud engineering, infrastructure, platform, Dev Ops, or SRE roles;
Bachelor’s in Computer Science or a related field, or equivalent practical experience. - Cloud & Infrastructure as Code: Deep, hands‑on AWS (ideally GCP) — EKS/GKE, Lambda, networking, IAM, S3, Kinesis, Open Search, Route
53 — managed declaratively with Terraform at scale. - Containers & Delivery: Solid experience with Kubernetes (EKS/GKE), Helm, Git Ops (ArgoCD), and CI/CD pipelines (Git Hub Actions or similar) in a monorepo or multi‑service environment.
- Observability & Operations: Modern observability tooling (Datadog or similar) for metrics, logs, traces, and alerting; supporting production systems, on‑call, and debugging complex distributed systems.
- Coding: Proficiency in at least one of Java, Type Script, Python, or Go for automation, tooling, and platform development; comfort with scripting and SQL.
- Communication & Leadership: Able to communicate technical concepts to technical and non‑technical audiences, lead through influence, and mentor other engineers.
- Experience with edge and zero‑trust networking (Cloudflare, Envoy/Gloo).
- Experience operating data and streaming infrastructure (Kinesis, Flink, Redis/Valkey, Confluent/Kafka).
- Experience enabling AI/ML…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).