Software Engineer , Platform
Listed on 2026-07-29
-
IT/Tech
Company description
Welcome to Our World We’ve been leading the charge in the affiliate industry from day one—establishing performance marketing and paving the way for future innovations. We're known for maintaining one of the largest, most reliable partnership platforms with impeccable, personalized service.
Founded in Santa Barbara, California in 1998, CJ (formerly Commission Junction) stands as the most trusted name in performance marketing. We specialize in building partnerships between top brands and reputable publishers to drive revenue and business growth. CJ’s industry-leading solutions make us the platform of choice for over 3,800 global brands across sectors like retail, travel, finance, technology, and home services.
As part of Publicis Groupe, our savvy data capabilities, cutting-edge tech, and strategic expertise facilitate genuine connections, allowing brands to reach consumers wherever they are.
A Quick Peek at Affiliate Marketing Think back to your last online purchase. Did an influencer tip you off about a great product and offer a discount? Or perhaps you relied on a trusted review site to make your decision? Whatever path you took, affiliate publishers likely played a role by influencing, informing, or helping you find the best deal. CJ connects brands with these publishers, creating valuable resources for shoppers like you.
OverviewThis is a hybrid role requiring 3 days a week
You must be work authorized in the United States without the need for employer sponsorship.
Must have Ad Tech / Mar Tech industry experience, specifically in e-commerce, travel, and finance.
As a Software Engineer 3 on the Engineering Experience (Eng Exp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable.
Eng Exp owns all of it. This is not just an infrastructure role - your value is in engineering judgment: how you evaluate systems, detect risk, and make decisions under uncertainty. You'll own meaningful pieces of these systems independently and drive changes from design through production. We want real depth in the systems below, not just familiarity with the tool names.
Eng Exp owns the systems below. You'll own pieces of them independently and be a credible reviewer of changes to them:
- Observability & monitoring — Prometheus, Alert manager, Grafana, and Open Telemetry across production regions. This is not dashboard-building: you'll own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store), and understand Alert manager HA and the blast radius of alert-routing config. Deep Prometheus and Alert manager knowledge is a core requirement.
- Kubernetes & cloud infrastructure — multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management. Spot failure modes before they happen (subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption).
- AWS networking — VPC and subnet design, CIDR management, VPC peering, Route
53, security groups, and NAT gateway topology across accounts and regions, plus the 24/7 networking alarms for prod networking between clusters and squad resources. - CI/CD & artifact management — Git Lab administration (runner fleet, cache, access - not just pipeline authoring), Git Ops delivery through ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows.
- Access & identity — Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management. Turn recurring access requests into self-service workflows that are hard to misuse.
- Cost observability — Open Cost,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).