Senior Platform Engineer; Core Infrastructure
Job in
Oklahoma City, Oklahoma County, Oklahoma, 73116, USA
Listed on 2026-09-03
Listing for:
Lambda
Full Time
position Listed on 2026-09-03
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
- Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance
- Architect, deploy, and operate Kubernetes clusters across AWS and Lambda’s bare-metal datacenters
- Build and maintain automation for cluster lifecycle management — provisioning, upgrades, and scaling
- Own the reliability, performance, and security of Kubernetes workloads in production
- Implement observability, logging, and alerting for clusters and critical workloads
- Partner with product teams to design scalable, cloud-native services and CI/CD pipelines
- Set the standards for resource management, networking, and RBAC across the platform
- Lead incident response, root-cause analysis, and post-mortems for platform issues
- Mentor engineers and raise the bar for platform engineering across the org
- Health, dental vision
- 401k match
- 5 sick days and 12 paid holidays
- Flexible PTO
- Paid parental, medical, and caregiver leave
Solid grounding in networking, service meshes, and container runtimes
Proficient with infrastructure-as-code (Terraform, Pulumi, or equivalent)
Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting)5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale
Strong with Helm, Kustomize, or similar, and Git Ops-based delivery
Practical security experience: network policies, secrets management, and image scanning
Strong coding skills in Go or Python for automation and tooling
Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure
Experience with multi-cluster, multi-cloud, or hybrid environments
Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows)
Exposure to cost optimization and capacity planning for large clusters
Contributions to CNCF or Kubernetes open-source projectsCKA/CKS certification#J-18808-Ljbffr
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×