Platform Reliability Engineer
Northern, Floyd County, Kentucky, USA
Listed on 2026-10-08
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Reports To:
Platform Reliability Engineering Manager
Location:
Remote (within the U.S.)
Environment:
Remote
Status:
Exempt;
Salaried
Recognized by Gartner in their Modern 4PL Market Guide, Redwood Logistics is at the forefront of industry innovation. Our cutting-edge supply chain technology pairs with the expertise of our brilliant minds to empower logistics execution across North America and Mexico.
Leveraging a comprehensive range of services, data‑centric network solutions, and a seamlessly integrated platform, we have established our prominence as a key player in the mid‑market segment within the freight tech industry.
Whether you’re just starting your career or are an established professional looking for your next opportunity, Redwood inspires innovation across teams to provide transformative solutions for our customers.
Purpose of Your Work:As a Platform Reliability Engineer, you will design, build, and evolve Redwood's internal engineering platform, enabling product teams to deliver secure, scalable, and highly reliable software. You will create shared cloud infrastructure, Kubernetes platforms, infrastructure automation, CI/CD, observability, and developer self‑service capabilities while driving operational excellence across the engineering organization.
Working closely with Software Engineering, Architecture, Security, Data Engineering, and MLOps, you will establish engineering standards, improve platform reliability, eliminate operational toil through automation, and enable developers to deliver software faster with confidence.
How You Make a Difference Everyday:- Design and maintain reusable Terraform modules, Helm charts, Git Ops templates, and shared platform services.
- Develop self‑service platform capabilities and golden paths that accelerate software delivery.
- Design, standardize, and evolve enterprise CI/CD pipelines incorporating automated testing, security validation, deployment governance, and release automation.
- Engineer highly available Kubernetes platforms with strong security, networking, scalability, and operational practices.
- Drive platform observability through metrics, logs, traces, dashboards, alerting, SLIs, SLOs, and error budgets.
- Support engineering response during production incidents, perform post‑incident reviews, and implement permanent reliability improvements.
- Continuously eliminate manual operational work through automation.
- Partner with engineering teams to improve application reliability, resilience, performance, and operational readiness.
- Optimize cloud cost, platform efficiency, and resource utilization.
- 5+ years in Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure.
- Production experience with Azure and/or AWS.
- Production Kubernetes experience (AKS, EKS, or GKE).
- Infrastructure as Code using Terraform or equivalent.
- Modern CI/CD platforms (Git Hub Actions, Azure Dev Ops, or equivalent).
- Git Ops using ArgoCD or Flux.
- Helm chart development.
- Linux systems administration and container technologies.
- Scripting using Power Shell, Bash, Python, or Go.
- Enterprise observability platforms such as Datadog, Prometheus, Grafana, Logic Monitor, or New Relic.
- Cloud networking and identity technologies including OIDC.
- Strong troubleshooting, communication, and cross‑functional collaboration.
- Strong communication and collaboration skills with experience partnering across Software Engineering, Cloud Infrastructure, Architecture, Security, Data Engineering, and MLOps teams.
- Access to experts and resources for your Learning & Development journey
- Opportunity for internal mobility
- Employee referral bonus program
- Employee Resource Groups (ERGs)
- Annual fundraising and volunteer…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).