×
Register Here to Apply for Jobs or Post Jobs. X

Senior Manager, Site Reliability Engineering – Paylo Platform

Job in Temple, Bell County, Texas, 76501, USA
Listing for: PDI Technologies
Full Time position
Listed on 2026-08-18
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Project Manager, Systems Engineer
Job Description & How to Apply Below

Senior Manager, Site Reliability Engineering

PDI Technologies is looking for a Senior Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI's payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world.

This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/Open Tofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.

Key Responsibilities

  • Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20-person SRE organization supporting the Paylo product suite.
  • Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.
  • Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.
  • Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.
  • Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.
  • Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.
  • Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.
  • Own adoption and standards for Git Ops-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.
  • Own the Infrastructure-as-Code strategy across teams (Terraform, Open Tofu), including module standards, state management, drift detection, and remediation.
  • Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.
  • Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.
  • Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.
  • Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.
  • Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.
  • Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.
  • Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.
  • Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.
  • Track and report cost, capacity, and operational KPIs to senior leadership.

Required Qualifications

  • 8+ years of experience in Site Reliability Engineering, Dev Ops, or Infrastructure/Platform…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary