Site Reliability Engineer
Listed on 2026-07-17
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Why GMF Technology?
At GM Financial, innovation drives everything we do. We’re not just adopting technology — we’re shaping the future of software delivery. From generative AI and cloud‑native platforms to advanced release engineering practices, our teams are redefining how financial technology operates. This role is central to that transformation, influencing how we build, release, and scale software globally.
Join us and discover a workplace where your ideas matter, your development is prioritized, and you can truly make a global impact.
Job DescriptionLocation: Arlington, TX
Work Arrangement: Hybrid – 2 days onsite, 3 days remote per week
Sponsorship Notice: At this time, we are unable to offer employment sponsorship for this position. This includes, but is not limited to, H‑1B, TN, L1, and OPT visa types.
ResponsibilitiesThe Site Reliability Engineer, under general direction from leadership, will assist in day‑to‑day tasks critical to the team's success. The position will support cloud infrastructure architecture and components, including hybrid cloud and public cloud platforms. Responsibilities include prototyping, initiating, and operationalizing public cloud solutions; supporting overall cloud transformation initiatives; configuring, optimizing, and ensuring performance of deployed public cloud solutions; and extending into advanced automation and software development disciplines.
- Build and demonstrate a foundational understanding of SRE concepts, including observability, monitoring, incident response, and the core systems owned by the team.
- Execute standard operational tasks independently using established processes, runbooks, tooling, and escalation paths; raise issues when scenarios become complex or unfamiliar.
- Perform initial troubleshooting for clear production or environment issues with limited guidance; contribute findings and next steps to the broader resolution effort.
- Demonstrate ownership of learning by seeking mentorship, asking questions, and contributing back to shared team knowledge.
- Help teams apply SRE operational readiness practices using the SRE Checklist—emphasis on detection/observability, performance, resiliency, automation, and operational readiness before go‑live.
- Assist with defining and implementing basic monitoring coverage aligned to Golden Signals (e.g., latency, traffic, errors, saturation/capacity) and validate telemetry appears correctly in monitoring platforms.
- Follow established standards for cloud‑based resources in Azure environment for automation and troubleshooting.
- Support logging and exception‑handling hygiene by aligning to known standards (e.g., ensuring correlation IDs and key dimensions are captured where required).
- Assist and provide systems administration setup/configuration as needed for supported services and environments.
- Contribute to toil reduction by helping implement/maintain repeatable operational mechanisms (e.g., health checks/probes and monitoring configuration) as defined in standards and patterns.
- Thorough command of both the Windows and Linux operating systems, with strong background in troubleshooting either.
- Knowledge of native Kubernetes or related enterprise container platforms such as Open Shift.
- Good understanding of the mechanics of this platform and the deployment pipeline that feeds it.
- Knowledge of public cloud governance frameworks, architectures, configurations, services, and solutions, specifically within Microsoft Azure, but may also include AWS and GCP.
- Knowledge in core Azure services like Azure Kubernetes Service, Cosmos DB, Azure Functions, Azure Storage and concepts, Azure CLI and Power Shell cmdlets.
- Knowledge of Azure organizational entities such as departments, accounts, subscriptions, resource groups, and management groups.
- Strong automation skills in Linux and Windows including bash, Python, and Power Shell.
- Extensive experience with Terraform plans and associated development.
- Knowledge of ARM templates and various related automation methods within Azure.
- Experience with modern source control repositories (e.g., Git) and Dev Ops toolsets (Jenkins, Ansible, etc.) and familiarity with Agile/Scrum…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).