SRE Architect
Job in
Albany, Dougherty County, Georgia, 31701, USA
Listed on 2026-07-24
Listing for:
Mphasis
Full Time
position Listed on 2026-07-24
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Job Title – SRE Architect – Thought Leadership & Enterprise Resilience
Mphasis is seeking a visionary SRE Architect to define, champion, and lead enterprise‑scale Site Reliability Engineering practices across its global delivery organization. This is a strategic, principal‑level role at the intersection of cloud architecture, operational excellence, and technical thought leadership. The successful candidate will shape SRE culture, author reference architectures, and partner directly with client C‑suite and VP‑level stakeholders to drive resilience transformation programs.
Yearsof Experience Needed
15+ Years
Role LevelPrincipal / Staff Architect
Enterprise SRE Strategy & Architecture- Design and own the enterprise‑wide SRE operating model, including toil elimination frameworks, error budget policies, and reliability governance structures across multi‑tenant, multi‑cloud environments.
- Architect resilience patterns for mission‑critical platforms spanning Open Shift on‑premise, AWS, and hybrid‑cloud topologies, with focus on fault isolation, graceful degradation, and zero‑downtime deployments.
- Define and institutionalize SLI/SLO/SLA hierarchies at organizational scale; translate business risk appetite into measurable reliability targets with clear ingress and escalation contracts.
- Lead the design of next‑generation CI/CD and Git Ops pipelines; establish blueprint architectures using Git Lab, Argo CD, Tekton, Terraform, and Ansible for repeatable, auditable delivery.
- Establish a Chaos Engineering practice – defining experiment frameworks, blast‑radius controls, and Game Day runbooks – to proactively validate system resilience before incidents occur.
- Architect observability platforms integrating distributed tracing, structured logging, and advanced metric pipelines (Open Telemetry, Prometheus, Grafana, Datadog, Sumo Logic) to enable proactive anomaly detection.
- Drive platform engineering strategy: define internal developer platforms (IDPs), golden paths, and self‑service capabilities that reduce cognitive load for application engineering teams.
- Architect cost‑optimized, auto‑scaling, and Fin Ops‑aligned cloud infrastructure on AWS (preferred) and hybrid environments; lead cloud cost governance programs including Savings Plans, resource tagging taxonomy, and rightsizing automation.
- Design and enforce data security and access control frameworks using AWS HSM, IAM, Secrets Manager, and zero‑trust network principles; engage enterprise security architects to address vulnerabilities surfaced by internal and external audits.
- Define backup, replication, and disaster recovery architectures for critical data and application components; own RTO/RPO targets and validate them through regular DR rehearsals.
- Provide architectural direction for API management, data streaming pipelines, and integration layers (Service Now, Version One, Sumo Logic), ensuring operational observability end‑to‑end.
- Lead capacity management and elastic infrastructure design to accommodate irregular traffic bursts; model growth scenarios and translate them into provisioning and scaling strategies.
- Maintain data integrity and access control using AWS security tools and services such as HSM, IAM, etc.; develop tools to monitor AWS billing, generate cost‑related reports, and implement cost optimization strategies.
- Architect cost‑optimized, auto‑scaling, and Fin Ops‑aligned cloud infrastructure on AWS (preferred) and hybrid environments; lead cloud cost governance programs including Savings Plans, resource tagging taxonomy, and rightsizing automation.
- Design and enforce data security and access control frameworks using AWS HSM, IAM, Secrets Manager, and zero‑trust network principles; engage enterprise security architects to address vulnerabilities surfaced by internal and external audits.
- Define backup, replication, and disaster recovery architectures for critical data and application components; own RTO/RPO targets and validate them through regular DR rehearsals.
- Provide architectural direction for API management, data streaming pipelines, and integration layers (Service Now, Version One, Sumo Logic), ensuring operational…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×