Senior Operations Analyst (SRE
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability
Senior Operations Analyst (SRE)
CGI's Advantage Cloud Operations is an SRE-driven operating model, anchored on an Operations Control Plane that unifies telemetry, event management, automation, and IT Service Management (ITSM). The Senior Operations Analyst is a senior Site Reliability Engineering (SRE) practitioner responsible for driving reliability engineering, problem management, and runbook automation across customer environments. This role also mentors the operations team in adopting proactive, data-driven operational practices while improving platform reliability, operational efficiency, and service quality.
This position can be performed from any CGI U.S. CSG office, with a preference for Lafayette, LA.
Your Future Duties and ResponsibilitiesReliability Engineering & Operations Leadership
- Serve as a senior SRE within the Cloud Operations team supporting the CGI Advantage platform.
- Own end-to-end reliability outcomes, including reducing MTTD and MTTR, minimizing escalations, and improving release safety.
- Lead incident response, Root Cause Analysis (RCA), and Problem Management activities.
- Design and implement automation and self-healing workflows with appropriate guardrails, approvals, and rollback capabilities.
- Act as the technical lead for the Operations technology stack and drive platform adoption.
- Mentor Operations Analysts while establishing operational standards, taxonomy, and Service Level Objective (SLO) discipline.
Observability & Incident Response
- Drive distributed tracing adoption using Open Telemetry with trace-to-log and metric correlation for priority services.
- Define SLO-aligned alerting strategies and optimize deduplication, suppression, and event correlation to reduce alert noise.
- Lead Sev 1 and Sev 2 incident bridges while ensuring evidence-based RCAs and corrective actions are completed.
- Standardize dashboards and observability signal sets across customer environments.
Automation & Problem Management
- Build and maintain an active automation backlog and deliver automated runbooks on a quarterly basis.
- Convert recurring incidents into problem records with RCA hypotheses, ownership, and due dates.
- Develop guard-railed self-service operational actions (diagnostics, restarts, environment tasks) utilizing RBAC and audit trails.
- Champion Configuration-as-Code practices, including version control, drift detection, and baseline versus override governance.
Release Safety & Security
- Enforce release gates, quality thresholds, and automated rollback to the last known good state.
- Implement time-bound privileged access (JIT/Break Glass), SSO integration, and audit-ready operational controls.
- Apply Policy-as-Code guardrails across Infrastructure as Code (IaC), Kubernetes admission controls, and CI/CD pipelines.
Leadership & Governance
- Mentor and upskill Operations Analysts while documenting and promoting "golden path" operational workflows.
- Report reliability KPIs including:
Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), First Touch Resolution, Automation Coverage. - Partner with Product, Delivery, ITSM, and Security teams within a matrix operating model.
• 6–9 years in cloud operations/SRE, including 3+ years operating production SaaS on Azure.
• Observability:
Open Telemetry, distributed tracing, log analytics, metrics, and dashboards.
• Kubernetes at scale: AKS, Calico network policy, KEDA autoscaling, Helm based deployments.
• IaC and configuration:
Terraform;
Git Ops with Git Hub Actions and Argo CD.
• Event management and orchestration: event correlation, auto remediation workflows.
• Scripting/automation: strong Python and Bash; API driven integration across ITSM and tooling.
• Security tooling:
Keycloak/ SSO, Azure Key Vault, cert manager, JIT access patterns.
• ITSM depth: incident/problem/change management, severity taxonomy, SLA/SLO reporting.
Preferred Experience• Policy as code (OPA, Kyverno, Conftest) and compliance grade audit reporting.
• Job/batch orchestration (JS7 or equivalent) and API gateway operations (APISIX/APIM).
• PostgreSQL operations, PgBouncer connection pooling, and managed database patterns.
• Multi tenant managed services…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).