Operations Analyst (SRE
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability
Operations Analyst (SRE)
CGI's Advantage Cloud Operations is an SRE-driven operating model, anchored on an Operations Control Plane that unifies telemetry, event management, automation, and IT Service Management (ITSM). The Operations Analyst is a hands-on Site Reliability Engineering (SRE) practitioner responsible for monitoring, triaging, and resolving incidents using correlated telemetry and automated runbooks. This role contributes to reducing operational toil through automation and disciplined problem management while supporting the reliability of CGI's Advantage platform.
This position can be performed from any CGI U.S. CSG office, with a preference for Lafayette, LA.
Your Future Duties and ResponsibilitiesOperations & Reliability
- Serve in a hands-on operations role (SRE track) within the Cloud Operations team supporting the CGI Advantage platform.
- Take first-line ownership of monitoring, event triage, incident handling, and service request execution.
- Work from the operator portal utilizing correlated dashboards, one-click runbooks, and standardized operational workflows.
- Contribute to automation, ticket hygiene, and knowledge base quality to reduce repeat work.
- Progress toward the Senior Operations Analyst role through the SRE realignment career path.
Monitoring & Event Triage
- Monitor environment health using standardized dashboards and SLO-aligned alerts.
- Triage events in IAP by applying deduplication and suppression context while routing incidents according to severity taxonomy.
- Utilize distributed traces and correlated logs and metrics to isolate faults before escalation.
- Meet First Touch Resolution and L1-to-L2 escalation reduction targets.
Incident & Service Request Execution
- Execute approved runbooks and guard-railed self-service actions, including diagnostics, restarts, and environment tasks.
- Maintain complete, well-classified tickets using guided intake fields and consistent operational taxonomy.
- Coordinate activities within automatically created Microsoft Teams incident channels while documenting timelines and handoffs.
- Draft clear customer-facing status updates using standardized communication templates.
Continuous Improvement
- Identify recurring incidents as problem management candidates and support Root Cause Analysis (RCA) efforts with evidence and operational data.
- Develop small automation scripts using Python, Bash, or Power Shell and recommend new runbook candidates based on recurring operational tasks.
- Report configuration drift and validate changes against approved baselines under change control.
- Maintain runbooks and knowledge articles while leveraging LLM-assisted triage tools to improve operational efficiency.
• 2–5 years of experience in cloud or application operations, Network Operations Center (NOC), or L1/L2 production support environments.
• Working knowledge of Microsoft Azure and Kubernetes (AKS), including:
Pods, Deployments, Logging, Scaling concepts
• Experience with monitoring and logging platforms, including:
Grafana-style dashboards, Log search and analytics, Alert management
• Basic scripting experience using:
Bash, Python, Power Shell
• Knowledge of IT Service Management (ITSM), including:
Incident Management, Service Request Lifecycle, SLA awareness, Quality ticket documentation
• Linux fundamentals
• Networking fundamentals, including: DNS, TLS, Load balancing
• Experience using Git
• Strong written and verbal communication skills for incident updates and operational handoffs.
Preferred ExperienceOpen Telemetry, Jaeger tracing, or Prometheus/Loki exposure.
Ansible playbooks, Terraform basics, or Git Hub Actions pipelines.
Event management/AIOps platforms and runbook automation tools.
PostgreSQL basics and SQL for operational queries and diagnostics.
Certifications:
AZ 900/AZ 104, CKA (entry), ITIL 4 Foundation.
CGI expects to accept applications for this position through 9/30/2026.
Other Information:
CGI is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).