Job Description & How to Apply Below
Cloud Architecture & Infrastructure
Architect, build, and govern scalable, secure, and cost-efficient AWS infrastructure — EC2, EKS, Lambda, API Gateway, VPCs, IAM, Security Groups, Load Balancers, Route 53, and related services.
Lead Infrastructure-as-Code (IaC) strategy using Terraform and / or AWS CDK; establish reusable module libraries and enforce IaC governance standards across all projects.
Drive multi-account AWS environment management using Control Tower, AWS Organizations, and Landing Zone patterns.
Apply AWS Well-Architected Framework principles across all five pillars — reliability, security, operational excellence, performance efficiency, and cost optimisation — on all new and existing workloads.
Evaluate and recommend new AWS services and toolchain improvements to keep EXL's platform capabilities current.
Team Leadership & People Management
Lead, mentor, and coach a team of Dev Ops Engineers.
Own sprint planning, workload distribution, and capacity management for the Dev Ops team.
Foster a culture of ownership, continuous learning, and engineering excellence within the team.
Act as the primary escalation point for critical production incidents and high-priority engineering decisions.
CI/CD, Automation & Dev Ops Practices
Design, implement, and own enterprise-grade CI/CD pipelines using Git Hub Actions, AWS Code Build, AWS Code Deploy, and ArgoCD.
Define and enforce branching strategies, deployment standards, release gates, and rollback procedures across all engineering teams.
Champion Git Ops practices and drive adoption of automated testing, security scanning (Dev Sec Ops ), and compliance checks within pipelines.
Create and maintain reusable templates, runbooks, and golden-path tool chains for developers to self-serve infrastructure and deployments.
Data Platform & Engineering Collaboration
Collaborate closely with data engineering teams to design and support platforms spanning AWS Glue, Databricks, DynamoDB, Redshift, and S3-based data lakes.
Ensure Dev Ops processes, pipeline reliability, and observability are appropriately applied to data and ML/analytics workloads.
Provide infrastructure guidance and developer-experience improvements to application and data engineers to enable faster, safer deployments.
Monitoring, Reliability & Security
Define and own the observability strategy — dashboards, alerting, log management, and SLO/SLA tracking using Cloud Watch, Grafana, Prometheus, and the ELK Stack.
Drive incident management processes — from detection and root-cause analysis through to post-mortems and corrective action follow-through.
Enforce cloud security best practices: IAM least-privilege, secrets management, vulnerability scanning, network segmentation, and compliance controls.
Lead Fin Ops initiatives — continuous cost visibility, anomaly alerting, rightsizing, and Reserved Instance / Savings Plan strategy.
Stakeholder Management & Governance
Serve as the primary Dev Ops point of contact for client-facing and internal technical stakeholders.
Produce and present infrastructure architecture decisions, platform roadmaps, and cloud health reports to senior leadership.
Define and track Dev Ops KPIs — deployment frequency, change failure rate, MTTR, lead time — and use data to guide team priorities.
Ensure compliance with organisational, client, and regulatory security and audit requirements.
Skills:
Must have:
Deep, production-grade AWS expertise across compute (EC2, EKS, Lambda), networking (VPC, Route 53, Load Balancers, API Gateway), security (IAM, KMS, Security Groups), and storage (S3, EBS, EFS).
Expert-level Infrastructure-as-Code proficiency — Terraform (modules, state management, remote backends) and / or AWS CDK.
Proven track record designing and operating CI/CD pipelines at scale — Git Hub Actions, Code Build, Code Deploy, ArgoCD; experience with Git Ops delivery models.
Hands-on container orchestration expertise — Kubernetes / Amazon EKS — including Helm, autoscaling, RBAC, and network policies.
Solid experience supporting data engineering platforms: AWS Glue, Databricks, Redshift, DynamoDB, or equivalent.
Strong scripting and automation skills in Python, Bash, or Power Shell…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×