Platform Engineer
Listed on 2026-07-27
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure, Azure
Job Summary:
We are seeking a skilled and driven Platform Engineer to join our Hosting Operations team. In this role, you will be the primary owner of the release pipeline infrastructure, responsible for designing, building, and maintaining CI/CD pipelines that enable reliable and consistent deployments across all hosted customer environments. You will manage pipeline configuration and YAML files in Git Hub, administer deployments through Azure Dev Ops (ADO), leverage Ansible for configuration management, and build proactive monitoring solutions on Azure Cloud and on-Premises Infrastructure.
This is a hands‑on engineering role with broad impact across the full software delivery lifecycle for our hosted platform.
Job Description:
KEY RESPONSIBILITIES
Release Pipeline Management
- Own and maintain all CI/CD release pipeline YAML files across Git Hub and Azure Dev Ops (ADO), ensuring pipelines are well‑structured, version‑controlled, and consistently documented.
- Design, build, and enhance automated deployment pipelines for non‑production and production hosted environments, supporting multiple customer Environments.
- Manage pipeline branching strategies, approval gates, environment‑specific variable groups, and deployment conditions in alignment with Change Management (CAB) processes.
- Continuously improve pipeline reliability, execution speed, and error handling to reduce deployment risk and increase release frequency.
- Maintain pipeline‑as‑code standards, ensuring all pipeline definitions are peer‑reviewed and stored in source control alongside application code.
Deployment & Environment Management
- Coordinate and execute deployments across all hosted environments development, test, staging, and production - with a focus on consistency, repeatability, and minimal customer impact.
- Manage environment‑specific configuration files and secrets, ensuring proper separation of configuration from code and secure handling of sensitive values.
- Implement blue/green, rolling, or canary deployment strategies where applicable to reduce downtime and deployment risk.
- Partner with DBA, IT, and application teams to align deployment windows with maintenance schedules and change management requirements.
- Maintain deployment runbooks, rollback procedures, and post‑deployment validation checklists for all hosted environments.
Git Hub & Azure Dev Ops (ADO) Administration
- Manage Azure Dev Ops organizations, projects, pipelines, service connections, agent pools, and variable groups.
- Define and enforce Git workflow standards (branching strategies, pull request policies, merge requirements) across the hosting operations team.
- Integrate Git Hub and ADO with downstream tools and services, including Azure cloud services, monitoring platforms, and notification channels.
- Maintain pipeline and repository security, ensuring secrets and credentials are managed via Azure Key Vault or equivalent secret management tooling.
Configuration Management with Ansible
- Develop and maintain Ansible playbooks and roles for automated configuration management across hosted Linux and Windows environments.
- Use Ansible to enforce baseline configurations, manage software installations, apply patches, and maintain environment consistency across all hosted nodes.
- Integrate Ansible automation into CI/CD pipelines to enable fully automated end‑to‑end deployment and configuration workflows.
- Maintain an organized Ansible inventory structure and variable hierarchy supporting multiple environments and customer configurations.
- Document all Ansible playbooks and roles to support team knowledge sharing, onboarding, and audit readiness.
Proactive Monitoring & Observability
- Design and implement proactive monitoring solutions for hosted environments using internal tools, 3rd party developed tools, and alerting frameworks.
- Build and maintain dashboards that provide real‑time visibility into environment health, pipeline execution status, deployment success rates, and infrastructure performance.
- Define alerting thresholds and escalation rules to detect and surface issues before they impact customers, integrating alerts with Microsoft Teams, Pager Duty, or equivalent notification channels.
- Implement…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).