Vice President, Reliability Engineering & Technology Operations
Listed on 2026-08-24
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Vice President, Reliability Engineering & Technology Operations
Antares Capital is a leading alternative credit manager and a trusted financing partner to private equity sponsors and middle-market companies. We are committed to building resilient, scalable, and modern technology platforms that support our business and clients.
As part of our continued technology transformation, we are seeking a Vice President, Reliability Engineering & Technology Operations to lead the evolution of our production operations, reliability engineering, observability, and operational automation capabilities.
This is a strategic leadership role for an engineering-minded leader who thrives at the intersection of software engineering, cloud infrastructure, platform operations, and operational excellence. The ideal candidate combines strong technical depth with exceptional execution skills and has experience building highly reliable systems while leading teams through modernization and transformation initiatives.
Technology is central to Antares' growth strategy. We are investing heavily in cloud platforms, engineering excellence, AI-enabled workflows, automation, and modern operational practices.
As the leader of Reliability Engineering & Technology Operations, you will be responsible for the availability, performance, scalability, and resilience of critical business platforms. You will partner closely with Engineering, Infrastructure, Cybersecurity, Data, and Business stakeholders to ensure our systems remain secure, observable, scalable, and operationally mature.
You will help shape the future of technology operations by introducing reliability engineering practices, expanding observability, leveraging AI-driven operational capabilities, and reducing operational overhead through automation.
This role is ideal for someone who has grown through engineering, cloud, platform, infrastructure, or Dev Ops leadership roles and understands how to bridge engineering and operations to deliver exceptional business outcomes.
Key Responsibilities Reliability Engineering Leadership- Establish and lead Antares' Reliability Engineering function.
- Define and implement strategies that improve system reliability, resiliency, scalability, and operational excellence.
- Partner with engineering teams to embed reliability practices throughout the software development lifecycle.
- Drive adoption of modern operational practices including service ownership, operational readiness reviews, SLOs, error budgets, and post-incident learning.
- Lead production operations across critical business applications and technology platforms.
- Establish clear support models, escalation paths, ownership boundaries, and service management processes.
- Oversee operational readiness, release support, change management, and platform health.
- Continuously improve operational maturity through metrics, automation, process simplification, and engineering collaboration.
- Provide technical leadership for cloud-based platforms running in Microsoft Azure.
- Partner with Infrastructure and Engineering teams to optimize reliability, scalability, and operational efficiency.
- Support containerized workloads and Kubernetes-based environments.
- Drive best practices around cloud architecture, capacity planning, platform resilience, security, governance, and cost optimization.
- Ensure cloud platforms are designed and operated to meet business continuity and availability objectives.
- Serve as the executive incident leader during major production events.
- Coordinate cross-functional teams during outages and high-severity incidents.
- Manage communications with business stakeholders and technology leadership.
- Establish disciplined root cause analysis processes and ensure corrective actions are executed.
- Drive long-term reduction in recurring incidents and operational risk.
- Champion an automation-first and AI-enabled approach to technology operations.
- Identify opportunities to leverage AI for incident triage, alert correlation, knowledge management, operational analytics, and runbook execution.
- Partner with engineering teams to develop intelligent automation and self-healing capabilities.
- Evaluate emerging AIOps, agentic AI, and automation technologies and drive adoption where appropriate.
- Reduce manual operational effort through scripting, workflow automation, orchestration platforms, and AI-assisted tooling.
- Lead and mentor high-performing technology operations and reliability engineering teams.
- Build strong partnerships across Engineering, Infrastructure, Cybersecurity, Architecture, Data, and Business teams.
- Translate operational challenges into actionable roadmaps and measurable initiatives.
- Drive accountability, execution excellence, and continuous improvement across the organization.
- Present operational trends, risks, recommendations, and performance metrics to technology…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).