Director, Production Engineering
Listed on 2026-09-12
-
IT/Tech
SRE/Site Reliability
In an increasingly complex world where people are starving for someone they can trust, we stand for something simple: always put the client first. We do well by doing good for those we serve. It's the ultimate measure. We believe in providing value beyond a doubt and in the notion that time will either expose you or promote you, based on your willingness to embrace change.
We serve financial advisors and investors through three entities, each headquartered in Omaha, Nebraska:
Carson Wealth, Carson Coaching and Carson Partners. We provide coaching and partnership services to advisor firms – and straightforward financial advice to the investing public. We all share a common mission to be the most trusted in financial advice.
Lead production engineering strategies, standards, and practices that ensure reliable, stable, observable, and resilient production technology services. Establish enterprise-wide governance for monitoring, telemetry, incident response, service reliability, operational risk management, and production readiness. Drive proactive operational engineering practices that support measurable service performance, business continuity, and customer expectations. Partner with internal technology leaders to improve service performance, operational excellence, and production outcomes across the organization.
WhatTo Expect
Production Reliability & Service Health: Lead the Production Engineering function by establishing strategies, standards, and practices that ensure highly available, reliable, and resilient production services. Own enterprise observability capabilities, including monitoring, logging, distributed tracing, alerting, telemetry, and operational analytics. Define and govern reliability engineering practices, including Service Level Objectives, Service Level Indicators, service health measurement, and continuous improvement initiatives. Establish governance frameworks for service ownership, configuration management, feature management, and operational controls to ensure accountability, consistency, and risk mitigation.
Own production readiness standards and processes to ensure applications and services are operationally prepared before deployment into production environments.
Incident Management: Serve as incident commander and own the enterprise incident management framework, including operational response standards, escalation processes, governance, and continuous improvement practices. Establish root cause analysis standards and ensure consistent post-incident review processes that drive organizational learning and long-term issue resolution. Lead efforts to improve incident response effectiveness, recovery times, and overall operational maturity. Partner with internal technology leaders to strengthen operational readiness and improve production outcomes across the organization.
AIOps & AI Platform Operations: Lead the AIOps strategy by leveraging automation, analytics, and intelligent operational capabilities to improve service management, issue detection, and operational efficiency. Drive the adoption of AI-enabled operational practices that enhance observability, incident response, and reliability engineering outcomes. Lead operational governance, observability, telemetry, and reliability practices for AI-enabled platforms and services, ensuring appropriate performance, operational transparency, and risk controls.
Partner with internal technology leaders to evaluate and implement emerging AI operational capabilities that improve service reliability and operational effectiveness.
Executive Reporting & Operational Metrics: Develop and deliver executive-level reporting and operational metrics that provide visibility into service health, reliability, operational risks, incidents, and overall technology performance. Establish operational performance indicators, dashboards, and reporting frameworks that support data-driven decision-making and technology investment priorities. Provide leadership with actionable insights into operational trends, service performance, reliability outcomes, and organizational risk exposure.
Leadership &
Collaboration:
Lead and develop a high-performing Production Engineering function focused on service stability, performance, and continuous improvement. Partner with internal technology leaders to align reliability, observability, and operational excellence with business objectives. Establish operational metrics, governance, and accountability frameworks that drive…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).