Production Support Lead
Fort Worth, Tarrant County, Texas, 76101, USA
Listed on 2026-09-20
-
IT/Tech
IT Support
About the role
We are seeking a Production Support Lead to own the end-to-end operational health, stability, and continuous improvement of a critical enterprise application. In this role, you will lead incident management efforts, drive incident reduction strategies, partner closely with engineering teams, and ensure seamless communication across stakeholders. You will play a key role in transforming reactive support operations into a proactive, data-driven support model by identifying recurring issues, driving root cause resolution, and improving overall application reliability.
Inthis role, you will:
- Own the overall operational health, availability, and supportability of the application.
- Lead and execute incident management processes, ensuring timely resolution of P1/P2 incidents and minimizing business impact.
- Drive incident reduction initiatives through root cause analysis (RCA), problem management practices, and collaboration with engineering teams.
- Facilitate and lead war-room calls during critical production incidents.
- Manage cross-functional escalations and coordinate activities between support teams, development teams, infrastructure teams, and business stakeholders.
- Partner with engineering teams to prioritize and implement permanent fixes through application improvement backlogs.
- Analyze incident trends, recurring failure patterns, and operational metrics to identify opportunities for process and system improvements.
- Develop and maintain operational dashboards, KPI reports, trend analyses, and executive health scorecards.
- Ensure adherence to ITIL-based Incident, Problem, and Change Management processes.
- Drive continuous service improvement initiatives that improve application stability, reduce incident volume, and shorten resolution times.
- Provide regular operational updates and communicate incident status effectively to leadership and business stakeholders.
- Foster a culture of accountability, operational excellence, and proactive support management.
We strive to provide flexibility wherever possible. Based on this role’s business requirements, this is a remote position open to qualified applicants in Dallas, TX or Dublin, OH. Regardless of your working arrangement, we are here to support a healthy work-life balance through our various wellbeing programs
The working arrangements for this role are accurate as of the date of posting. These may change based on the project you are engaged in, as well as evolving business and client requirements. We will always be transparent about role expectations.
Please note:
Some roles may require in-person interviews at Cognizant offices or client locations based on project needs.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.
- Proven experience leading production support teams in high-volume enterprise application environments.
- ITIL certification with hands-on experience in Incident Management, Problem Management, and Change Management processes.
- Strong working knowledge of:
- Microsoft SQL Server (MSSQL)
- VB6
- Ability to review application code, perform diagnostic analysis, execute database queries, and collaborate effectively with engineering teams during incident resolution.
- Demonstrated success in reducing incident volumes through structured problem management and root cause elimination.
- Experience managing client-facing escalations and leading cross-functional incident resolution efforts.
- Strong analytical capabilities with experience creating operational reports, KPI dashboards, and trend analyses using tools such as Excel and/or Power BI.
- Excellent verbal and written communication skills, with the ability to engage effectively with technical and non-technical stakeholders.
- Ability to thrive in a fast-paced environment and manage multiple priorities under pressure.
- Experience supporting large-scale enterprise applications with high transaction volumes and complex integrations.
- Strong expertise in operational excellence and continuous improvement methodologies.
- Proven ability to convert recurring production issues into long-term engineering solutions.
- Experience developing executive-level operational scorecards and service health reporting.
- Strong stakeholder management skills with experience engaging senior leadership and business executives.
- Knowledge of monitoring, observability, and application performance management tools.
- Experience…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).