×
Register Here to Apply for Jobs or Post Jobs. X

Production Support Engineer

Job in Aliso Viejo, Orange County, California, 92656, USA
Listing for: Compunnel, Inc.
Full Time position
Listed on 2026-09-17
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Support, Systems Administrator
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below

The Lead - Production Support Engineer is a hands-on technical leadership role responsible for the operational health, reliability, observability, automation, and continuous improvement of APM and Env Ops services. The role requires active participation in production incidents, SWAT engagements, deployments, troubleshooting, migrations, monitoring onboarding, automation initiatives, and operational escalations while also providing technical leadership, coaching, governance, and stakeholder management. The engineer will support services including Datadog, SQL Server, Oracle, SSIS, SSRS, Informatica, CTU, APM services, and Precise and Team Quest tools.

Key Responsibilities
  • Participate in daily production and operational support activities.
  • Assist engineers with troubleshooting, issue resolution, and technical escalations.
  • Support P0, P1, and P2 production incidents and SWAT bridges.
  • Lead technical investigations and Root Cause Analysis (RCA) activities.
  • Support deployments, change implementations, migrations, and operational escalations.
  • Perform monitoring onboarding, configuration, and ongoing optimization.
  • Define monitoring standards and governance for Datadog and observability services.
  • Create and maintain Datadog monitors and dashboards.
  • Drive alert optimization, noise reduction, monitoring coverage, and observability maturity.
  • Partner with application, database, engineering, and vendor teams on observability and platform improvements.
  • Assess business and operational impact during incidents and coordinate technical response activities.
  • Provide leadership and stakeholder updates during critical incidents and ensure appropriate follow-up actions.
  • Lead RCA efforts, reduce recurring incidents and alert fatigue, and improve service reliability.
  • Identify technical debt and reliability improvement opportunities and drive corrective action plans through closure.
  • Identify and prioritize operational automation opportunities.
  • Partner with automation engineers on operational workflows and support AI-enabled operational solutions.
  • Promote automation-first thinking to reduce manual effort and improve scalability.
  • Coach and mentor engineers and facilitate cross-training and knowledge sharing.
  • Reduce single points of failure and develop backup coverage models across the team.
  • Drive technical capability development across the engineering team.
  • Own KPI/SLA reviews, operational reporting, governance dashboards, and service metrics.
  • Conduct documentation and runbook reviews and ensure operational knowledge is maintained.
  • Track risks, issues, action items, and service improvements.
  • Support executive and stakeholder reporting.
  • Support strategic migration activities and observability improvement initiatives.
  • Contribute to automation transformation roadmaps and initiatives focused on operational efficiency and scalability.
Required Qualifications
  • 4+ years of experience in Production Support, Application Support, IT Operations, Database Operations, Observability, or Infrastructure Support.
  • 1+ years of experience leading technical teams, incident response, operational governance, stakeholder management, and service ownership.
  • Strong experience with SQL Server administration and production support.
  • Experience with Oracle support and operations.
  • Experience with Datadog administration and observability.
  • Experience with SSIS, SSRS, and Informatica.
  • Strong knowledge of incident and problem management processes.
  • Experience with performance monitoring and tuning.
  • Experience supporting production environments and participating in SWAT engagements.
  • Strong Root Cause Analysis and troubleshooting skills.
  • Experience with monitoring and alert management.
  • Strong documentation and runbook management skills.
  • Strong leadership, coaching, stakeholder…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary