×
Register Here to Apply for Jobs or Post Jobs. X

Devops Engineer

Job in Irving, Dallas County, Texas, 75084, USA
Listing for: Prodapt Solutions Private Limited
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 110000 - 160000 USD Yearly USD 110000.00 160000.00 YEAR
Job Description & How to Apply Below

Overview

Prodapt is the largest and fastest-growing specialized player in the Connectedness industry, recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider across North America, Europe and Latin America. With its singular focus on the domain, Prodapt has built deep expertise in the most transformative technologies that connect our world. Prodapt is a trusted partner for enterprises across all layers of the Connectedness vertical.

Prodapt designs, configures, and operates solutions across their digital landscape, network infrastructure, and business operations – and craft experiences that delight their customers. Today, Prodapt’s clients connect 1.1 billion people and 5.4 billion devices, and are among the largest telecom, media, and internet firms in the world. Prodapt works with Google, Amazon, Verizon, Vodafone, Liberty Global, Liberty Latin America, Claro, Lumen, Windstream, Rogers, Telus, KPN, Virgin Media, British Telecom, Deutsche Telekom, Adtran, Samsung, and many more.

A“Great Place To Work®Certified™” company, Prodapt employs over 6,000 technology and domain experts in 30+ countries across North America, Latin America, Europe, Africa, and Asia. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 30,000 people across 80+ locations globally.

Unlike a traditional production support role, this position requires strong engineering aptitude and a proactive operational mindset. The ideal candidate combines customer/agent facing application ecosystem knowledge, application monitoring expertise, incident management experience, automation skills, and a passion for leveraging AI and GenAI technologies to improve operational efficiency.

You will partner closely with engineering teams, release management, security, compliance, and business stakeholders to identify risks, reduce incidents, enhance monitoring capabilities, and drive platform reliability across a complex ecosystem supporting hundreds of enterprise applications.

Responsibilities Site Reliability & Operations
  • Serve as a key member of the Engineering Operations organization supporting self-assist Web & Mobile App and related business applications.
  • Provide Tier 1/2 operational support for production systems in a 24x7 environment.
  • Monitor application health, performance, availability, and customer experience across the platform.
  • Drive proactive issue detection and prevention rather than relying solely on customer-reported incidents.
  • Participate in incident response, triage, war rooms, major incident management, and post-incident reviews.
  • Perform root cause analysis (RCA) and identify opportunities to improve platform stability and resiliency.
  • Partner with Tier 1, Tier 2, Tier 3, infrastructure, security, and application teams to rapidly resolve issues.
  • Create and maintain operational runbooks, knowledge articles, and support documentation.
Observability & Monitoring
  • Build, maintain, and optimize monitoring dashboards, alerts, and health checks.
  • Analyze application logs, API activity, transactions, and performance metrics.
  • Utilize observability and monitoring platforms including:
    • Dynatrace
    • ELK Stack (Elasticsearch, Logstash, Kibana)
    • Catchpoint or other synthetic monitoring solutions
    • Quantum Metrics or other user session replay solutions
  • Reduce alert fatigue through automation, threshold tuning, and intelligent event correlation.
  • Develop and enhance monitoring strategies to provide end-to-end visibility across Digital ecosystem and integrated systems.
Incident Management & Problem Management
  • Act as a technical responder during production incidents and service disruptions.
  • Coordinate issue resolution efforts across multiple technical teams and stakeholders.
  • Manage incident lifecycle activities including:
    • Detection
    • Triage
    • Resolution
    • Communication
    • Root cause analysis
  • Identify recurring issues and lead problem management initiatives to eliminate operational inefficiencies.
Automation & AI Enablement
    • Develop innovative approaches to reduce manual operational effort through automation.
    • Leverage AI, GenAI, agentic workflows, and intelligent operational tooling where appropriate.
    • Create automation solutions to…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary