Platform Operations Engineer
Listed on 2026-09-12
-
IT/Tech
SRE/Site Reliability
Select how often (in days) to receive an alert:
As an NRG employee, we encourage you to take charge of your career and development journey. We invite you to explore exciting opportunities across our businesses. You’ll find that our dynamic work environment provides variety and challenge. Your growth is key to our ongoing success—take the lead in shaping your career development, goals and future!
Position SummaryThe Platform Operations Engineer is responsible for coordinating and improving the operational effectiveness of the Home Services technology platform. This role partners across Engineering, Product, Architecture, Quality Assurance, and Platform teams to improve platform reliability, operational visibility, and service performance through coordination, reporting, operational excellence, and continuous improvement.
The Platform Operations Engineer supports production operations, operational reporting, environment coordination,monitoring and observability, and AI-enabled operational capabilities to ensure engineering teams have the visibility, processes, and operational support necessary to deliver reliable technology solutions.
Key Responsibilities Production Support & OperationsCoordinate production support activities across Engineering, Product, and Platform teams tofacilitatetimelyincident resolution and effective communication.
Support incident management processes by coordinating investigations, root cause analysis activities, corrective actions, and post-incident follow-up with the appropriate engineering teams.
Coordinate operational readiness activities for releases, maintenance events, and platform changes.
Track recurring operational issues and coordinate continuous improvement initiatives with responsible teams.
Operational Metrics & PerformanceDevelop,maintain, and communicate operational dashboards, KPIs, and service performance metrics.
Analyze operational data toidentifytrends, risks, and opportunities for operational improvement.
Coordinate recurring operational reviews and provide visibility into platform health, service levels, and operational performance.
Support the definition, measurement, and reporting of operationalobjectivesand service quality indicators.
Monitoring & ObservabilityPartner with engineering teams to ensure appropriate instrumentation, logging, monitoring, and alerting are implemented across platform services.
Identify gaps in observability and coordinate improvements with engineering teams.
Supportadoptionof monitoring standards and operational reporting practices that improve proactive issue detection and platform visibility.
Promote consistent telemetry and operational reporting across platform services.
Coordinate environment planning, scheduling, availability, andutilizationacross multiple
Facilitate environment requests, refreshes, deployments, and conflict resolution activities.
Maintain visibility intoenvironmentreadiness, dependencies, and operational risks.
Communicate environment status, planned activities, and potential impacts to stakeholders.
AI & Operational InnovationIdentify opportunities toleverageAI and automation to improve production support, operational efficiency, and platform reliability.
Partner with engineering teams to implement AI-assisted operational workflows, monitoring, and support processes.
Support adoption of AI-enabled operational tools that improve issue detection, operational insights, knowledge management, and engineering productivity.
Evaluate emerging AI capabilities and recommend practical applications that enhance platform operations.
Operational ExcellenceSupport development and maintenance of operational documentation, runbooks, standard operating procedures, and knowledge resources.
Identify opportunities to improve…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).