×
Register Here to Apply for Jobs or Post Jobs. X

Senior Network Reliability Engineer, Incident Management

Job in Mountain View, Santa Clara County, California, 94043, USA
Listing for: Skylo
Full Time position
Listed on 2026-07-26
Job specializations:
  • IT/Tech
    IT Support, Cybersecurity, Network Security, Network Engineer
Job Description & How to Apply Below

Network Reliability Engineer (NRE) For Incident Management

As a Network Reliability Engineer (NRE) for Incident Management in the Global Product Support & Customer Success organization, you are the operational nerve center of Skylo's 24x7 incident response function. You own the bridge, from first alert to full-service restoration. On a network where every subscriber is roaming over satellite, structured incident coordination directly determines whether millions of devices stay connected.

You operate across all three Skylo hubs (Mountain View, Espoo, Bengaluru) and across every network domain: RAN, 5G Core, Cloud Infrastructure, and OSS, ensuring incidents are triaged, escalated, and resolved with speed and discipline.

At Senior NRE level you independently command bridge calls for Sev 1-4 incidents, support subject matter experts as incident coordinator and function as the communication lead during Sev 1 events, drive the execution of runbooks without supervision, manage the full incident lifecycle end-to-end in the ticketing system, and continuously improve the procedures you operate against. You are not a ticket router, you are the structured coordinator who ensures every incident has a clear owner, a live timeline, and a documented resolution path.

Assuring Skylo's network availability commitments to MNO partners and meeting SLA obligations is your primary measure of success.

Incident Command & Bridge Coordination
  • Serve as the central command point during all network degradations, service disruptions, and subscriber-impacting events, opening the bridge, initiating the incident management process, and maintaining command from first alert through full-service restoration.
  • Own end-to-end incident lifecycle management for Sev 1-4: initiate bridge calls, identify the cause and impacted domain, page the correct on-call NRE, maintain bridge discipline with clear ownership and timelines, and drive to restoration.
  • Prioritize incidents according to urgency and business impact, classifying severity accurately using alarm signatures, subscriber impact data, and domain KPI telemetry from available OSS systems.
  • Escalate to subject matter experts in Operations and Engineering teams when critical or time-sensitive resolution is required, providing full technical context, a structured problem statement, and a documented timeline.
  • Engage and interface with vendor support teams (RAN vendor, Core vendor, cloud infrastructure) when incident resolution requires external escalation; track vendor SLA response and escalate vendor delays to the domain NRE.
  • Support hypercare operations during major network launches, high-risk change windows, and special events, maintaining readiness and acting as first responder for any degradation during the hypercare window.
Incident Documentation & Lifecycle Management
  • Ensure all trouble tickets are created promptly in the incident management system (Jira/Service Now) with complete technical details, troubleshooting steps, MOPs followed, and outcome documentation, never closing an incident with an incomplete ticket.
  • Produce a structured incident timeline artifact within two hours of closure: sequence of events, alarms triggered, actions taken, bridge participants, and all open action items with assigned owners and due dates.
  • Manage the open incident backlog at optimum levels: track ageing tickets, escalate stalled items, and ensure no incident closes without a documented resolution path or a justified deferral.
  • Coordinate post-incident review (PIR) scheduling: compile the incident record, gather logs from in-house observability tools and other relevant sources, and deliver a structured problem statement to the domain NREs owning the root cause analysis.
  • Handle internal, external, and MNO partner incident escalations and follow-ups; interface with Market Operations, OEM contacts, and partner NOCs for joint incident resolution, ensuring external-facing communications are approved before transmission.
Network Availability & SLA Assurance
  • Assure that Skylo's operated network meets agreed availability KPIs and MNO partner SLA commitments, proactively tracking availability metrics and flagging…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary