Who are we?
Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet.
A place where tech thinkers and future builders turn bold ideas into breakthrough experiences, we welcome your unique perspective.Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.
Job Summary
The Principal Network Engineer is a senior technical leader within the Site Reliability Engineering (SRE) organization, responsible for designing, scaling, and maintaining highly reliable, self-healing network platforms. This role integrates AI driven operations, predictive analytics, observability, and advanced automation to ensure the network meets strict availability, performance, and resiliency standards. The engineer translates business and technical requirements into scalable network capabilities, defines SLOs/SLIs for mission critical services, and delivers engineering artifacts for implementation teams.
This role mentors engineers, supports architectural decisions, and drives continuous improvement.
Responsibilities
Define, maintain, and own network SLIs, SLOs, and error budgets across critical services
Lead major incident response efforts with AI assisted diagnostics and blameless postmortems
Maintain network observability frameworks including telemetry pipelines, distributed tracing, and ML-based anomaly detection
Reduce operational delays through automation, AI copilots, LLM generated playbooks, and self-healing workflows
AI-Driven Network Engineering
Integrate AIOps platforms for predictive detection, anomaly correlation, and automated remediation
Implement generative AI workflows for script creation, documentation, and troubleshooting
Develop ML based analytics pipelines for traffic modeling, baselining, and forecasting
Deployment & Engineering Leadership
Translate requirements into scalable network capacity and capability designs.
Deliver engineering artifacts and guide implementation teams
Lead live deployments of multi-vendor network solutions and large-scale migrations
Maintain guidelines, templates, and automation for SRE and Ops teams
Build repeatable workflows that support automation-first operations
Train and mentor engineering teams through workshops and technical boot camps
Product Lifecycle Engagements
Provide feedback across Product, Development, and QA based on operational insights
Participate in design reviews, functional specs, and cross-functional architecture discussions
Support POCs, device testing, API validation, etc.
Provide deployment feedback, validate automation logic, and assess system readiness
Service Delivery & Fulfillment
Deliver fulfillment for network and datacenter services using automated workflows
Ensure high-quality provisioning and seamless end-to-end service operations
Collaborate with escalation teams, order fulfillment, and professional services
Analyze fulfillment KPIs and drive improvements in efficiency and reliability
Identify and resolve root causes impacting customer experience
Innovation & Continuous Improvement
Lead innovation initiatives across products, processes, and engineering teams
Deploy features and automation frameworks that improve reliability and customer satisfaction
Present at industry events and train internal teams in new developments
Provide cross-industry insights to keep engineering practices aligned with modern trends
Mentorship & Leadership of SRE Engineers
Mentor SRE engineers in networking, automation, observability, and reliability engineering best practices
Coach engineers in Service Level Objective (SLO) ownership, incident response, and resilience engineering
Review automation scripts and design proposals, elevating engineering quality
Lead training in AI-assisted operations, AIOps tooling, and automated remediation
Develop skill progression plans and identify growth opportunities
Mentor during complex incidents and post-incident analysis
Foster a culture of learning, excellence, and cross-team collaboration
Qua…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: