×
Register Here to Apply for Jobs or Post Jobs. X

Principal Core Engineer — Infra ​/ SRE

Job in Denver, Denver County, Colorado, 80202, USA
Listing for: Edgescale Ai
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below

Core Engineer

We're looking for a Core Engineer at the Principal Infra / SRE level to own the reliability, scalability, upgradeability, and operational excellence of our edge platform at fleet scale.

In this role, you'll be the technical authority for designing and operating compound capabilities that span software, infrastructure, networking, security, data, and hardware—ensuring we can reliably deploy, upgrade, and manage fleets of thousands of devices with the highest technical rigor. You will set and enforce production standards, and you have the authority to stop changes that would put fleet safety or reliability ing high-severity incidents, you are the technical owner—leading root-cause analysis and driving fixes across teams.

This is a hands-on role for someone who thrives in a high-ownership setting and wants to build the infrastructure that makes real-world AI possible. You'll operate in an AI-native way, using AI to assist diagnostics and operations while ensuring all production changes remain governed, reviewed, and auditable.

Responsibilities
  • Own platform-wide reliability and scalability architecture across the fleet, including upgradeability, rollback safety, resilience, observability, and incident response.
  • Lead the design and delivery of compound capabilities that span multiple specialist domains (hardware, networking, security, data, infrastructure, and AI runtime).
  • Set and enforce production-grade standards for operational excellence, including SLOs/SLIs, error budgets, on-call readiness, change management, incident management, and postmortem practices, with the authority to stop changes that introduce unacceptable risk.
  • Serve as the technical owner during high-severity incidents, leading diagnosis, root-cause analysis, and coordinated remediation across teams.
  • Design and operate secure, automated fleet lifecycle systems for deployment, updates, configuration management, and health management at scale.
  • Drive the evolution of observability and telemetry systems (metrics, logs, traces, audit, fleet state) so issues are detectable, diagnosable, and preventable.
  • Partner with engineering and commercial teams to translate real-world constraints into platform-level requirements and prioritization decisions.
  • Operate in an AI-native way: develop and use AI systems to accelerate diagnostics, automate operational workflows, and increase engineering velocity, while ensuring all production changes remain governed, reviewed, and auditable.
  • Mentor senior engineers across domains, review technical designs, and raise the quality bar for architecture and reliability across the organization.
Success Metrics

In your first 3 months, you will have:

  • Taken full ownership of a platform-wide reliability, upgradeability, or incident reduction initiative and delivered measurable improvements in fleet stability, deployment safety, and operational clarity.
  • Established or strengthened production standards that reduce risk and improve consistency across releases and fleet operations.
  • Demonstrated strong incident ownership by leading at least one high-severity investigation through root cause and durable remediation.

In your first year, you will be:

  • Owning the fleet-scale operational architecture end-to-end, with clear accountability for reliability, upgradeability, scalability, and security posture across thousands of deployed systems.
  • Delivering step-function improvements in platform resilience and operational excellence through durable systems (automated lifecycle management, observability, incident reduction, reliability standards).
  • Raising engineering rigor across the organization by enforcing standards, mentoring technical leaders, and driving cross-domain architectural decisions that compound over time.
Qualifications
  • 10+ years building and operating production infrastructure and distributed systems, including reliability engineering at scale across complex, multi-tenant or fleet environments.
  • Deep experience with SRE practices: SLOs/SLIs, error budgets, observability, incident response, postmortems, and operational automation (e.g., Kubernetes-based platforms, Linux systems, and automation through…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary