Platform Delivery & Reliability Engineer; Remote
Colorado Springs, El Paso County, Colorado, 80509, USA
Listed on 2026-07-18
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Project Manager
Platform Delivery & Reliability Engineer (Remote) - 29337
Enlighten, honored as a Top Workplace from USA Today, is a leader in big data solution development and deployment, with expertise in cloud-based services, software and systems engineering, cyber capabilities, and data science. Enlighten provides continued innovation and proactivity in meeting our customers’ greatest challenges.
We recognize that the most effective environment for your projects doesn’t always look the same. Our hybrid work approach ensures that you can make lasting relationships with your team and collaborate in-person to get the job done—while having the flexibility to work from home when needed to achieve focused results.
Why Enlighten? At Enlighten, our team’s unwavering work ethic, top talent and celebration of innovative ideas have helped us thrive. We know that our employees are essential to our company’s success, so we seek to take care of you as much as you take care of us. Here are a few highlights of our benefits package:
- 100% paid employee premium for healthcare, vision and dental plans.
- 10% 401k benefit.
- Generous PTO + 10 paid holidays.
- Education/training allowances.
Anticipated Salary Range: $-$. The salary range for this role is intended as a good faith estimate based on the role's location, expectations, and responsibilities. When extending an offer, Enlighten takes a variety of factors into consideration which include, but are not limited to, the role's function, internal equity and a candidate's education or training, work experience, certifications and key skills.
Occasionally positions/roles may include additional non-recurrent compensation and will be addressed by the recruiter during the interview process.
Enlighten is looking for a Senior Platform Delivery & Reliability Engineer to own the rollout of our data lakehouse platform across a large, multi-site government enterprise, currently ~50 production Kubernetes clusters and growing. This role is a rare hybrid of platform engineer, SRE, and delivery lead. You can deploy the platform, debug anything you encounter in the field, feed what you learn back to the engineering teams, and help fix underlying issues in the code base.
Just as importantly, you can step back from any individual issue and fix the system that produced it by building the processes, tooling, and communication channels inside Enlighten that make every deployment faster and less painful than the one before it.
You will be a full member of the Infrastructure team, working daily with our Ingest, Query, Application and Testing teams as well as government customers and site personnel. Success in this role looks like: rollouts across the enterprise happen predictably and efficiently, issues found in the field are cataloged, communicated, and resolved quickly, and the friction that slows deployments steadily disappears.
- Plan, coordinate, and execute deployments and upgrades of the data lakehouse platform across 50 production Kubernetes clusters in customer environments.
- Debug and troubleshoot critical issues anywhere in the stack (infrastructure, Kubernetes, platform services, data services, and applications) and drive them to root cause.
- Contribute patches, configuration changes, and automation improvements directly back to the platform.
- Catalog and triage issues discovered in the field, communicate them clearly to the Infrastructure, Ingest, Query, and Application teams, and maintain a living knowledge base of failure modes, fixes, and runbooks.
- Enable and improve site reliability for fielded environments: monitoring, alerting, incident response, and continuous reliability improvement.
- Identify friction and dysfunction in how deployments happen, such as unclear handoffs, communication gaps, and repeated manual work; then design, implement, and institutionalize the processes that eliminate them (release checklists, readiness reviews, escalation paths, cross-team communication cadences).
- Continuously improve deployment tooling and automation so rollouts become faster, safer, and more repeatable.
- Coordinate with a large set of stakeholders (the engineering teams, government programs,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).