Lead Engineer, Internal Developer Platform
Listed on 2026-08-22
-
Software Development
DevOps, Cloud Engineer - Software
This role is responsible for designing, building, and operating cloud platform and Dev Ops capabilities across AWS and GCP. The engineer will enable reliable, secure, and scalable delivery by developing Kubernetes and Infrastructure-as-Code foundations, building and maintaining CI/CD pipelines, and driving automation, observability, and operational best practices across engineering teams.
In this flex office/home role, you will be expected to work a minimum of 10 days per month from one of the following office locations:
Madison, WI 53783;
Boston, MA 02110, Denver, Co. Candidates must reside within a 50-mile radius of the office location (or 35-mile radius for Boston).
Position Compensation Range: $ - $
Pay Rate Type:
Salary
Compensation may vary based on the job level and your geographic work location. Relocation support is offered for eligible candidates.
Core Qualifications & Responsibilities:
Developer Experience (DevX) & Backstage Development:
- Proven experience in designing, implementing, and managing an internal developer portal, preferably with Backstage.io.
- A strong desire to improve the developer experience by creating self-service tools and streamlining workflows.
- Expertise in developing custom Backstage plugins and customizing the platform using Type Script, Node.js, and React to meet the unique needs of our developers.
- Experience configuring and managing the Backstage Software Catalog to provide a unified view of services, documentation, and tooling.
- Ability to create and maintain software templates using Backstage's Scaffolder to standardize project creation and enforce best practices.
Cloud and Infrastructure Management:
- Extensive experience with at least one major cloud provider (AWS, GCP) and familiarity with the other. This includes expertise in core services like VPC, IAM, S3, and managed Kubernetes (EKS, GKE).
- Proficiency in Infrastructure as Code (IaC) principles and hands‑on experience with tools like Terraform to automate the provisioning and management of cloud resources.
- Experience with Crossplane.ioto manage cloud services and infrastructure through a unified, Kubernetes-native API. This includes creating and managing composite resources to abstract away the complexity of the underlying infrastructure.
- Strong understanding of containerization technologies, including Docker and Kubernetes, and their ecosystems.
CI/CD & Dev Ops Mindset:
- Solid experience with CI/CD principles and proficiency in building and maintaining pipelines using Git Lab CI/CD.
- Experience integrating various tools into CI/CD pipelines to automate builds, testing, and deployments.
- Proficiency in managing software artifacts and dependencies using JFrog Artifactory. This includes creating and managing repositories and integrating Artifactory with our CI/CD pipelines.
- A strong scripting and automation skill set using languages like Python, Go, or Bash.
Observability & Site Reliability Engineering (SRE):
- A solid understanding of SRE principles, including Service Level Objectives (SLOs), error budgets, and a focus on reliability, scalability, and performance.
- Hands‑on experience with Datadog for full‑stack observability, including setting up monitors, creating dashboards, and using APM for distributed tracing.
- A proactive approach to identifying and resolving platform issues before they impact developers.
- Experience in troubleshooting complex, distributed systems and a methodical approach to problem‑solving.
Software Development Fundamentals:
- Strong programming skills in languages such as Python, Type Script, and JavaScript.
- Experience with web development fundamentals and API design (REST, GraphQL).
- A solid understanding of software development lifecycle (SDLC) best practices, including version control with Git.
- Demonstrated knowledge of modern SDLC practices and operating models (e.g., Agile, product‑based teams, platform‑as‑a‑product).
- Demonstrated experience delivering and operating cloud platforms in both AWS and GCP, including core services for compute, storage,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).