Full-Stack Software Engineer; Infrastructure - remote
Austin, Travis County, Texas, 78716, USA
Listed on 2026-07-27
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software)
About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers.
As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Company Description
About Mirantis Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers.
As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Job Description
We are looking for an experienced Full-Stack Software Engineer to design and implement the Infrastructure Services that power our GPU-as-a-Service platform. You will build the control plane that turns high-level API calls into real infrastructure actions — enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware.
You will own the full lifecycle of the infrastructure-level services—from the Server and Machine Type APIs down to the provisioning workflows and reconciliation loops that keep the platform's view of hardware consistent with physical reality.
Key Responsibilities- Infrastructure API Design:
Design, build, and maintain the versioned REST and gRPC APIs for bare-metal server lifecycle, Machine Type definitions, and cluster CRUD operations. - Provisioning & Workflow Development:
Develop the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers. - Full-Stack Implementation:
Implement and maintain the consoles and interfaces that visualize hardware inventory, provisioning progress, and cluster health. - System Reliability:
Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably.
- API Development:
Strong experience designing RESTful APIs or gRPC services. You understand API versioning and gateway patterns. - Language Stack:
Proficiency in Go (preferred for backend/Kubernetes ecosystem) and modern Type Script/React (for the Console). - Kubernetes Knowledge:
Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions. - Bare-Metal Provisioning:
Hands-on experience with bare-metal provisioning flows — BMC/Redfish, PXE/iPXE, image management, and hardware inspection. - Asynchronous Systems:
Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures.
- State Reconciliation:
Experience building informers or reconciliation bridges that keep an external datastore consistent with Kubernetes resource state. - Multi-Tenancy:
Experience building…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).