Sr SRE & Automation Engineer; Customer Facing
Listed on 2026-10-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centres across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit
Job DescriptionNeo Cloud is building an AI-operated GPU cloud --- and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery.
You are the SRE who stands between raw infrastructure and the customer's experience --- designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.
- End-to-end reliability of the customer-facing GPU cloud service --- availability, job completion, provisioning latency, and tenant experience.
- Production Kubernetes clusters optimized for GPU workloads at scale (100--10,000 GPUs) as the runtime substrate for customer workloads.
- Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity --- placing customer jobs on the right hardware.
- Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (name spaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
- Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
- SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
- Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
- Monitoring & observability stack:
Prometheus, Grafana, Alert manager, Pager Duty --- tenant-aware dashboards and alerting. - GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling --- minimizing customer-visible impact.
- Infrastructure-as-code:
Terraform providers/modules, Helm, and Git Ops (ArgoCD/Flux) across GPU clusters. - Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.
- You are accountable for the customer's reliability experience --- when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
- Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
- Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
- Build self-service observability that lets customers answer their own questions --- status, quota, job health --- reducing support load.
- The remediation-actuator and workflow engine land here --- you make the control plane safe for…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).