×
Register Here to Apply for Jobs or Post Jobs. X

Cloud Service Security Platform DevOps & Maintenance

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: Bitdeer Technologies Group
Full Time position
Listed on 2026-08-17
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit

Job Description

You are the paved road for every AI product we ship — the CI/CD, IDP, and MLOps substrate that lets model teams deploy without opening a ticket.
This is an entry-to-mid level role

Position Summary

Bitdeer is building an AI-operated GPU cloud, and the AI Cloud team ships the products on top of it. As Cloud Senior Dev Ops Engineer you build the paved road those products travel down: CI/CD, infrastructure-as-code, MLOps pipelines, and the Internal Developer Platform that lets model, data-science, and product teams deploy at speed without stepping around governance. You are the bridge between research and production, and the operator of the platform substrate the AIOps team plugs into.

Key Responsibilities
  • CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
  • Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
  • High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
  • Infrastructure as Code (IaC): Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
  • Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics. Feed the AIOps substrate with clean, well-labeled telemetry.
  • Internal Developer Platform (IDP): Build the paved road that lets product, model, and data-science teams deploy without opening a ticket. Make golden paths so obvious that shortcuts feel harder than doing it right.
  • Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency.
  • Governance, Security & Compliance: Establish and enforce system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance (e.g., SOC2, ISO
    27001).
  • Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, thorough RCA, and preventative remediation — and turn each incident into an automation that stops the next one before it pages a human.
Basic Qualifications
  • Experience &

    Education:

    Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in Dev Ops, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
  • Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
  • Containerization & Orchestration: Deep mastery of Docker and Kubernetes…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary