×
Register Here to Apply for Jobs or Post Jobs. X

SRE L1 Support​/Cloud Platform Ops Engineer

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: Bitdeer Technologies Group
Full Time position
Listed on 2026-08-17
Job specializations:
  • IT/Tech
    IT Infrastructure, Cloud Computing: Infrastructure & Operations, IT Support, Systems Administrator
Salary/Wage Range or Industry Benchmark: 65000 - 90000 USD Yearly USD 65000.00 90000.00 YEAR
Job Description & How to Apply Below

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit

Position Overview

You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.

Neo Cloud is building an AI-operated GPU cloud. That doesn't mean fewer humans — it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for Neo Cloud's US GPU DCs during the 8AM–8PM PST shift. You execute SOPs, elevate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.

What

you'll own
  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (Service Now/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
Feed the AIOps substrate
  • Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation.
  • Every runbook you touch should get closer to being executable by the platform, not by you.
  • Your handoff notes are structured signal, not free-form email.
Why this role is different from a NOC job
  • You are not the last line of defense — the platform is. You are the training signal.
  • Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.
Job Requirement:
  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (Service Now, Jira Service Management)
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary