More jobs:
Platform Engineer
Job in
Sunnyvale, Santa Clara County, California, 94087, USA
Listed on 2026-09-24
Listing for:
GCB Services LLC
Full Time
position Listed on 2026-09-24
Job specializations:
-
IT/Tech
Unix/Linux, Systems Engineer, SRE/Site Reliability, IT Support
Job Description & How to Apply Below
Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence. Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies. Participates in team stand-ups on projects, incidents, and daily priorities. Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring. Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.
Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset. Must have:
Architecture, design patterns, reliability, and scaling of new and existing systems. Incident command experience - driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through. Observability built from the ground up - defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do. Linux kernel internals - scheduler, memory allocation, driver subsystems. High-quality code in at least one language (Python, Go, or similar).
System-level debugging - kdump, kernel panic analysis. IaC (Ansible, Terraform, Kubernetes) and CI/CD (Git Lab CI, AWX, etc.) for bare-metal or cloud infrastructure. TCP/IP and network programming. Distributed storage systems - object, block, and/or file storage paradigms. Strong communication skills. Nice to have:
Hardware and GPU troubleshooting. OVN/OVS-based networking stack exposure. Sourcing note:
This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical "L1" label implies - screen for genuine engineering depth, not helpdesk/NOC-tier breadth. Role 2 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team:
Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have:
Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: o Storage: block, blob, and file storage; distributed storage;
performance diagnostics and data-path optimization. Critical screening filter - operator vs. SRE:
Crusoe explicitly does not want another "operator" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure. Role 3 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team:
Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have:
Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: o Compute:
Linux systems, KVM/QEMU, Cloud Hypervisor, kernel tuning, CPU/memory/VM optimization. Critical screening filter - operator vs. SRE:
Crusoe explicitly does not want another "operator" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure. Role 4 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team:
Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have:
Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: SDN: OVS/OVN, network…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×