More jobs:
Senior Site Reliability Engineer; SRE, Compute Node Team
Job in
1011, Amsterdam, North Holland, Netherlands
Listed on 2026-09-23
Listing for:
Jobgether
Full Time
position Listed on 2026-09-23
Job specializations:
-
IT/Tech
SRE/Site Reliability, Unix/Linux, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (SRE, Compute Node Team) based in Netherlands.
This is a senior Site Reliability Engineering role focused on the infrastructure that runs and manages virtual machines across a large-scale cloud platform.
You will work close to the Linux operating system, hypervisor, and node-level services that form the foundation of the compute environment.
The role combines deep Linux systems engineering, virtualization, containerization, observability, and production reliability.
You will investigate complex issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance across user and kernel space.
You will also help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortem processes.
Collaboration with platform, kernel, hypervisor, GPU, and infrastructure teams will be central to improving system design and operability.
This is an opportunity to influence critical compute infrastructure supporting demanding AI and cloud workloads at significant scale.
Accountabilities
Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.
Analyze and debug complex Linux systems across both user space and kernel space.
Investigate system capabilities, limitations, dependencies, and trade-offs across different layers of the operating system and infrastructure stack.
Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling.
Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.
Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.
Support and improve containerized workloads using Linux-native mechanisms such as name spaces and cgroups.
Design and evolve observability for the compute node layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
Build reliability signals that provide clear and actionable insight into system behavior.
Lead or contribute to incident response, ensuring production issues are diagnosed and resolved efficiently.
Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.
Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.
Identify opportunities to automate operational processes and improve the resilience of compute infrastructure.
Collaborate closely with platform, kernel/hypervisor, GPU, and infrastructure teams on system design and operational improvements.
Contribute to improving the operability, scalability, and maintainability of node-level services.
Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.
Help establish reliability and observability as core capabilities of the compute platform.
Requirements
Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
Deep expertise in Linux, including strong understanding of both user space and kernel space.
Knowledge of important Linux kernel subsystems, including scheduling, memory management, file systems, cgroups, and name spaces.
Strong understanding of system boundaries, constraints, dependencies, and trade-offs across different infrastructure layers.
Hands-on experience with QEMU/KVM and a solid understanding of virtualization technologies.
Understanding of virtual machine life cycles, performance characteristics, resource management, and failure modes.
Practical experience with…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×