×
Register Here to Apply for Jobs or Post Jobs. X

Operations Engineering Manager; m​/f​/d

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Northern Data Group
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 90000 - 130000 GBP Yearly GBP 90000.00 130000.00 YEAR
Job Description & How to Apply Below
Position: Operations Engineering Manager (m/f/d)
Location: Greater London

Job Description The Operations Engineering Manager will own the operational reliability of the company’s GPU‑accelerated HPC infrastructure and lead a team of Operations Engineers. This role combines technical leadership with people management and Agile delivery ownership. The successful candidate will set the vision for operational excellence, manage and develop the internal team, and work closely with Platform, Network, and Infrastructure teams to build operational excellence.

The role will also play a key part in implementing and maturing the company’s scaled Agile Framework across Operations Engineering, ensuring alignment, transparency, and continuous improvement.

YOUR RESPONSIBILITIES Team Leadership & Management
  • Lead, coach, and develop an internal team of Operations Engineers.
  • Set clear goals, priorities, and expectations for the team.
  • Manage workload and resource allocation across the team.
  • Support regular 1:1s, performance reviews, and development plans.
  • Build a collaborative team culture focused on ownership and accountability.
Operational Ownership & Reliability
  • Own the reliability, performance, and availability of the GPU‑accelerated HPC infrastructure from an Operations perspective.
  • Oversee proactive system monitoring, incident trend analysis, and root cause analysis.
  • Define, track, and report on key operational metrics.
  • Ensure strong operational control through effective processes, runbooks, and change management.
  • Drive proactive improvements in reliability, automation, and performance.
Agile Ways of Working & Framework Oversight
  • Champion and oversee the scaled Agile Framework within the Operations function.
  • Collaborate with Product, Platform, and Network teams to align priorities and manage backlogs.
  • Support Agile ceremonies such as planning, stand‑ups, reviews, and retrospectives.
  • Improve delivery flow, predictability, and cross‑team coordination.
  • Ensure work is prioritised and delivered in line with Agile principles.
Process, Automation, and Documentation
  • Own the creation and maintenance of operational documentation, SOPs, and troubleshooting guides.
  • Drive automation initiatives to reduce manual effort and improve consistency.
  • Promote best practices in scripting, configuration management, and observability.
  • Maintain effective knowledge sharing across the team.
  • Stay informed on relevant trends and assess opportunities for improvement.
People & Stakeholder Communication
  • Provide regular status updates and reporting to leadership.
  • Represent Operations in cross‑functional planning and strategic discussions.
  • Act as an escalation point for major incidents and complex technical issues.
  • Coordinate with Platform, Network, and third‑party support teams during critical events.
  • Communicate clearly to support alignment, risk management, and decision‑making.
YOUR QUALIFICATIONS
  • Required 5+ years in infrastructure/operations, with 2+ years managing a technical team.
  • Advanced Linux administration in production, ideally at scale.
  • Proven experience running incident/problem management and working with third‑party or external support teams.
  • Hands‑on with automation (Ansible or equivalent) and monitoring/observability tools (e.g. Grafana, Prometheus).
  • Experience with Agile ways of working and exposure to scaled Agile frameworks.
  • Excellent communication and stakeholder management skills, able to work closely with Platform, Network, and leadership.
  • Nice to Have Experience in HPC or GPU‑accelerated environments (NVIDIA GPUs, Infini Band/RDMA, parallel file systems).
  • Scripting skills in Python and/or Bash for automation and tooling.
  • Understanding of performance tuning for HPC/GPU systems.
  • Experience with CI/CD pipelines and modern Dev Ops tooling.
  • Background designing or improving on‑call rotations, runbooks, and incident readiness.
WHAT WE OFFER
  • With us, you will work towards the future of HPC:
    From new, sustainable building methods for data centers to cooling concepts to software solutions for accelerated compute. Your approaches count:
    In official exchange formats or spontaneously at the coffee machine. At Northern Data, it's the best idea that counts - not the hierarchy. We’re looking forward to getting your inputs! You make the…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary