×
Register Here to Apply for Jobs or Post Jobs. X

Technical Service Operations Lead; TSO Lead), Germany

Job in New York City, Richmond County, New York, USA
Listing for: Xsolla
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Support
Job Description & How to Apply Below
Position: Technical Service Operations Lead (TSO Lead), Germany

Technical Service Operations Lead (TSO Lead)

We are looking for an operationally driven, collaborative, analytical, and strong communicator to join our Global Technical Operations (GTO) team. The best candidate will thrive in a fast-paced, highly collaborative, and dynamic setting and is excited to help coordinate incident response alongside cross-functional teams, identify trends and patterns in production issues, improve how we communicate with partners during incidents, and drive continuous improvement through post-incident reviews.

Strong incident management experience, ITIL knowledge, and observability/monitoring expertise are essential, along with experience in technical operations, SRE, or NOC environments supporting high-availability platforms in the gaming industry. The ability to communicate clearly and effectively in English — both written and verbal — across technical and executive audiences will be key to your success in this role.

If you're passionate about driving operational excellence and platform reliability at scale and love ensuring the reliability and uptime of commerce and payment solutions that game developers and players depend on, we would love to hear from you!

Responsibilities:

  • Serve as Incident Commander for major incidents — coordinating cross-functional response teams, driving investigation, making escalation decisions, and ensuring incidents are resolved within SLA targets.
  • Own all incident communications: draft and send clear, timely updates to senior leadership, Customer Success, and partner/customer contacts throughout the incident lifecycle, and manage customer-facing status page updates ().
  • Facilitate blameless Post-Incident Reviews (PIRs) for major incidents — leading root cause identification, assigning corrective actions with clear owners and deadlines, and tracking them to closure.
  • During non-incident periods, proactively analyze incident trends, recurring issues, and production bugs — identify patterns, create Problem tickets, and report findings and recommendations to product and engineering teams on a regular cadence.
  • Enforce the incident management framework across the organization, including the severity model, priority matrix, SLA targets, escalation procedures, and deployment readiness gates.
  • Oversee and mentor the Operations Engineer on your shift — coaching on triage, investigation, runbook execution, and documentation quality while conducting regular knowledge transfer sessions to build depth across the service portfolio.
  • Produce shift handoff reports and deliver regular operational reporting: incident trends, KPI performance (MTTD, MTTA, MTTR), SLA adherence, proactive detection rate, and repeat incident analysis.
  • Audit service catalogue completeness on a regular cadence and govern JIRA Service Management workflows for incident, PIR, and problem management.
  • Cover for the Operations Engineer role during absences, breaks, or surge incidents. Participate in weekend on-call rotation for major incidents.

Qualifications:

  • Previous experience working at a gaming company is required — you understand the pace, player expectations, live operations dynamics, and the operational demands of the gaming industry.
  • 6+ years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high-availability, high-transaction systems.
  • Proven incident management experience — coordinating multi-team response, making real-time escalation decisions, and communicating with executive stakeholders under pressure.
  • Excellent written and verbal communication skills in English — ability to draft clear, concise executive updates at 3 AM under pressure, facilitate blameless PIRs, present operational metrics to senior leadership, and communicate incident status to customers and partners with clarity and professionalism.
  • Strong ITIL foundation — understanding of incident, problem, and change management life cycles with practical experience implementing or operating ITIL-aligned workflows.
  • Technical depth across the observability stack — ability to read and interpret logs, traces, and metrics in Datadog (or equivalent: Grafana, Splunk, New…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary