×
Register Here to Apply for Jobs or Post Jobs. X

Senior Infrastructure & Site Reliability Engineer – Datacentre AI Engineering KSA

Job in Riyadh, Riyadh Region, Saudi Arabia
Listing for: Qualcomm
Full Time position
Listed on 2026-09-01
Job specializations:
  • IT/Tech
    SRE/Site Reliability, AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 420000 - 660000 SAR Yearly SAR 420000.00 660000.00 YEAR
Job Description & How to Apply Below
Position: Staff/Senior Staff Infrastructure & Site Reliability Engineer – Data centre AI Engineering KSA

Company

Qualcomm Middle East Information Technology Company LLC

Job Area

Engineering Group, Engineering Group >
Software Test Engineering

About the Role

The role focuses on the configuration, operation, and continuous improvement of large-scale AI inference systems in a datacenter environment. The engineer will support critical AI use cases by ensuring Qualcomm’s AI infrastructure is reliable, scalable, and production-ready for advanced machine-learning workloads.

The role requires strong systems and software engineering fundamentals, hands‑on execution, and the ability to work independently on complex problem areas while collaborating closely with cross‑functional teams across hardware, software, and machine learning.

Ideal candidate will have 8+ years of experience in SRE

Key Responsibilities Will Include
  • AI Infrastructure
  • Define and apply best-practice configurations, deploy, and operate large-scale AI inference systems supporting critical AI workloads.
  • Ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
  • Develop and maintain software tools and support infrastructure around AI software stacks.
  • AI & ML Engineering
  • Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads.
  • Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
  • Work with models, systems, and software teams to improve model performance on AI100 deployments.
  • Identify and implement optimizations for workloads running on multi‑SoC and multi‑card systems.
  • Site Reliability Engineering (SRE)
  • Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization.
  • Support production ML systems using MLOps tools and operational best practices.
  • Contribute to incident reviews, operational documentation, and continuous reliability improvements.
  • Observability & Tooling
  • Define and implement optimal configuration strategies for large-scale AI inference systems, including deployment, operation, and support of critical AI workloads and maintain observability tools, dashboards, and alerts to monitor system health and reliability.
  • Monitor infrastructure and services using tools such as Prometheus, Grafana, Cloud Watch, and custom telemetry.
  • Create and maintain technical documentation, runbooks, and knowledge‑base articles.
  • Automation & CI/CD
  • Develop automation to reduce manual operational tasks and improve system reliability.
  • Support CI/CD pipelines for AI service and agent deployment.
  • Apply Infrastructure‑as‑Code practices using tools such as Terraform and Ansible.
Required Skillset
  • AI & Deep Learning
  • Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems.
  • Understand ML inference concepts including batching, token streaming, and performance considerations.
  • Hands‑on experience with PyTorch and familiarity with modern ML frameworks.
  • Familiarity with distributed inference, checkpointing, and accelerator‑based compute environments.
  • AI Operations
  • Experience supporting AI or ML applications in production environments.
  • Familiarity with LLM inference pipelines and AI service operations.
  • Programming, Software Configuration & Configuration Strategy
  • Strong programming skills in Python with experience building and supporting production systems.
  • Experience with scripting and automation using Python and Bash.
  • Familiarity with configuration management, orchestration tools, and best‑practice configuration strategies for production AI systems.
  • Systems & Infrastructure
  • Strong Linux fundamentals include shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
  • Experience working with cluster schedulers such as Slurm or equivalent systems.
  • Experience operating distributed systems with high availability and fault tolerance.
  • Observability & Monitoring
  • Hands‑on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
  • Understanding of incident management, service health metrics, and system reliability monitoring.
  • Dev Ops & SRE Practices
  • Solid understanding of SDLC, release processes, and operational reliability practices.
  • Familiarity with CI/CD pipelines…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary