More jobs:
Platform Site Reliability Engineer
Job in
Abu Dhabi, UAE/Dubai
Listed on 2026-06-05
Listing for:
Dicetek LLC
Full Time
position Listed on 2026-06-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Site Reliability Engineer (SRE) Role Overview
We are seeking an Site Reliability Engineer to own the "Production Readiness" of our cloud-based AI solutions. This hybrid role combines automated software testing and Site Reliability Engineering (SRE). You will build the automated frameworks that validate our AI outputs and ensure the underlying Azure/AWS infrastructure is resilient, performant, and compliant with banking standards.
Key Responsibilities- Resiliency Engineering (SRE):
Implement "Chaos Engineering" and load testing to ensure web/mobile backends can handle banking-scale traffic. Maintain high availability through automated recovery scripts. - Automated Regression:
Build CI/CD‑integrated test suites using Python that validate both the application logic and the infrastructure state (IaC validation). - Observability & SLIs:
Define and monitor Service Level Indicators (SLIs) and Objectives (SLOs). Set up advanced alerting in Azure Monitor or AWS Cloud Watch to catch performance degradation before users do. - Security & Compliance Testing:
Automate security scans and compliance checks to ensure all AI data handling meets strict banking data residency and privacy protocols.
- Automation Stack:
High proficiency in Python (for AI testing) and framework automation (PyTest, Selenium, or Robot Framework). - Cloud Infrastructure:
Strong hands‑on experience with Azure or AWS, specifically regarding networking, scaling, and serverless reliability. - AI/ML Understanding:
Understanding of Prompt Engineering and how to evaluate AI model outputs (RAG evaluation, ROUGE/BLEU scores, or custom LLM‑benchmarks). - Monitoring Tools:
Experience with Grafana, Prometheus, or native cloud monitoring tools to build real‑time reliability dashboards. - Fin Ops Awareness:
Ability to identify expensive failing tests or inefficient cloud resource usage during the testing phase.
- Languages:
Python (Mandatory), Bash scripting. - Tools:
Git Hub Actions (CI/CD), Terraform (reading/validating), K6 or JMeter (Performance). - AI Frameworks:
Deep Eval, Ragas, or Lang Smith (for automated AI evaluation).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×