Associate MLOps Engineer
AppliedAI is a pioneering AI technology company headquartered in Abu Dhabi, UAE. We are committed to innovation and excellence in artificial intelligence solutions across regulated industries such as healthcare, insurance, government, and financial services.
AppliedAI is at the forefront of redefining the future of work through cutting‑edge AI solutions. We empower organizations by automating complex, document‑heavy processes, ensuring unparalleled efficiency and accuracy. Our commitment to synergizing human intelligence with artificial intelligence helps our clients excel in highly regulated industries, including healthcare, finance, and insurance.
We're seeking an Associate ML Ops Engineer to join our growing team. In this role, you'll work at the intersection of operations and development, supporting our platform's reliability, performance, and security while helping our development teams build and maintain scalable solutions. This is a strong opportunity for someone early in their SRE/MLOps career to build technical depth alongside a senior team.
Key Responsibilities- Help monitor and maintain production, development and staging environments, supporting high availability and performance of our architecture
- Collaborate with Dev Ops, MLOps and Development teams to help troubleshoot and resolve issues
- Support our observability stack, with guidance from senior engineers
- Help maintain and improve our incident response processes
- Assist with capacity planning and performance optimization for systems handling up to 10,000 requests per minute
- Support compliance with security standards and regulatory requirements across our infrastructure
- Participate in on‑call rotation during regular business hours, with senior engineer backup for escalations
- Help implement and maintain SLOs, SLIs, and SLAs
- Support cloud cost optimization efforts alongside system performance and reliability
- Support our CI/CD pipelines and deployment processes
- Familiarity with infrastructure best practices such as the AWS and Azure Well‑Architected Framework
- Basic understanding of infrastructure patterns for high availability, fault tolerance, and disaster recovery
- Understanding of infrastructure security fundamentals (principle of least privilege, network segmentation, encryption at rest/in transit)
- Awareness of infrastructure compliance and governance frameworks
- Interest in cost optimization strategies and Fin Ops practices
- Exposure to Infrastructure as Code concepts (modularity, reusability, versioning)
- Understanding of observability patterns (logging, metrics, tracing)
- 1-2 years of experience in SRE, Dev Ops, or a similar role (internships and hands‑on project experience will be considered)
- Some experience with AWS services, including:
- Compute:
Lambda, ECS, Fargate - Networking: ALB, ELB, API Gateway, Route
53, Cloud Front, App Sync - Messaging:
Event Bridge, SNS, SQS - Security:
Security Groups, Secrets Manager (SM), Systems Manager (SSM), IAM - Exposure to monitoring and observability tools
- Basic knowledge of infrastructure as code (CDK and/or Terraform)
- Understanding of event‑driven architecture concepts
- Some experience with containerization and microservices
- Basic scripting and automation skills
- Good problem‑solving abilities and a systematic approach to debugging
- Experience working in Agile environments is a plus
- Experience with Next.js, Node.js and Python
- Familiarity with authentication systems (Auth0, SSO)
- Awareness of regulatory compliance requirements (SOC 2, HIPAA, GDPR, PCI DSS)
- Interest in ML/LLM operations
- Exposure to multi‑region AWS deployments
- Interest in working with high‑traffic systems
- Basic experience with database management and optimization
- Aw…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).