More jobs:
Senior Site Reliability Engineer – AI & Automation
Job in
Miami, Miami-Dade County, Florida, 33222, USA
Listed on 2026-09-04
Listing for:
Veriipro
Full Time
position Listed on 2026-09-04
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Primary Responsibilities
- Design and implement AI agents, LLM-based solutions, and operational automation for SRE and Dev Ops environments.
- Support production systems, perform incident response, troubleshooting, and 24x7 operational triage.
- Develop automation and integrations using Python.
- Build and maintain monitoring and observability solutions using Splunk, Datadog, New Relic, or App Dynamics.
- Design and manage cloud infrastructure across AWS, Azure, and/or GCP.
- Implement Infrastructure as Code (IaC) using Terraform, Cloud Formation, and/or AWS CDK.
- Apply AI/ML, prompt engineering, and agentic AI concepts to operational use cases.
- Develop runbook automation, anomaly detection, and intelligent incident management capabilities.
- Collaborate with engineering, operations, and business stakeholders to improve system reliability and operational processes.
- Mentor engineers and help establish SRE, Dev Ops, automation, and observability best practices.
- 5+ years of SRE, Dev Ops, or Site Reliability Engineering experience.
- Strong Python development skills for automation, tooling, and integrations.
- Experience with AI agents, LLMs, and agentic AI frameworks.
- Experience with Splunk, Datadog, New Relic, App Dynamics, or similar observability platforms.
- Hands-on experience with AWS, Azure, and/or GCP.
- Experience with Terraform, Cloud Formation, and/or CDK.
- Strong production support and incident response experience.
- Knowledge of AI/ML, prompt engineering, and operational AI automation.
- Strong communication, stakeholder management, and mentoring skills.
- Ability to work effectively in a 24x7 operational environment.
- Experience with Adobe Experience Manager (AEM).
- Experience supporting CMS platforms and digital applications.
- Knowledge of incident management, runbook automation, and anomaly detection.
- Familiarity with Atlassian Rovo, AI operational tooling, and modern observability platforms.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×