More jobs:
SRE AIOps Engineer
Job in
Miami, Miami-Dade County, Florida, 33101, USA
Listed on 2026-09-04
Listing for:
Apex Systems
Full Time
position Listed on 2026-09-04
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Engineer (Applied/Software)
Job Description & How to Apply Below
SRE AIOps Engineer
Location:
Miami, Florida (Hybrid)
Role Overview
The AIOps Engineer operates in a dual capacity, designing and building AI-powered automation and tools that streamline SRE workflows, while also providing specialized L2 triage for application-specific technologies including Adobe Experience Manager (AEM) and other platforms. This role is responsible for an LLM wrapper project or AI modeling project, potentially overseeing a team. The ideal candidate is equally comfortable writing production-grade automation code and triaging complex application-specific incidents.
Key Responsibilities
- Design and implement AI/ML and LLM-based solutions to automate incident triage, generate summaries, and recommend remediation actions.
- Develop Python-based services, CLIs, and scripts to automate repetitive SRE tasks and process logs, metrics, and events.
- Implement and manage Infrastructure as Code (Terraform, Cloud Formation) for SRE tooling and observability.
- Work with internal AI tools and Atlassian Rovo to advance AI-driven operational capabilities.
- Serve as an L2 point of contact for AEM Sites, AEM Assets, Adobe Target, and related integrations.
- Perform initial triage of application issues, including reviewing error logs, checking system behavior, and validating configurations.
- Provide L3/engineering teams with detailed incident descriptions, impact analysis, and steps to reproduce.
- Monitor health dashboards and alerts from tools like Splunk and New Relic to proactively identify anomalies.
- Maintain and refine runbooks, knowledge base articles, and standard operating procedures for common incidents.
Required Qualifications
Experience:
- 6–8 years in a lead SRE, Dev Ops, or related role with a strong emphasis on automation and tooling.
- Experience with either SRE or AEM is required.
- Hands-on experience with AI/ML/LLM-based solutions and tools is highly desirable.
- 3–6 years supporting CMS-based platforms.
Technical
Skills:
- Strong Python development skills for services, CLIs, and integrations.
- Proficiency with Infrastructure as Code (IaC), specifically Terraform.
- Advanced use of monitoring and observability tools such as Splunk or New Relic.
- Experience with monitoring and supporting applications.
- Familiarity with Atlassian Rovo (Jira or Confluence).
- Proficiency with cloud platforms (AWS, Azure, or GCP).
- Understanding of AEM concepts, including Sites, authoring/publishing flows, and diagnosing application issues.
Education:
- A bachelor's degree in Computer Science, IT, Computer Engineering, or a related field is preferred.
Preferred Qualifications
- Experience with both SRE and AEM.
Work Environment
- This role follows a hybrid work model, requiring onsite presence in Miami, Florida, from Monday to Thursday, with the option to work remotely on Friday.
- Participation in a 24x7 on-call rotation is required, occurring approximately one to two times per month, including evenings, weekends, and holidays.
- The position supports a global, always-on multi-brand eCommerce platform, requiring readiness to respond to high-severity incidents outside standard working hours.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×