×
Register Here to Apply for Jobs or Post Jobs. X

AIOps​/Observability Engineer

Job in Phoenix, Maricopa County, Arizona, 85003, USA
Listing for: Cloud Security Web
Full Time position
Listed on 2026-08-29
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below
Position: AIOps / Observability Engineer

We are seeking an experienced AIOps / Observability Engineer to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred.

Key Responsibilities
  • Design, implement, and maintain enterprise observability and monitoring solutions.
  • Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.
  • Build and integrate APIs to connect monitoring, alerting, and automation platforms.
  • Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.
  • Automate operational tasks using scripting languages and infrastructure-as-code tools.
  • Monitor production environments to ensure high availability, performance, and reliability.
  • Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.
  • Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.
  • Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.
  • Implement best practices for observability, logging, tracing, and performance monitoring.
  • Participate in on-call production support and continuous service improvement initiatives.
Required Skills
  • 8+ years of experience in AIOps, Observability, SRE, or Production Support.
  • Strong experience with observability and monitoring platforms such as Dynatrace, Splunk, App Dynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack
    .
  • Hands-on experience with REST APIs and API integrations.
  • Experience with cloud platforms including AWS
  • Strong scripting and automation experience using Python, Shell, Power Shell, Ansible, Terraform, or similar technologies
    .
  • Experience with incident management, troubleshooting, and root cause analysis.
  • Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.
  • Experience supporting mission-critical production environments.
  • Excellent analytical, communication, and problem-solving skills.
Preferred Qualifications
  • Experience with CI/CD pipelines and Dev Ops practices.
  • Knowledge of Kubernetes, Docker, and container observability.
  • Familiarity with ITSM tools such as Service Now.
  • Experience implementing AI/ML-driven monitoring and predictive analytics.
Nice to Have
  • Kafka or messaging platform monitoring.
  • Infrastructure as Code (Terraform/Cloud Formation).
  • Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary