×
Register Here to Apply for Jobs or Post Jobs. X

AI​/ML Observability Engineer

Job in Fort Worth, Tarrant County, Texas, 76102, USA
Listing for: StradIT
Full Time position
Listed on 2026-07-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, DevOps, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 130000 - 170000 USD Yearly USD 130000.00 170000.00 YEAR
Job Description & How to Apply Below

VISA INDEPENDENT CANDIDATES ONLY !

Overview

We are seeking a passionate and hands‑on AI/ML Engineer to accelerate our Enterprise Observability strategy. This role will design, build, and operationalize AI/ML capabilities that enhance end to end telemetry pipelines, anomaly detection, intelligent alerting, and proactive system resiliency.

You will work at the intersection of AI/ML engineering, Observability platforms, and automation, developing solutions that improve detection, diagnosis, and prevention of operational issues across distributed systems.

Key Responsibilities
  • Design and deploy AI/ML models supporting anomaly detection, baselining, event correlation, and predictive operational analytics.
  • Build and integrate AI‑enabled capabilities into enterprise Observability platforms, including Grafana, APM/RUM tools, network telemetry systems, and data observability tools.
  • Develop AI Agents that can autonomously triage issues, recommend corrective actions, and initiate automated remediation workflows to reduce recovery time and improve system resilience.
  • Implement self‑healing automation using AI‑driven decisioning, integrating with orchestration frameworks, service APIs, and infrastructure automation pipelines.
  • Engineer and maintain real‑time and batch data pipelines using Snowflake ML Jobs, Snowflake Cortex, streams, tasks, and UDFs.
  • Implement and manage Open Telemetry‑based telemetry ingestion for logs, metrics, traces, and spans across distributed systems.
  • Build asynchronous Python APIs and services for model inferencing and operational integration.
  • Enhance observability intelligence with AI‑powered capabilities such as root‑cause acceleration, chatbot/search enablement, and automated insights.
  • Contribute to SLO/SLI modeling, Golden Signals instrumentation, and Observability NFR adoption.
  • Collaborate across engineering, SRE, platform and business teams to embed proactive intelligence and Observability standards throughout the ecosystem.
Required

Skills & Qualifications Core Technical Skills
  • Strong proficiency in Python and data science/ML libraries:
    • Num Py
    • Pandas
    • scikit learn
    • Tensor Flow
    • Py Torch
    • Matplotlib
    • Seaborn
  • Experience with Generative AI, LLM fine tuning, prompt engineering, RAG pipelines, and LLM evaluation frameworks.
  • Expertise in developing and deploying ML models in production (batch & streaming).
  • Strong understanding of statistics, time series modeling, and anomaly detection.
Observability & Telemetry
  • Experience with Open Telemetry for logs, metrics, traces, spans.
  • Familiarity with Observability concepts:
    • Golden Signals
    • SLO/SLI design
    • APM
    • RUM
    • Synthetics
    • event correlation
    • baselining
  • Experience with Observability tools such as:
    • Grafana (Alloy agents, dashboards, ML capabilities)
    • Dynatrace
    • Monte Carl0 (Data Observability)
    • Netscout
    • Thousand Eyes
    • Solar Winds
    • Net Brain
Cloud, Data & Platform
  • Hands on with AWS (Sage Maker, Bedrock), Snowflake ML, Snowflake/Openflow, Snowflake AI Observability tooling.
  • Experience building Snowflake data pipelines (streams, tasks, UDFs) – plus for Cortex features.
  • Strong understanding of distributed systems and microservices telemetry requirements.
Automation & Engineering Quality
  • Experience with automation pipelines, CI/CD, and infrastructure as code patterns supporting Observability adoption.
  • Ability to build asynchronous Python APIs or services for model inference and operational integration.
Preferred Qualifications
  • Experience developing agentic AI systems that analyze telemetry, generate action recommendations, or execute automated operational responses.
  • Experience building self‑healing patterns, including automated rollback, service restarts, configuration corrections, and predictive maintenance.
  • Experience in Snowflake ML workflows, Snowflake Cortex Agents, and data pipeline automation.
  • Exposure to AI‑enabled alerting, RCA automation, and operational self‑healing concepts.
  • Experience with large‑scale operational telemetry and multi‑cloud ecosystems.
Soft Skills
  • Strong analytical thinking and problem solving.
  • Excellent communication skills for cross functional collaboration with infrastructure, SRE, engineering, business, and leadership teams.
  • Curiosity, continuous learning mindset, and passion for applied AI and Observability.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary