×
Register Here to Apply for Jobs or Post Jobs. X

Observability and Evaluation Engineer

Job in Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listing for: Mphasis
Full Time position
Listed on 2026-09-14
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer, AI QA / Validation Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Core Responsibilities

  • Build Observability Pipelines: Instrument applications and LLM/agent pipelines using telemetry tools (like Open Telemetry, Arize, Galileo, or Lang Smith) to capture execution traces, run logs, and latency metrics. [1]
  • Design Evaluation Frameworks: Create offline and online evaluation suites to benchmark accuracy, groundedness, toxicity, tool-use correctness, and reasoning-chain validity. []
  • Implement Regression & Drift Testing: Build automated test harnesses and continuous evaluation gates to detect model degradation, data drift, or output anomalies before releases reach production. [1]
  • Root-Cause Analysis: Investigate execution traces and multi-turn interaction failures to diagnose erratic system behaviors, API misparameters, or bottlenecks. []
  • Optimize Performance & Cost: Monitor and balance operational telemetry relating to token consumption, execution speed, and infrastructure costs. [1]
Key Skills & Requirements
  • Programming: Strong proficiency in Python or Go for building custom evaluation scripts and hooking into telemetry SDKs.
  • Tools & Stacks: Experience with Open Telemetry, Prometheus, Grafana, or dedicated AI observability platforms (e.g., Arize Phoenix, Lang Chain/Lang Smith, Galileo, Braintrust).
  • Testing Methodologies: Designing adversarial prompts, red-teaming protocols, and golden datasets for LLM/system validation.
  • System Design: Understanding of distributed microservices or LLM orchestration layers (Lang Chain, Llama Index, Auto Gen). [1, 2, 3]

An Observability and Evaluation Engineer (often specialized in AI/LLM systems) designs the tracking, tracing, and testing frameworks that monitor production behavior, measure output quality, and catch regressions in complex software or agentic AI workflows. [1, 2, 3, 4]

If you are tailoring this for a specific application, would you like me to focus this job description more heavily on traditional cloud/microservices observability or Generative AI / LLM agent evaluation
?

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary