More jobs:
Observability and Evaluation Engineer
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-09-14
Listing for:
Mphasis
Full Time
position Listed on 2026-09-14
Job specializations:
-
Software Development
AI Reliability/ Performance Engineer, AI QA / Validation Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Core Responsibilities
- Build Observability Pipelines: Instrument applications and LLM/agent pipelines using telemetry tools (like Open Telemetry, Arize, Galileo, or Lang Smith) to capture execution traces, run logs, and latency metrics. [1]
- Design Evaluation Frameworks: Create offline and online evaluation suites to benchmark accuracy, groundedness, toxicity, tool-use correctness, and reasoning-chain validity. []
- Implement Regression & Drift Testing: Build automated test harnesses and continuous evaluation gates to detect model degradation, data drift, or output anomalies before releases reach production. [1]
- Root-Cause Analysis: Investigate execution traces and multi-turn interaction failures to diagnose erratic system behaviors, API misparameters, or bottlenecks. []
- Optimize Performance & Cost: Monitor and balance operational telemetry relating to token consumption, execution speed, and infrastructure costs. [1]
- Programming: Strong proficiency in Python or Go for building custom evaluation scripts and hooking into telemetry SDKs.
- Tools & Stacks: Experience with Open Telemetry, Prometheus, Grafana, or dedicated AI observability platforms (e.g., Arize Phoenix, Lang Chain/Lang Smith, Galileo, Braintrust).
- Testing Methodologies: Designing adversarial prompts, red-teaming protocols, and golden datasets for LLM/system validation.
- System Design: Understanding of distributed microservices or LLM orchestration layers (Lang Chain, Llama Index, Auto Gen). [1, 2, 3]
An Observability and Evaluation Engineer (often specialized in AI/LLM systems) designs the tracking, tracing, and testing frameworks that monitor production behavior, measure output quality, and catch regressions in complex software or agentic AI workflows. [1, 2, 3, 4]
If you are tailoring this for a specific application, would you like me to focus this job description more heavily on traditional cloud/microservices observability or Generative AI / LLM agent evaluation
?
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×