Observability & Evaluation Engineer; NC)
Listed on 2026-09-06
-
Software Development
AI Reliability/ Performance Engineer, AI Engineer (Applied/Software), DevOps
Observability & Evaluation Engineer
We are seeking an experienced Observability & Evaluation Engineer to build and operationalize observability, telemetry, and evaluation capabilities for LLM applications and AI agent systems. The engineer will implement end-to-end tracing, metrics, dashboards, automated evaluation suites, SLOs, runbooks, and production-readiness evidence for priority AI and agent releases.
The ideal candidate will have strong experience with LLM/agent evaluation, observability, Python, test automation, production operations, and model/prompt performance analysis.
Key Responsibilities- Implement telemetry, logging, metrics, and distributed tracing for LLM applications and AI agents.
- Build dashboards to monitor AI application health, model behavior, agent performance, latency, errors, and usage.
- Develop automated LLM and AI agent evaluation suites for quality, reliability, safety, and performance.
- Create reusable Python-based test automation and evaluation frameworks.
- Analyze prompt and model performance and identify optimization opportunities.
- Define and monitor SLIs, SLOs, operational metrics, and alerting for AI services.
- Develop runbooks for troubleshooting, incident response, and production support.
- Establish production-readiness criteria and maintain release-readiness evidence for priority agent releases.
- Integrate evaluation and observability checks into CI/CD pipelines.
- Troubleshoot production issues involving model performance, latency, availability, failures, and unexpected agent behavior.
- Collaborate with AI/ML engineers, SRE, platform, product, and governance teams.
- Strong Python development experience.
- Hands-on experience with LLM, Generative AI, or AI agent evaluation.
- Experience developing automated test and evaluation frameworks.
- Strong knowledge of telemetry, metrics, logging, and distributed tracing.
- Experience with monitoring dashboards and alerting.
- Understanding of SLOs, SLIs, and production operations.
- Experience analyzing prompt and model performance.
- Strong troubleshooting and production support experience.
- Open Telemetry
- Prometheus / Grafana
- Datadog / Splunk
- LLM observability and evaluation platforms
- AI agent and multi-step workflow monitoring
- CI/CD
- AWS, Azure, or GCP
- Kubernetes / Docker
- API testing and microservices
- SRE and incident management practices
Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).