×
Register Here to Apply for Jobs or Post Jobs. X

Director: AI Systems Reliability, Testing & Performance

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Novartis
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    AI Evaluation, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 100000 - 186000 GBP Yearly GBP 100000.00 186000.00 YEAR
Job Description & How to Apply Below

Salary Range: £ - £

Job Description Summary

Director: AI Systems Reliability, Testing & Performance

#LI-Hybrid

Location:

London

Novartis is unable to offer relocation support for this role: please only apply if this location is accessible for you.

The Director, AI Systems Reliability, Testing & Performance is a senior technical AI leadership role within Data Science & AI, responsible for defining how AI systems across Novartis Development are evaluated, tested, validated, benchmarked, monitored, and continuously improved throughout their lifecycle.

This role owns the technical evidence required to determine whether AI systems are reliable, robust, secure, and fit-for-use across Novartis Development. The role defines common approaches for AI evaluation, benchmarking, testing, validation, monitoring and production-readiness across agentic AI systems, predictive models, retrieval systems, digital twins, and other AI capabilities.

The Director provides technical leadership in AI evaluation science, reliability engineering, validation, adversarial testing, and performance assessment. Key areas of focus include model and agent evaluation, benchmarking, failure-mode analysis, drift detection, digital twin validation, AI red teaming, observability, traceability, and technical evidence generation supporting regulated and business-critical AI systems.

This is not a Governance, Product Management, PMO, or infrastructure operations role. Governance owns policies, risk frameworks, and approval processes. Product teams own roadmaps, adoption, and value realization. DDIT and engineering teams own platforms, infrastructure, and operational services.

Success means Development AI systems are supported by objective evidence demonstrating how they perform, where they fail, and whether they remain fit-for-use over time.

Job Description Major Accountabilities AI Evaluation & Benchmarking

Define evaluation methodologies for AI systems, models, agents, digital twins, and simulation environments across Development.

Establish benchmark suites, evaluation datasets, and testing harnesses used across AI initiatives.

Define objective measures for quality, reliability, robustness, grounding, agent effectiveness, and task success.

Ensure evaluation approaches remain scientifically rigorous, reproducible, and comparable across AI systems.

Build a common evidence framework for assessing AI capabilities, limitations, and fitness-for-use.

AI Testing & Validation

Establish approaches for hallucination testing, failure-mode analysis, robustness testing, and behavioral validation.

Define validation methodologies for agentic systems, digital twins, simulation environments, and human-in-the-loop workflows.

Develop production readiness criteria for AI systems operating in Development environments.

Support technical validation activities required for GxP-relevant and regulated AI systems.

Ensure AI systems are tested under realistic operating conditions and supported by objective validation evidence.

Reliability, Monitoring & Performance

Define how AI system reliability, degradation, drift, and operational performance are measured over time.

Establish monitoring requirements and performance assessment frameworks across Development AI systems.

Partner with engineering teams to ensure required evaluation, monitoring, tracing, and observability signals are available.

Define methods for identifying, diagnosing, and assessing AI system failures.

Promote evidence-based improvement of deployed AI capabilities.

AI Security & Adversarial Testing

Define approaches for AI red teaming, adversarial testing, and security evaluation.

Establish testing methodologies for prompt injection, jailbreaks, retrieval attacks, tool misuse, and agent manipulation scenarios.

Assess…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary