Principal AI/ML Engineer
Listed on 2026-09-26
-
Software Development
AI Engineer (Applied/Software), Software Architect, AI Reliability/ Performance Engineer
As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable value for our clients. You will lead the architecture, engineering, operationalization, and ongoing reliability of advanced AI systems, ensuring they can scale across enterprise environments while meeting rigorous standards for performance, security, resilience, and responsible AI.
This role sits at the critical intersection of AI research, engineering, product development, and operations. You will partner closely with world-class AI researchers, product leaders, and engineering teams to accelerate the journey from prototype to production. Your work will span some of the most advanced areas of AI, including Large Language Models (LLMs), Trustworthy AI, agentic systems, and emerging AI technologies.
You will mentor engineers, shape architecture, guide production support strategy, and serve as a thought leader for scaling AI across the organization. In addition to building and scaling AI solutions, you will help establish an engineering culture that emphasizes operational excellence, ownership, reliability, and continuous improvement.
Key Responsibilities:AI Architecture & Technical Leadership
- Define and lead the technical architecture for enterprise-scale AI and ML platforms.
- Design scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads.
- Establish architectural standards, engineering patterns, and best practices for AI deployment and operations.
- Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure.
- Partner closely with AI researchers to transform cutting‑edge prototypes into production‑grade solutions.
- Lead efforts to operationalize advanced AI capabilities across areas such as:
- Large Language Models (LLMs)
- Trustworthy and Responsible AI
- Agentic AI Systems
- Establish repeatable pathways that accelerate innovation‑to‑production cycles.
- Ensure production solutions maintain scientific rigor while meeting enterprise engineering standards.
- Bridge the gap between research breakthroughs and sustainable business value.
- Solve the organization's most complex AI engineering and scalability challenges.
- Design systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost.
- Drive adoption of MLOps, LLMOps, and AI platform engineering best practices.
- Improve the robustness, maintainability, observability, and operational readiness of our AI products.
- Identify and eliminate architectural bottlenecks that impact scale, reliability, or client experience.
- Raise standards through coaching, architecture reviews, design guidance, and technical leadership.
- Own the operational excellence, reliability, performance and availability of our products.
- Lead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages.
- Serve as the senior technical escalation point for the team's most challenging production challenges.
- Establish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs).
- Mentor and lead junior engineers in troubleshooting, root cause analysis, operational decision‑making, and incident response.
- Drive post‑incident reviews focused on learning, continuous improvement, and long‑term corrective actions.
- Develop operational processes that ensure AI solutions remain secure, scalable,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).