Principal, AI Engineering Lead
Listed on 2026-09-01
-
Software Development
AI Engineer (Applied/Software), Software Architect
Principal AI Engineer
The Principal AI Engineer is a senior technical leader within the central AI Engineering function, sitting at the heart of a hub-and-spoke Data & AI organization. This role owns the design and delivery of the enterprise generative AI platform, the foundational infrastructure that enables every AI use case across investment and corporate teams. You will architect and build production-grade platform capabilities including a multi-LLM gateway, hybrid RAG retrieval services, agentic orchestration frameworks, model registry, prompt governance, MCP integrations, and the end-to-end sandbox-to-production deployment pipeline: all on Databricks on Azure with Unity Catalog governance and MNPI-compliant access controls required for a private equity environment.
This is a hands-on technical leadership role. You will write production code, define engineering standards consumed by vertical spoke teams, and partner directly with AI Product Management, Data Engineering, and business stakeholders across the firm. Use cases span the full enterprise: deal execution (CIM review, IC memo drafting), portfolio operations (covenant monitoring, performance analytics), investor relations (LP reporting, fund commentary), legal and compliance workflows, and internal productivity tooling.
You operate as the platform's technical authority, setting the patterns the organization builds on, not just reviewing them.
The AI Engineering function (hub) builds and operates the shared platform. Vertical spoke teams in each investment vertical consume platform services and build vertical-specific use cases within the guardrails the hub establishes. The Principal AI Engineer is the primary technical authority for platform architecture and a key collaborator to vertical AI engineers.
AI Platform Architecture & Engineering- Design and build the enterprise generative AI platform on Databricks on Azure, covering model serving, retrieval infrastructure, agent orchestration, and deployment pipelines
- Architect and operate a multi-LLM gateway (e.g., LiteLLM or equivalent) with routing logic, cost tracking, rate limiting, and model failover across Azure OpenAI and other providers
- Build hybrid RAG retrieval services: embedding models (e.g., BGE, OpenAI, or Cohere), Databricks Vector Search with Unity Catalog, structured extraction (Delta tables), and query routing across analytical, semantic, and hybrid modes
- Develop a reusable agentic orchestration layer using multi-stage patterns (e.g., orchestrator, section, and editor agents) with schema-constrained outputs, token budgets, and verbosity controls, generalized to serve use cases across deal execution, portfolio operations, LP reporting, legal review, and productivity workflows
- Implement a model registry, prompt library, and A2A (agent-to-agent) workflow framework as reusable platform primitives
- Build and maintain the data gateway link: integrating AI retrieval services with Gold-layer data products from the Data Engineering function
- Establish sandbox-to-production deployment pipelines for AI use cases, including evaluation frameworks, staged rollout, and rollback capabilities
- Implement MNPI controls and information barrier enforcement at the retrieval and inference layer, ensuring deal-context separation across verticals in Unity Catalog
- Design audit logging, access controls, and retrieval permissioning aligned with Legal, Compliance, Risk, and Cyber governance requirements
- Support the AI governance gate process: producing technical evidence packages for ARB and regulatory sign-off on new use case deployments
- Maintain observability across all LLM calls (e.g., Langfuse or equivalent): latency, token spend, retrieval quality, hallucination flags, and per-use-case cost attribution (AI Fin Ops)
- Define reusable platform patterns, APIs, and SDKs that vertical spoke teams consume to build investment-specific AI use cases without rebuilding core infrastructure
- Provide technical guidance and code review to vertical AI engineers, enforcing platform standards for chunking strategy, embedding choice, retrieval patterns, and agent design
- Collaborate with AI Product Management on use case intake, feasibility assessment, and translating business workflows into platform capabilities spanning investment, operations, compliance, and firm-wide functions
- Own the engineering standards for the AI platform across the organization: LLM integration patterns, RAG architecture conventions, agent design principles, evaluation criteria, and prompt governance
- Drive technical decisions on model selection, framework adoption, and infrastructure rationalization maintaining a lean, production-grade stack
- Mentor senior and mid-level AI engineers; serve as the escalation point for cross-cutting technical problems affecting multiple verticals or platform stability
- Produce ARB-ready architecture artifacts: reference diagrams,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).