More jobs:
AI-Ops Engineering Lead - Director
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-08-13
Listing for:
Sumitomo Mitsui Financial Group, Inc.
Full Time
position Listed on 2026-08-13
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software)
Job Description & How to Apply Below
SMFG's shares trade on the Tokyo, Nagoya, and New York (NYSE: SMFG) stock exchanges.
In the Americas, SMBC Group has a presence in the US, Canada, Mexico, Brazil, Chile, Colombia, and Peru. Backed by the capital strength of SMBC Group and the value of its relationships in Asia, the Group offers a range of commercial and investment banking services to its corporate, institutional, and municipal clients. It connects a diverse client base to local markets and the organization's extensive global network.
The Group's operating companies in the Americas include Sumitomo Mitsui Banking Corp. (SMBC), SMBC Nikko Securities America, Inc., SMBC Capital Markets, Inc., SMBC MANUBANK, JRI America, Inc., SMBC Leasing and Finance, Inc., Banco Sumitomo Mitsui Brasileiro S.A., and Sumitomo Mitsui Finance and Leasing Co., Ltd.
Role Description
As the AI-Ops Engineering Lead in the Platform Engineering team, you will define and drive the strategy for operationalizing, monitoring, and governing the AI/GenAI platform and the models, pipelines, and agents that run on it. You will set the standards, reference patterns, and operating model that the broader engineering organization adopts, and partner with Azure, Databricks, and other infrastructure providers to build and operate the MLOps/LLMOps backbone of the platform.
As a senior technical leader, you will influence architecture, technology, data, and business stakeholders, shape the AI-Ops roadmap, and champion operational excellence-reliable, observable, secure, and cost-efficient AI systems-across the enterprise. This role begins as a hands-on technical leader and is expected to grow into building and leading a dedicated AI-Ops team over time.
This is a unique opportunity to own and lead the operational excellence of the GenAI technology stack-bridging the gap between one-off experiments and production-grade AI systems-while shaping the practices, governance, and culture that ensure industrial-grade reliability, compliance, and efficiency in a high-stakes financial environment.
Role Objectives
- Set AI-Ops strategy and standards:
Define the enterprise vision, operating model, and reference patterns for MLOps/LLMOps on Databricks and Azure Cloud Services, and drive their adoption across engineering, architecture, and data teams. - Operationalize the AI Platform:
Own the MLOps/LLMOps backbone for the AI platform, standardizing how models, prompts, pipelines, and agents are built, promoted, and run reliably in production. - Build CI/CD and release engineering:
Establish automated CI/CD pipelines and infrastructure-as-code for models, prompts, and agents using Databricks Asset Bundles across DEV/QA/REL/PROD, with canary, blue/green, shadow, and automated-rollback deployment strategies. - Own governance, versioning and auditability:
Implement end-to-end lineage and version control across data, prompts, retrievals, models, and responses using MLflow (Prompt Registry, Tracing, Experiments/Runs), delivering audit-ready artifacts and enforceable quality gates for internal and regulatory review. - Monitoring, drift and cost governance:
Build observability for data quality, data and model drift, retrieval and hallucination/grounding health, application performance, and business KPIs, with cost visibility, inference optimization, and Fin Ops-aligned governance. - Testing, evaluation and validation:
Establish automated regression, A/B, canary, shadow, and champion-challenger validation with golden datasets, evaluation rubrics, and human-in-the-loop review to certify quality and safety before and after release. - Drive responsible AI, security and governance adoption:
Partner with architecture, risk, security, and business leaders to embed responsible-AI guardrails (bias/harm detection, explainability, safety) and security/privacy controls into the enterprise path-to-production. - Operational readiness and run management:
Ensure reliable day-2 operations through model cards, API/SLA contracts, runbooks, incident response, and escalation readiness. - Evaluate emerging technology:
Proactively identify and evaluate emerging AI-Ops tooling and integrate those that improve reliability, observability, and cost efficiency. - Technical leadership and team building:
Mentor and uplift broader engineering teams on MLOps/LLMOps best practices, establish the AI-Ops discipline, and build and eventually lead a dedicated AI-Ops team as the function scales.
- Bachelor's degree in Computer Science, Machine Learning, Data Science, or related field (advanced…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×