Lead ML Platform Engineer; SRE/FTE/Onsite
Listed on 2026-08-31
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Select how often (in days) to receive an alert:
Create Alert
Date: Aug 28, 2026
Location: Charlotte, NC, US
Company: NTT DATA Services
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization,
We are currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to join our team in Charlotte
, North Carolina (US-NC),
United States (US).
The Lead ML Platform Engineer provides architecture and hands-on engineering leadership for the Cortex Predictive AI Platform across cloud and on-premises environments. This role will establish and implement reusable, secure, scalable standards that enable data scientists, ML engineers, and application teams to build, validate, deploy, monitor, and operate predictive models efficiently and reliably.
The successful candidate will lead technical design and engineering decisions across the ML platform lifecycle, including governed data and feature access, model development environments, training and validation workflows, model delivery pipelines, real-time and batch inference, observability, reliability, and operational readiness. This role will also mentor engineering teams and transfer knowledge to support sustainable platform operations and adoption.
Key Responsibilities- Define and lead the target architecture for predictive AI and ML platform capabilities spanning public cloud and on-premises environments.
- Design, build, and operate reusable platform services supporting the end-to-end ML lifecycle: governed data and features, model development, training, validation, deployment, inference, monitoring, and operations.
- Establish scalable reference architectures, engineering standards, reusable templates, and implementation patterns for ML workloads across the Cortex portfolio.
- Lead platform engineering for GCP and multi-cloud environments, including secure connectivity, identity, network controls, compute, storage, and managed AI/ML services where applicable.
- Design and operate Kubernetes-based ML platforms using GKE, Open Shift, and associated container, workload orchestration, and resource-management capabilities.
- Implement and improve MLOps capabilities for experiment tracking, model packaging, validation, approval gates, model registry integration, deployment automation, rollback, and lifecycle management.
- Build CI/CD pipelines and infrastructure automation for platform services, ML workflows, model delivery, and environment provisioning.
- Enable model migration from legacy environments into standardized Cortex platform patterns, minimizing delivery risk and operational disruption.
- Engineer production-grade real-time and batch inference capabilities, including API-based serving, scalable runtime patterns, resiliency, performance, and operational support.
- Partner with data engineering, data governance, security, privacy, risk, model validation, and application teams to ensure data protection and control requirements are embedded into platform design.
- Implement platform observability, including logs, metrics, traces, dashboards, alerts, service-level indicators, service-level objectives, and operational runbooks.
- Drive reliability engineering practices for ML platform services, including capacity planning, high availability, disaster recovery, incident management, root-cause analysis, and continuous improvement.
- Ensure platform designs meet enterprise security requirements for authentication, authorization, secrets management, encryption, data access, auditability, and environment isolation.
- Provide technical leadership, architecture reviews, code reviews, design guidance, and mentoring to ML platform engineers and adjacent delivery teams.
- Produce clear technical documentation, reference implementations, operational procedures, and knowledge-transfer materials to enable self-service adoption and long-term support.
- 8+ years of experience in platform engineering, cloud engineering, infrastructure engineering, SRE, MLOps, or related technical roles.
- 4+ years of experience designing, building, or operating enterprise AI/ML or data platforms.
- Demonstrated experience leading architecture and engineering delivery for complex, production-grade cloud and/or on-premises platforms.
- Strong hands-on experience with GCP and working knowledge of multi-cloud or hybrid-cloud architecture.
- Experience with Kubernetes-based platforms, including GKE and Open Shift, in production environments.
- Strong experience implementing MLOps capabilities, model lifecycle workflows, or ML platform services.
- Proficiency in Python for platform automation, integration, operational tooling, or ML workflow development.
- Experience with CI/CD, Git-based development, automated testing, deployment automation, and infrastructure-as-code practices.
- Strong understanding of enterprise security, data protection, identity and access management, secrets management,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).