Digital - Principal SRE; AI Engineer
Job in
Columbus, Franklin County, Ohio, 43224, USA
Listed on 2026-07-26
Listing for:
Huntington
Full Time
position Listed on 2026-07-26
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Digital
- Principal SRE (AI Engineer)
Skip to main content#Digital
- Principal SRE (AI Engineer) page is loaded## Digital
- Principal SRE (AI Engineer)
Apply remote type:
Office locations:
Columbus, OHtime type:
Full time posted on:
Posted Todayjob requisition :
R0069689#
** Description
* **** Huntington will not sponsor applicants for this position for immigration benefits, including but not limited to assisting with obtaining work permission for F-1 students, H-1B professionals, O-1 workers, TN workers, E-3 workers, among other immigration statuses. Applicants must be currently authorized to work in the United States on a full-time basis.
*** The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for designing, deploying, and maintaining AI-driven solutions while ensuring the reliability, scalability, and performance of digital platforms and services. The ideal candidate will work closely with Digital SRE engineers, data scientists, Dev Ops, and operations teams to deliver robust, efficient, and automated systems that support business goals.
Job Description
*
* Summary:
** The IS Technical Specialist provides technical and consultative support on the most complex technical matters. This role typically reports to the Head of Digital SRE and may involve on-call responsibilities. The position provides opportunities to work on cutting-edge AI solutions, collaborate with cross segment teams, and drive reliability for mission-critical digital services
*
* Duties and Responsibilities:
*** Design, develop, and implement AI-driven systems and automation tools to enhance the reliability and efficiency of digital platforms.
* Monitor the health, availability, and performance of AI-enabled applications and infrastructure using SRE best practices.
* Collaborate with cross-functional teams to integrate machine learning models into production environments, ensuring seamless deployment and operation.
* Establish and enforce service-level objectives (SLOs), error budgets, and incident response procedures for AI-driven services.
* Identify, troubleshoot, and resolve complex incidents related to AI systems, leveraging observability and monitoring tools.
* Drive continuous improvement by analyzing post-incident reviews, automating manual tasks, and optimizing system performance.
* Stay up to date with advancements in AI, SRE, and cloud technologies, recommending innovative solutions to enhance digital reliability.
* Document processes and runbooks for operational transparency and knowledge sharing.
* AI Platform Integration:
Develop abstraction layers across AI providers (Google, OpenAI, etc. ) to enable seamless integration and enablement.
* Conduct design workshops, POCs, and code-with sessions to shape data-driven agent workflows with stakeholders, fostering trust and adoption.
* Measure & Improve:
Define and use key metrics, test harnesses, and evaluation plans to measure agent accuracy, latency, safety, and cost effectiveness.
* Knowledge Sharing:
Craft reusable patterns, documentation, and best practices to influence internal assets and client roadmaps.
*** Huntington will not sponsor applicants for this position for immigration benefits, including but not limited to assisting with obtaining work permission for F-1 students, H-1B professionals, O-1 workers, TN workers, E-3 workers, among other immigration statuses. Applicants must be currently authorized to work in the United States on a full-time basis.
****
* Basic Qualifications:
*** Bachelor’s degree in computer science, Engineering, Data Science, or a related field and experience.
* 5+ years Hands-on with AI/ML engineering, SRE, Dev Ops, or related roles. + Hands-on programming skills in Python, Java, or similar languages, with Hands-on with in developing and deploying machine learning models. + Hands-on with cloud platforms (e.g., AWS, GCP) and containerization technologies (Docker, Kubernetes). + Familiarity with observability tools (Prometheus, Grafana, ELK stack) and Service Now incident management platforms.
+ Solid understanding of SRE…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×