Engineering Manager, Automation and Service Reliability
Listed on 2026-08-18
-
IT/Tech
AI Engineer (Applied/Software), IT Project Manager
Please Note: This role is based at the Stanford Historic Campus, Pine Hall.
The Engineering Manager, Automation & Service Reliability leads a team of four engineers who design, build, and support the integration platforms, robotic process automation (RPA), and data services for Stanford University's IT Infrastructure team. This is a hybrid technical working manager and people leadership role. The manager actively contributes to technical projects, sets technical direction, and reviews engineering work while also leading the team's day-to-day people management, including coaching, performance management, hiring, and career development.
The manager partners closely with the Director of Communications Technologies Service Support and represents the team with internal University IT stakeholders and healthcare partners including Stanford Health Care (SHC) and Stanford Medicine Children's Health (SCH).
This role also carries a forward-looking mandate: evolve the team's automation practice beyond traditional RPA and integration work toward agentic AI. The manager will identify, pilot, and scale agentic AI solutions — AI agents and LLM-powered workflows that can reason, plan, and take multi-step action across existing systems — to solve problems for clients across the broader IT Infrastructure organization, not just the team's existing service lines.
Key Responsibilities People Management- Lead, manage and grow a team of one Technical Lead and three Software Engineers including training, goal-setting, performance reviews, and individual development plans.
- Recruit, hire, onboard, and mentor engineering talent; build a culture of technical excellence and accountability.
- Balance workload across the team, manage on-call/support coverage, and resolve interpersonal or performance issues.
- Guide architecture and design decisions across AI solutions, Mulesoft integrations, UiPath automations, Oracle APEX applications, and data pipelines.
- Review code and designs at a level sufficient to ensure reliability, security, and maintainability, without necessarily being the primary hands-on developer.
- Own service reliability for production systems (SIPdb, TDC, data feeds) including incident response, root-cause analysis, and preventive engineering.
- Establish and maintain an intake process for new automation requests, prioritizing the team's backlog against Service Now requests and strategic initiatives.
- Define and drive a roadmap for applying agentic AI — AI agents and LLM-powered workflows capable of multi-step reasoning and action — across the team's existing platforms (Mulesoft, UiPath, Oracle APEX) and newly identified use cases.
- Identify high-value opportunities for agentic AI across the IT Infrastructure organization, working with peer teams to surface pain points suited to AI-driven automation.
- Stand up pilots for agentic AI use cases (e.g., ticket triage and resolution, incident summarization, service desk drafting, data feed and reporting automation) and define success metrics before scaling to production.
- Evaluate AI tooling, platforms, and integration patterns (including LLM APIs and MCP-style tool/agent frameworks), and build organizational guardrails around security, data handling, and responsible use.
- Upskill the team in AI-assisted and agentic development practices, and champion the team's evolution toward an AI-forward automation practice (e.g., an Automation Center of Excellence model).
- Serve as the team's primary technical point of contact for IT Infrastructure, University Campus and Healthcare partners.
- Collaborate with the Director and broader Communications Technologies leadership on change management for next-generation voice and contact center initiatives.
- Communicate roadmap, risk, and status clearly to both technical and non-technical audiences.
- Lead and manage all business, technical, and education activities(e.g. application development, installing, configuring, and maintainingservers, routers, firewalls, workstations, and networkequipment).
- Exercise full management responsibility for a technical group,including…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).