Lead Cloud Software Engineer
Job in
Foster City, San Mateo County, California, 94420, USA
Listed on 2026-07-19
Listing for:
Coupa
Full Time
position Listed on 2026-07-19
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Software Architect, Cloud Engineer - Software, DevOps
Job Description & How to Apply Below
Job Description
- AI is reshaping how the world builds, learns, and connects. At our company, we believe the next generation of enterprise software will be powered by AI-driven platforms that are secure, scalable, and built for constant innovation.
- As our Lead of AI LLM Platform Dev Ops team, you will build solutions that enable this vision.
- From CI/CD pipelines, ML Ops, GenAI, Agentic AI platforms to AI governance and observability, you will define how our engineers build, test and ship the technologies that power global enterprises.
- This is your chance to lead & shape the evolution of Coupa’s AI platform. You will lead system design, code development, deployment, and automation for our LLM/GenAI/Agentic AI platforms ensuring reliability and scalability of infrastructure.
- Lead Platform Architecture and Strategy:
Define and drive the long-term technical vision for Agentic AI platform and LLM infrastructure, making strategic decisions on technology choices, architectural patterns and platform evolution. - Design Enterprise-Scale Platform Solutions:
Architect complex, multi-tenant platform systems that support high availability, scalability, and security requirements across global multi-cloud environments. - Drive Technical Leadership and Mentorship:
Lead and mentor a team of platform Dev Ops engineers, providing technical guidance, code reviews, and career development while fostering a culture of engineering excellence. - Establish Platform Engineering Standards:
Define and implement best practices, coding standards and architectural principles that ensure consistency, reliability, and maintainability across all platform components. - Evaluate new products and technologies such as RAG, MCP servers, AI Agents, and Agentic Workflows.
- Deployment and life cycle management of LLM & Embedding Models.
- Build and manage MLOps pipelines.
- Infrastructure as Code (IaC), e.g., Terraform, for multi-cloud deployments.
- Ensure technical quality through code reviews, architecture discussions, and engineering standards.
- Collaborate with the Principal Architect and AI/ML Engineering Lead on technical direction.
- Collaborate with platform, Dev, & QE teams to plan & deploy releases.
- Participate in design reviews, code reviews, troubleshooting incidents and write & review Incident RCAs.
- Manage AWS core & AI/GenAI services (S3, IAM, EKS, Bedrock, etc.) and be an expert in IaC (Terraform).
- Demonstrate architectural decision making:
Proven ability to make complex technical trade‑offs, evaluate emerging technologies, and drive architectural decisions that balance current needs with long‑term scalability. - Lead in fast‑moving environments, motivating teams and embracing innovation.
- Experience building enterprise‑grade, scalable, cost‑efficient, secure AI/ML platforms with a focus on resiliency, automation, and observability.
- Strong communication skills across geographies.
- Senior Platform Engineering Leadership: 10+ years of hands‑on platform engineering experience, including 3+ years in technical leadership roles, managing complex cloud infrastructure at enterprise scale.
- Advanced Automation and Programming:
Expert‑level programming skills in multiple languages (Python, Go, Java, etc.) with extensive experience building sophisticated automation frameworks and platform tools. - Enterprise Platform Design:
Proven track record designing and implementing large‑scale platform solutions, including microservices architectures, container orchestration, and distributed systems patterns. - Advanced Cloud Architecture Expertise:
Deep expertise in cloud‑native architectures, multi‑cloud strategies, and advanced services across AWS or GCP, with proven experience designing systems that handle millions of transactions. - BS/MS in Computer Science or equivalent experience.
- Deployment and life cycle management of LLM & Embedding Models (OpenAI, AWS Bedrock, Amazon Sage Maker) and vector databases like LanceDB.
- Strategic CI/CD and Dev Ops:
Deep expertise in enterprise CI/CD design, MLOps pipelines, and Dev Ops transformation initiatives, with experience implementing these practices across large engineering organizations. - Enterprise AI platform infrastructure development experience and exposure to evaluation of new products and technologies such as RAG, MCP servers, AI Agents, Agentic Workflows, Lang Chain, etc.
- Containerization and Orchestration:
Deep knowledge of Docker and Kubernetes (especially GKE or EKS) for scaling inference services, building and deploying workloads.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×