Principal Technical Lead, Software Engineering; AI Infrastructure - North America Software Center
Listed on 2026-07-26
-
Software Development
AI Engineer (Applied/Software)
Join TSMC Washington and help power the future of technology. At TSMC, we don't just make semiconductors; we innovate to transform industries and enhance lives. As the world’s leading semiconductor foundry, we partner with top tech companies to drive advancements in industries such as healthcare, automotive, consumer electronics, and renewable energy. At TSMC Washington, you'll thrive where innovation meets precision manufacturing, and integrity guides our high standards and customer trust.
Our visionary leaders collaborate with clients to achieve groundbreaking results, ensuring our leadership in the semiconductor sector. Explore career opportunities with TSMC Washington and join a company with a commitment to excellence and innovation.
We develop next-generation AI infrastructure, cloud platforms, data infrastructure, developer platforms, and site reliability engineering (SRE) capabilities supporting engineering teams and manufacturing facilities worldwide. As TSMC accelerates AI adoption across the enterprise, we are building enterprise-scale AI platforms that enable secure, scalable, and reliable AI services across global regions.
We are seeking a talented and experienced AI Infrastructure Tech Leader to define and build TSMC's enterprise AI Infrastructure Platform.
As the Principal Technical Lead - AI Infrastructure, you will serve as one of the senior technical leaders responsible for defining the architecture and technical strategy of TSMC's next-generation AI platform.
You will design and lead enterprise-scale AI infrastructure that enables thousands of engineers across multiple global regions to securely develop, deploy, and operate AI applications using both internal and external foundation models.
You will collaborate with engineering teams in North America and Taiwan to build a highly scalable AI platform supporting model serving, inference, AI gateways, GPU infrastructure, observability, governance, and platform automation.
This role combines deep expertise in distributed systems, cloud infrastructure, AI platforms, and software architecture.
Key Responsibilities:- Define AI Platform Architecture.Lead the architecture and technical strategy for enterprise AI infrastructure. Drive long-term technical direction for: AI Platform, Enterprise AI Gateway, Multi-model inference platform, GPU infrastructure, AI governance, AI observability.
- Build Enterprise AI Infrastructure:Design highly scalable platforms supporting LLM serving, GPU scheduling, Model routing, Model lifecycle management, RAG infrastructure, Vector databases, AI orchestration.
- Design Distributed Systems:Lead architecture for high availability AI services, multi-region deployment, disaster recovery, service mesh, distributed caching, event-driven architecture, global load balancing.
- AI Platform Engineering:Drive engineering best practices for:
Kubernetes, Platform-as-a-Service, AI deployment automation. - Technical Leadership:Define technical strategy across multiple engineering teams. Lead architecture reviews. Mentor senior engineers. Driving engineering excellence. Influence technical decisions across North America and Taiwan. Partner with product managers, architects, and engineering leaders to define long-term roadmaps.
- BS/MS/PhD in Computer Science or related field.
- 12+ years of software engineering experience.
- 5+ years leading architecture for distributed systems or cloud infrastructure.
- 3+ years hands‑on experience designing enterprise AI platforms.
- 3+ years’ experience as a Principal Engineer, Staff Engineer, Distinguished Engineer, or Technical Lead.
- Strong technical communication skills for collaborating with global cross‑functional teams.
Experience:
AI Infrastructure, strong experience with:
- LLM inference
- Enterprise AI Gateway
- AI serving infrastructure
- GPU scheduling
- Multi-model AI architecture
- Model lifecycle management
- NVIDIA GPU ecosystem
- vLLM
- TensorRT-LLM
- Triton Inference Server
- Ray
- Kubeflow
- MLflow
- Kubernetes
- Docker
- Helm
- ArgoCD
- Service Mesh
- AWS or Azure or GCP
- Strong programming…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).