GCP AI Architect
Listed on 2026-08-24
-
Software Development
Software Architect, AI Engineer (Applied/Software)
Technical Architect
We currently have a career opportunity for a Technical Architect to join our team located in New York.
Job Overview:
We are seeking a distinguished technical leader to architect the next generation of our platform on Google Cloud (GCP). You will not just select services; you will design fault-tolerant, planetary-scale distributed systems that integrate state-of-the-art AI into the core of our engineering culture. You will bridge the gap between "Research" and "Production," applying SRE rigor to AI life cycles.
The Technical Architect provides technology direction, ensures project implementation compliance, and utilizes technology research to innovate, integrate, and manage technology solutions. As a Technology Architect, you will significantly contribute to identifying best-fit architectural solutions for one or more projects; you will collaborate with some of the best talent in the industry to create and implement innovative high quality solutions, participate in Sales and various pursuits focused on our clients' business needs.
This role is considered part of the Business Unit Leadership team and may mentor Junior Architects and/or development team members.
Perficient is always looking for the best and brightest talent and we need you! We're a quickly-growing, global digital consulting leader, and we're transforming the world's largest enterprises and biggest brands. You'll work with the latest technologies, expand your skills, and become a part of our global community of talented, diverse, and knowledgeable colleagues.
Responsibilities- Architecting for "Google-Scale" & Reliability
- Design highly resilient, cloud-native architectures rooted in SRE principles. You will define the strategy for GKE (Kubernetes) multi-cluster orchestration, service mesh (Istio/Anthos) connectivity, and global load balancing. You will enforce strict SLOs and Error Budgets at the architectural level to ensure high availability (99.99%+) while managing the trade-offs between consistency (Spanner) and availability.
Product ionizing Vertex AI & Large Models
- Design highly resilient, cloud-native architectures rooted in SRE principles. You will define the strategy for GKE (Kubernetes) multi-cluster orchestration, service mesh (Istio/Anthos) connectivity, and global load balancing. You will enforce strict SLOs and Error Budgets at the architectural level to ensure high availability (99.99%+) while managing the trade-offs between consistency (Spanner) and availability.
- Orchestrate the end-to-end MLOps lifecycle leveraging the Vertex AI ecosystem.
- You will architect pipelines that move models from notebooks to production seamlessly, managing feature stores (Vertex Feature Store) and model serving endpoints. You will lead the integration of Gemini/PaLM models, defining patterns for RAG (Retrieval-Augmented Generation) using Vector Search and Big Query to deliver low-latency, context-aware AI.
- Implementing "Responsible AI" & Safety Guardrails
- Operationalize AI Safety and Governance. You will architect rigorous evaluation frameworks to detect model drift, hallucination, and bias before deployment. You are responsible for implementing input/output guardrails and ensuring data sovereignty within GCP, treating model governance with the same rigor as IAM security policies.
- Performance Engineering & Compute Strategy
- Define the Compute Strategy for AI workloads, optimizing the mix of TPUs, GPUs, and CPUs. You will own the architectural economics (Fin Ops), utilizing Spot instances, Autoscaling profiles, and custom machine types to maximize throughput-per-dollar. You will dive deep into profiling distributed training jobs to eliminate bottlenecks in data ingestion (Dataflow/Pub Sub).
- Cross-Functional "Design Doc" Culture:
- Drive technical consensus through a rigorous Design Doc (RFC) culture. You will act as the technical authority between Product and Engineering, translating ambiguous business requirements into concrete, engineering-ready architectural blueprints. You will lead "Architecture Review Boards" to ensure every microservice adheres to the broader ecosystem vision.
- Codifying the "Paved Path" (Platform Engineering)
- Build the "Golden Path" for developers. Instead of manual policing, you will architect Infrastructure-as-Code (Terraform) modules and policy-as-code (OPA/kpt) that make doing the right thing the easiest thing. You will define the standard libraries and distinct abstractions that allow feature teams to ship code without reinventing the wheel.
- Bachelor's degree in Computer Science, Computational Linguistics, Mathematics, or equivalent practical engineering experience.
- 5+ years of software engineering experience with fluency in one or more languages (Python, Go, Java, or C++). You must be capable of code reviews and architectural prototyping.
- 2+ years of hands-on experience designing and deploying LLM-backed applications in high-throughput production environments. Experience goes beyond API calls to include fine-tuning, quantization, and context-window optimization.
- Demonstrated mastery of modern agentic frameworks (Lang Chain, Llama Index, AutoGPT) with a portfolio of complex, multi-turn conversational agents that execute deterministic actions (tool use/function calling).
- Solid understanding of distributed systems design, microservices, and API design principles (gRPC/REST). You understand CAP theorem, latency vs. throughput…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).