×
Register Here to Apply for Jobs or Post Jobs. X

AI Systems & Platform Internals - Technical Architect

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: Worky
Full Time position
Listed on 2026-08-15
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 260000 - 320000 USD Yearly USD 260000.00 320000.00 YEAR
Job Description & How to Apply Below

Accelloris an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities.

Our mission isto operationalize AI at scale and unlock sustained enterprise value.

Our offerings spanAI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. Byleveragingdesignthinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability.

With a proven track recordof enabling Fortune 100 enterprises and global innovators,Accellorstands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear:to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation.

Technical Architect — AI Systems & Platform Internals

Experience:

10-12 Years
Role Type:
Technical Architect / Staff-Level Systems Architect

Role Summary

Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads.

This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.

The ideal candidate is a senior hands-on architect who can reason across the full AI platform - from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.

Key Responsibilities:
  • 1. AI Systems Architecture

    Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.

    Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.

    Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.

  • 2. Inference Runtime & Model Serving

    Architect high-throughput, low-latency inference systems across large-scale GPU clusters.

    Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.

    Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.

    Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.

  • 3. GPU, Kernel & Distributed Performance

    Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.

    Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.

    Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.

  • 4. Context Engineering

    Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.

    Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.

    Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.

  • 5. Cost Optimization Frameworks

    Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.

    Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.

    Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.

    Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.

  • 6. Training & Research Infrastructure

    Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.

    Contribute to architecture around distributed training, checkpointing,…

  • To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
    (If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary