×
Register Here to Apply for Jobs or Post Jobs. X

AI Engineer, Model Training, Inference & Infra

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: Agentrys
Full Time, Apprenticeship/Internship position
Listed on 2026-09-04
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 170000 - 210000 USD Yearly USD 170000.00 210000.00 YEAR
Job Description & How to Apply Below
About Agentrys

Agentrys is building the next generation of design automation for the semiconductor industry.

Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. Agentrys Studio combines AI agents, engineering knowledge, agent-native tools, advanced models, and continuous learning to automate complex chip-design workflows.

Our team brings deep experience in artificial intelligence, electronic design automation, semiconductor design, GPU-accelerated computing, and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity, design quality, and time to market.

The Role

We are looking for an exceptional AI Engineer to own the model training, inference, and infrastructure that power Agentrys' agentic design workforce.

You will drive the full model lifecycle: data pipelines, pretraining and post-training, reinforcement learning, evaluation, and high-performance serving. You will operate the GPU infrastructure that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems, model training, and the self-improving, self-evolving learning loops that let our models and agents get better over time from real execution feedback.

Your work directly determines how capable, fast, and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills, and who wants the models they train and serve deployed in real semiconductor design environments—not left in notebooks or benchmarks.

What You'll Do
  • Train, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning, RLHF/RLAIF, reinforcement learning, and distillation.
  • Build scalable data pipelines for pretraining, post-training, and evaluation, including sparse, private, and domain-specific engineering data.
  • Design and operate distributed training on multi-node GPU clusters, using data, tensor, pipeline, and sequence parallelism (for example FSDP, Deep Speed, or Megatron-style approaches).
  • Build high-throughput, low-latency inference systems with continuous batching, KV-cache management, paged attention, quantization, and speculative decoding.
  • Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance across CPUs and GPUs.
  • Build core model infrastructure: cluster orchestration, job scheduling, checkpointing, fault tolerance, reproducibility, observability, and cost and utilization tracking.
  • Build automated evaluation systems, benchmarks, and reward models that measure agent capability, reliability, and regression across complex engineering tasks, including problems where design data is private or customer-specific.
  • Design self-improving and self-evolving algorithms and learning loops, where models and agents learn from execution feedback, outcomes, and new data to improve continuously over time.
  • Integrate models with agent runtimes, tool use, retrieval, and the production serving stack.
  • Improve reliability, throughput, and cost efficiency across the training and inference platform.
  • Translate promising research ideas into reliable, scalable product capabilities.
  • Collaborate with research, product, platform, and solutions teams across San Jose, Austin, and Taiwan.
  • Contribute to patents, publications, technical presentations, and the broader development of Agentic Design Automation.
What We're Looking For
  • PhD or master's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
  • Strong programming skills in Python and proficiency in at least one systems language such as C++ or Rust.
  • Deep experience with machine learning frameworks such as PyTorch or JAX.
  • Hands-on experience with one or more of the following:
    • Large-scale or distributed model training
    • High-performance model inference and serving
    • GPU programming and performance optimization
    • ML infrastructure and platform engineering
    • Automated evaluation, reward…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary