×
Register Here to Apply for Jobs or Post Jobs. X

AI Inference Engineer

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: F5
Full Time position
Listed on 2026-09-02
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, Backend Developer, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.

Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.

Job Description

The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy.

This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.

Key Responsibilities High-Performance AI Serving
  • Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.
  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.
Hardware Acceleration and Optimization
  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.
  • Collaborate with hardware teams to maximize utilization and performance across various computational environments.
Inference Orchestration and Scalability
  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.
  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.
Performance Monitoring and Observability
  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).
  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.
Technical Requirements

Required Skills:
  • Programming

    Languages:

    Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Inference Tools:
    Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Infrastructure Expertise:
    Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary