×
Register Here to Apply for Jobs or Post Jobs. X

AI Engineers

Job in Mountain View, Uinta County, Wyoming, 82939, USA
Listing for: Accellor
Full Time position
Listed on 2026-07-21
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below
Location: Mountain View

Role Summary

We are looking for an AI Engineer with strong problem‑solving ability, solid Python engineering skills, and hands‑on understanding of machine learning, deep learning, GPUs, and CUDA fundamentals. The ideal candidate may be early in their career but must be technically sharp, curious, implementation‑driven, and capable of learning complex AI systems quickly.

This role is suited for someone who can build models, debug training issues, optimize GPU workloads, understand tensor operations, and work closely with research, platform, and engineering teams.

Key Responsibilities
  • Design, build, train, and evaluate AI/ML models using Python, Tensor Flow, PyTorch, or JAX
    .
  • Develop clean training pipelines with data loading, checkpointing, logging, validation, and experiment tracking.
  • Work on deep learning models including CNNs, Transformers, LLMs, embeddings, and attention‑based architectures
    .
  • Debug model issues such as poor convergence, overfitting, unstable loss, NaNs, tensor‑shape errors, and GPU memory failures.
  • Run and optimize models on GPUs with awareness of CUDA execution, memory usage, batching, mixed precision, and kernel performance
    .
  • Profile training and inference workloads to identify bottlenecks in compute, memory, data loading, and communication.
  • Support development of custom operators, fused kernels, CUDA/Triton‑based optimizations, or framework‑level performance improvements.
  • Build AI inference pipelines with focus on latency, throughput, reliability, cost, and quality.
  • Create evaluation pipelines for model accuracy, robustness, hallucination risk, safety, latency, and regression testing.
  • Read research papers, implement ideas, run experiments, and clearly document findings, trade‑offs, and limitations.
Qualifications
  • Strong programming experience in Python.
  • Good understanding of data structures, algorithms, numerical programming, and clean code practices.
  • Hands‑on experience with Tensor Flow, PyTorch, or JAX.
  • Experience training and debugging models on CUDA‑enabled GPUs
    .
  • Comfortable with Linux, Git, Docker, shell scripting, and experiment management.
  • Ability to reason from first principles and solve ambiguous technical problems.
Fundamentals
  • Deep learning
  • Back propagation
  • Gradient descent
  • Optimizers
  • Loss functions
  • Regularization
  • CNNs
  • Transformers
  • Attention mechanisms
  • Embeddings
  • Model evaluation
GPU & CUDA Knowledge
  • GPU memory
  • CUDA kernels
  • Threads, blocks, and warps
  • Memory bandwidth
  • Kernel launch overhead
  • Mixed precision training
  • CUDA out‑of‑memory debugging
Preferred Skills
  • CUDA C/C++ or Triton kernel development.
  • C++ exposure for performance‑critical AI systems.
  • Distributed training using PyTorch DDP, Tensor Flow Distributed Strategy, JAX, NCCL, MPI, or similar tools.
  • Experience with GPU profiling tools such as:
    • NVIDIA Nsight Systems
    • NVIDIA Nsight Compute
    • PyTorch Profiler
    • Tensor Flow Profiler
  • Experience with inference optimization, quantization, TensorRT, ONNX, XLA, or model compilation.
  • Exposure to LLMs, RAG systems, agents, embeddings, multimodal AI, or AI evaluation frameworks.
  • Understanding of FP32, FP16, BF16, INT8, tensor cores, and numerical stability.
What Excellent Looks Like
  • Can implement a model from a research paper without blindly copying code.
  • Can debug why a model is not learning.
  • Can explain why GPU utilization is low.
  • Can identify whether a workload is compute‑bound, memory‑bound, or communication‑bound.
  • Can write clean, reproducible, and testable Python code.
  • Can profile before optimizing.
  • Can work across model logic, GPU execution, data pipelines, and inference systems.
  • Can communicate technical findings clearly and precisely.
Ideal Candidate Profile
  • The ideal candidate is an early‑career AI engineer who is not limited to using ML libraries as black boxes.
  • They understand how models train, how tensors move, how GPUs execute workloads, and how performance, quality, reliability, and safety come together in real AI systems.
  • The strongest candidates will show hands‑on projects in model implementation, GPU acceleration, CUDA/Triton operators, distributed training, LLM evaluation, inference optimization, or research paper reproduction.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary