×
Register Here to Apply for Jobs or Post Jobs. X

AI Infra Staff Researcher

Job in Morrisville, Wake County, North Carolina, 27560, USA
Listing for: Lenovo
Full Time position
Listed on 2026-08-13
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Job Description & How to Apply Below

Staff Researcher In Ai Compute And Data Infrastructure

* Please Note
* This is a hybrid role in Morrisville, NC. This candidate will be required to work onsite three days a week.

* This candidate MUST be a US citizen or US national; US permanent residents or candidates requiring sponsorship cannot be considered.

Position Overview

The Staff Researcher in AI Compute and Data Infrastructure will conduct applied research and hands-on development for intelligent, efficient, and resilient Hybrid AI systems. This position works across AI algorithms, computer systems, distributed computing, and data infrastructure to address performance, scalability, reliability, and energy-efficiency challenges.

The successful candidate will independently own substantial research and development work streams, build production-quality software, characterize AI workloads, diagnose infrastructure issues, and develop cross-layer optimization technologies spanning GPUs and other accelerators, CPUs, memory, storage, networking, system software, data pipelines, and AI frameworks.

Key Responsibilities

  • Research and develop technologies for AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
  • Design and implement production-quality software, system components, services, APIs, diagnostic tools, and scalable data-processing pipelines.
  • Characterize AI training, inference, and data-processing workloads using profiling, tracing, benchmarking, telemetry, logs, and hardware performance counters.
  • Diagnose performance bottlenecks and reliability issues across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, operating systems, runtimes, and AI frameworks.
  • Develop hardware/software co-optimization solutions for GPU utilization, workload scheduling, resource allocation, memory and cache management, communication, data movement, storage access, and model execution.
  • Optimize large-scale data ingestion, preprocessing, transformation, storage, retrieval, and delivery for AI training, inference, and analytics workloads.
  • Build intelligent infrastructure diagnostics for anomaly detection, root-cause analysis, performance regression detection, system health assessment, capacity forecasting, and predictive maintenance.
  • Develop fault-tolerance and resilience mechanisms, including fault detection and isolation, checkpointing, recovery, retry, failover, graceful degradation, and automated remediation.
  • Apply machine learning and deep learning to workload modeling, performance prediction, resource optimization, failure prediction, and operational decision-making.
  • Apply time-series analysis and signal processing to infrastructure telemetry, event detection, change-point detection, workload forecasting, and system health monitoring.
  • Apply causal inference to performance attribution, root-cause analysis, intervention evaluation, and infrastructure optimization.
  • Develop knowledge graphs to model infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
  • Optimize systems for throughput, latency, scalability, availability, resource utilization, energy consumption, and total cost of ownership.
  • Collaborate with hardware, systems, software, architecture, and product teams to transition research technologies into Enterprise AI and Personal AI products.
  • Contribute to patents, invention disclosures, technical publications, internal reports, and reusable software assets.
  • Provide technical guidance and mentorship to junior researchers and engineers.

Minimum Qualifications

  • Master's or PhD degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
  • Four or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
  • Strong programming skills in Python, C++, Java, Go, Rust, Scala, or a comparable language.
  • Demonstrated ability to design, implement, test, debug, profile, and optimize reliable software…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary