AI Infra Researcher
Listed on 2026-08-25
-
IT/Tech
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
We are Lenovo. We do what we say. We own what we do. We WOW our customers.
Lenovo is a US $83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services.
Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).
This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit , and read about the latest news via our Story Hub.
- Please Note
* This is a hybrid role in Morrisville, NC. This candidate will be required to work onsite three days a week.
This candidate MUST be a US citizen or US national; US permanent residents or candidates requiring sponsorship cannot be considered.
Position OverviewThe Staff Researcher in AI Compute and Data Infrastructure will conduct applied research and hands‑on development for intelligent, efficient, and resilient Hybrid AI systems. This position works across AI algorithms, computer systems, distributed computing, and data infrastructure to address performance, scalability, reliability, and energy‑efficiency challenges.
The successful candidate will independently own substantial research and development work streams, build production‑quality software, characterize AI workloads, diagnose infrastructure issues, and develop cross‑layer optimization technologies spanning GPUs and other accelerators, CPUs, memory, storage, networking, system software, data pipelines, and AI frameworks.
Key Responsibilities- Research and develop technologies for AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
- Design and implement production‑quality software, system components, services, APIs, diagnostic tools, and scalable data‑processing pipelines.
- Characterize AI training, inference, and data‑processing workloads using profiling, tracing, benchmarking, telemetry, logs, and hardware performance counters.
- Diagnose performance bottlenecks and reliability issues across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, operating systems, runtimes, and AI frameworks.
- Develop hardware/software co‑optimization solutions for GPU utilization, workload scheduling, resource allocation, memory and cache management, communication, data movement, storage access, and model execution.
- Optimize large‑scale data ingestion, preprocessing, transformation, storage, retrieval, and delivery for AI training, inference, and analytics workloads.
- Build intelligent infrastructure diagnostics for anomaly detection, root‑cause analysis, performance regression detection, system health assessment, capacity forecasting, and predictive maintenance.
- Develop fault‑tolerance and resilience mechanisms, including fault detection and isolation, checkpointing, recovery, retry, failover, graceful degradation, and automated remediation.
- Apply machine learning and deep learning to workload modeling, performance prediction, resource optimization, failure prediction, and operational decision‑making.
- Apply time‑series analysis and signal processing to infrastructure telemetry, event detection, change‑point detection, workload forecasting, and system health monitoring.
- Apply causal inference to performance attribution, root‑cause analysis, intervention evaluation, and infrastructure optimization.
- Develop knowledge graphs to model infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
- Optimize systems for throughput, latency, scalability, availability,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).