Data Center AI Systems Engineer (Competitive Analysis KSA
Listed on 2026-07-31
-
IT/Tech
AI Engineer (Applied/Software), Systems Engineer
Company
Qualcomm Middle East Information Technology Company LLC
Job AreaEngineering Group, Engineering Group >
Systems Engineering
About Us
As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation experiences and drives digital transformation to help create a smarter, connected future for all.
As a Qualcomm Datacenter AI Systems Engineer focused on AI Inference Platforms and Competitive Analysis, you will evaluate, benchmark, optimize, and validate end-to-end AI inference solutions built on Qualcomm AI accelerators. You will analyze real-world customer deployments, assess competitive platforms, and develop fact-based insights on performance, scalability, ecosystem readiness, and operational simplicity across modern GenAI workloads including LLMs, VLMs, RAG, and agentic applications.
AboutThe Role
In this role, you will work across hardware, software, frameworks, and deployment architectures to identify differentiators, gaps, and opportunities that shape Qualcomm's AI platform roadmap and customer engagement strategy. You will collaborate closely with architecture, product, systems, customer engineering, business development, and ecosystem teams to establish Qualcomm AI accelerators as a leading platform for enterprise AI inference deployments.
Minimum Qualifications- Bachelor’s degree in engineering, computer science, or Data Science with 8+ years of experience
- Master's degree in Engineering, Computer Science, Data Science or related field and 4+ year of related work experience.
OR
- PhD in Engineering, Computer Science, or related field and 4+ years of Experience.
- Strong proficiency in Python and common AI/ML frameworks such as PyTorch, Tensor Flow, or ONNX.
- Solid understanding of ML model development, deployment, and AI inference concepts for datacenter workloads.
- Strong foundation in system performance profiling, benchmarking, parallel computing, and workload analysis.
- Working knowledge of Linux-based development, containerized deployment, and cloud-native infrastructure concepts.
- Strong analytical, communication, and cross-functional collaboration skills with the ability to summarize technical findings clearly.
- Master's or PhD degree in Engineering, Computer Science, Information Systems, Electrical Engineering, Physics, or a related technical field with 4+ years of experience.
- Strong understanding of GenAI and enterprise inference workloads, including LLMs, VLMs/LVMs, embeddings, diffusion models, RAG pipelines, agents, and multi-turn application patterns.
- Hands-on experience evaluating AI accelerator platforms and deployment ecosystems, including Qualcomm AI inference accelerators and competitive platforms such as NVIDIA and AMD.
- Experience with production inference frameworks and serving stacks such as vLLM, Triton Inference Server, TensorRT-LLM, SGLang, KServe, Ray Serve, Torch Serve, or equivalent model-serving technologies.
- Deep understanding of inference performance fundamentals, including prefill/decode behavior, KV-cache utilization, batching, scheduling, quantization, memory bandwidth, latency, throughput, TTFT, TPOT, and scaling efficiency.
- Experience designing and executing fact-based competitive benchmarks across single-node and multi-node deployments using representative customer workloads, model configurations, and deployment patterns.
- Knowledge of distributed inference architectures, including tensor parallelism, pipeline parallelism, expert parallelism/MoE serving, disaggregated serving, routing, autoscaling, and cluster-level orchestration.
- Experience with containerized and cloud-native deployment workflows using Docker, Kubernetes, Helm, Git Ops, CI/CD, observability stacks, and infrastructure automation.
- Ability to assess ecosystem readiness, operational simplicity, tooling maturity, model onboarding complexity, reference architectures, observability, supportability, and deployment documentation quality.
- Background in compiler/runtime optimization, graph lowering, operator fusion, kernel optimization, or hardware-aware model optimization for ML workloads is a plus.
- Experience conducting…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).