More jobs:
Job Description & How to Apply Below
We are seeking a Senior AI Research Scientist with deep expertise in modern foundation models and advanced neural-network architectures. The ideal candidate understands large language models, diffusion and flow-based generative models, state-space models (SSMs), Transformer-SSM hybrids, mixture-of-experts systems, synthetic-data training, and distributed neural-network inference.
This is a hands-on research role for someone who can move comfortably between mathematical theory, rigorous experimentation, prototype implementation, and production collaboration. You will investigate new architectures and training methods while helping build systems that operate efficiently across GPUs, CPUs, NPUs, personal computers, edge devices, private servers, and distributed clusters.
Your work should lead to measurable improvements in model quality, reasoning, latency, throughput, memory use, energy efficiency, privacy, reliability, and total cost of operation. You will have significant influence over product's long-term technical direction and research roadmap.
Key Research Areas
The role will contribute across several of the following areas, with deep specialization expected in at least two:
- Large language models and multimodal foundation models
- Transformer alternatives and attention-efficient architectures
- State-space models, including selective SSMs and Mamba-style architectures
- Hybrid Transformer-SSM, recurrent, sparse, and modular model designs
- Diffusion, discrete diffusion, flow matching, and multimodal generative systems
- Mixture-of-experts models, expert routing, modular networks, and conditional computation
- Synthetic data, model-generated supervision, self-training, and knowledge distillation
- Distributed inference across heterogeneous and intermittently available devices
- Memory-efficient inference, KV-cache management, long-context execution, and model sharding
- Quantization, sparsity, pruning, low-rank adaptation, and dynamic adapter routing
- Agentic models, tool use, planning, reasoning, and multi-agent coordination
- Continual learning, personalization, privacy-preserving learning, and edge AI
Core Responsibilities
Frontier Model Architecture Research
- Design, implement, and evaluate new neural-network architectures for language, reasoning, multimodal generation, and agentic workloads.
- Research the strengths and limitations of Transformers, SSMs, diffusion models, recurrent architectures, mixture-of-experts models, and hybrid designs.
- Develop architectures that combine attention, state-space mechanisms, recurrence, memory, sparse routing, retrieval, and modular components where appropriate.
- Explore efficient long-context methods, external and recurrent memory, adaptive computation, speculative execution, and improved reasoning techniques.
- Investigate diffusion and flow-based approaches for text, image, audio, video, structured data, and multimodal generation.
- Translate promising research papers and mathematical concepts into reproducible prototypes and production-relevant experiments.
Synthetic Data and Model-Generated Training
- Design scalable pipelines for creating high-quality synthetic examples, reasoning traces, preferences, critiques, simulations, and task-specific training data.
- Develop teacher-student, self-training, rejection-sampling, curriculum-learning, process-supervision, and knowledge-distillation approaches.
- Evaluate alignment and post-training methods such as supervised fine-tuning, preference optimization, reinforcement learning, and AI-generated feedback.
- Build filtering, scoring, deduplication, provenance, contamination-detection, and quality-control systems for synthetic datasets.
- Study and reduce the risks of feedback loops, bias amplification, hallucinations, overfitting, reward hacking, and model collapse caused by poorly controlled synthetic data.
- Establish methods for combining synthetic, public, licensed, customer-authorized, and human-generated data while maintaining privacy and traceability.
Distributed Inference and Neural Systems
- Develop algorithms that partition, route, and execute neural-network workloads across heterogeneous devices and infrastructure.
- Research tensor, pipeline, expert, sequence, and context parallelism for inference in resource-constrained and geographically distributed environments.
- Design efficient methods for model sharding, layer placement, distributed KV-cache management, cache-aware scheduling, and dynamic workload migration.
- Improve inference across mixed hardware, including data-center GPUs, consumer GPUs, CPUs, NPUs, Apple Silicon, integrated graphics, edge devices, and browser-based runtimes where appropriate.
- Develop routing and scheduling strategies that account for memory capacity, bandwidth, latency, thermal limits, energy use, device availability, privacy rules, and workload priority.
- Create resilient inference methods that tolerate node loss, unreliable connectivity, changing resource availability, and partial…
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×