Member of Technical Staff - ML Infrastructure & Performance
Listed on 2026-07-01
-
IT/Tech
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Cloud Computing: Infrastructure & Operations, Data Engineering
Moonlake AI
Introducing Moonlake, AI for creating real-time interactive content
MissionImprove throughput, latency, and cost - deploying our models 2–10× faster and cheaper without quality regressions.
Scope of Work- GPU performance: CUDA/Triton kernels, Flash Attention family, paged attention, CUDA Graphs.
- Serving stack:
TensorRT-LLM/Triton Inference Server, vLLM/TGI; continuous batching; on-GPU KV reuse; speculative decoding/medusa; mixture-of-agents routing.
- Parallelism: FSDP/ZeRO, TP/PP/expert parallel; NCCL tuning.
- Quantization/PEFT: AWQ/GPTQ/FP8;
LoRA/DoRA serving.
- Systems:
Ray/k8s/Argo, observability (Prom/Grafana/Open Telemetry), autoscaling, A/B infra, canary + rollback.
Previous experience at infra-heavy startups such as Databricks, Roblox
We are committed to being an on-site, in-person team currently based in San Mateo
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).