AI Engineer, Applied AI — Smart Vision
Listed on 2026-10-05
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
About Arlo:
At Arlo, we're passionate about creating innovative and reliable solutions that help people protect what matters most to them. Our team is dedicated to delivering products that exceed our customers' expectations, while always pushing the boundaries of what's possible in the world of protection technology. We believe that everyone deserves to feel safe and secure, whether they're at home or away, and we're committed to providing our customers with the peace of mind they need to live their lives without worry.
Arlo’s deep expertise in AI- and CV-powered analytics, cloud services, user experience, product design, and innovative wireless and RF connectivity enables the delivery of a seamless, smart security experience for Arlo users that is easy to set up and interact with every day. Smart Vision is the AI team behind Arlo's intelligence layer: object and person detection, animal/vehicle/package recognition, custom-trained detections, video captioning and scene description, and natural-language search over a user's video library.
Our models run across the edge, the cloud, and third-party foundation models, and they process events from millions of cameras every day.
As a Staff AI Engineer for Applied AI, you'll be a technical owner of the models behind Arlo's smart features — from computer vision detectors running on-camera to vision-language models that describe what happened, to the retrieval and agent layers that let customers ask questions about their video. You'll pick the right approach for each problem (train, fine-tune, prompt, or retrieve), prove it with solid evals, and take it all the way to production at consumer scale.
This is a hands‑on applied role: you ship models, not papers.
- Build, train, and fine‑tune computer vision models for detection, classification, tracking, re‑identification, and video understanding, and improve them against real‑world customer footage — night, weather, motion blur, odd camera angles, edge compute limits.
- Own our video‑understanding pipeline built on vision‑language models: frame selection and temporal context, prompt and output‑schema design, grounding and hallucination control, multi‑event reasoning, and quality tuning for captioning and scene description.
- Adapt models to our domain: SFT, LoRA/QLoRA, preference tuning, distillation into small deployable models, and knowing when a 200M‑parameter specialist beats a frontier model.
- Own the data and evaluation loop — dataset curation, labeling strategy, hard‑negative and failure mining, active learning, benchmark suites, and offline/online metrics that reliably predict customer‑perceived quality.
- Own the embedding and retrieval stack behind video search: multimodal/video embeddings, vector index design and tuning, hybrid search and re‑ranking, and natural‑language queries over a user’s library.
- Build agentic experiences on top of the vision stack: tool/function calling, multi‑step reasoning over event history, RAG and memory, guardrails, and tracing/observability for agent runs.
- Keep production inference fast and economical — serving stack choice and tuning (vLLM/TensorRT-LLM/Triton), quantization, batching, GPU utilization, and routing between hosted foundation models and self‑hosted open models against clear cost and latency targets.
- Partner with product, data, backend, and firmware/edge teams to turn ambiguous product ideas into shipped AI features, with safe rollout (canary, A/B, feature flags) and real production telemetry.
- Raise the engineering bar — architecture and design reviews, MLOps practices, documentation, and mentoring senior and mid‑level engineers.
- BS in Computer Science or a related technical field with 8+ years of experience;…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).