×
Register Here to Apply for Jobs or Post Jobs. X

Embedded ML Engineer, Edge AI

Remote / Online - Candidates ideally in
Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: SimpliSafe
Remote/Work from Home position
Listed on 2026-07-20
Job specializations:
  • Software Development
    Embedded Systems/ Firmware/ IoT, Unix/Linux, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 185500 - 244600 USD Yearly USD 185500.00 244600.00 YEAR
Job Description & How to Apply Below
Position: Staff Embedded ML Engineer, Edge AI

Overview

We’re a high-tech home security company that’s passionate about protecting the life you’ve built and our mission of keeping Every Home Secure. We foster a collaborative, growth‑oriented culture with opportunities to make a meaningful impact. We embrace a hybrid work model with core in‑office days and flexible remote work for the remainder of the week.

About the Role

We are seeking a highly motivated and experienced Embedded Machine Learning Engineer to join our Edge AI team. As a key contributor, you will lead on‑device inference and performance optimization of ML models powering outdoor monitoring in the home security space. The role focuses on making models fast, power‑efficient, stable, and shippable on real embedded hardware (outdoor cameras and doorbells). You will operate across the stack from model runtime integration down to kernel/operator optimization, memory movement, scheduling, and accelerator utilization to deliver reliable real‑time behavior under tight compute, memory, bandwidth, and thermal constraints across device tiers.

Responsibilities
  • Own the embedded deployment and performance of on‑device ML inference for outdoor monitoring workloads (real‑time video/event pipelines).
  • Optimize end‑to‑end inference performance across CPU/DSP/NPU/GPU (as applicable): latency, throughput (FPS), memory footprint, power, thermals, startup time, and stability.
  • Perform kernel/operator‑level optimization: vectorization (e.g., SIMD/NEON), tiling, cache‑friendly memory layouts, reducing bandwidth and memory copies, optimizing post‑processing, fusing ops, minimizing synchronization/overhead, and thread scheduling.
  • Integrate and maintain ML models within embedded pipelines: model import/export validation, operator compatibility, graph transforms, robust error handling, watchdogs, and safe fallback behavior.
  • Drive quantization and deployment readiness from an embedded perspective: validate INT8/FP16 paths, calibration flows, numerical accuracy checks, and debug quantization edge cases and operator mismatches on target runtimes.
  • Build tooling for profiling, benchmarking, and regression tracking on devices: per‑layer timing, memory tracking, thermal/perf tests, CI gating, and automated performance regression gating across device tiers and firmware versions.
  • Partner closely with ML engineers to translate model changes into deployment impact; provide constraints and design guidance that improve deployability and performance.
  • Provide staff‑level leadership: set performance standards, lead technical reviews, mentor engineers, and influence platform roadmap for on‑device ML.
Qualifications
  • 8+ years of experience in embedded systems and/or performance engineering, with experience shipping production software on constrained devices.
  • Strong C/C++ expertise with deep knowledge of low‑level performance topics: CPU architecture, memory hierarchy, concurrency, and real‑time considerations.
  • Demonstrated experience optimizing ML inference on embedded targets, including operator/kernel tuning and end‑to‑end pipeline optimization.
  • Familiarity with modern vision model families (transformer‑based detectors and CNN‑based detectors) and their execution characteristics.
  • Experience with on‑device inference runtimes and deployment workflows (e.g., TFLite, ONNX Runtime, TensorRT or vendor runtimes), including operator support constraints and graph‑level transformations.
  • Strong debugging and profiling skills (perf, flame graphs, hardware counters, tracing) and ability to drive performance investigations to closure.
  • Ability to lead cross‑functional efforts across ML, firmware, and hardware teams; comfortable defining benchmarks/KPIs and making tradeoffs.
Bonus Points
  • Experience with embedded accelerators and vendor tool chains (DSP/NPU compilers, delegates, GPU compute, custom runtimes).
  • SIMD expertise (ARM NEON/SVE), hand‑tuned kernels, or familiarity with libraries like XNNPACK/QNNPACK/oneDNN/CMSIS‑NN.
  • Experience with quantized inference (INT8) at scale: calibration strategies, numerical debugging, and accuracy‑performance tradeoffs.
  • Experience with camera/doorbell pipelines: ISP/video decode/encode, DMA/zero‑copy buffers,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary