Senior Fleet Software Engineer
Listed on 2026-08-16
-
Software Development
DevOps, Cloud Engineer - Software, Software Engineer, Python
The Role
You'll build the software and infrastructure that lets a small team operate and monitor a growing fleet of deployed robots. The observability, alerting, and automation you build are what the team watches during a shift and what on-call responds to when something goes wrong. You'll contribute to the fleet's operational software end to end, from the telemetry we collect on every robot to the dashboards, pipelines, and tooling that act on it, and work closely with the operations and response teams so the fleet gets easier to run as it scales.
WhatYou'll DoObservability and telemetry
Build and own fleet observability: the metrics, logs, traces, and dashboards that give the team full visibility into live robots
Design telemetry and data pipelines that reliably move robot data to the cloud for monitoring, debugging, and model training
Build fleet-health dashboards and reports that make performance and regressions easy to spot
Build alerting and on-call tooling that catches issues fast and routes them to the right responder with the right context
Automate diagnostics and incident capture so responders can start debugging instead of gathering data
Automate manual, error-prone work and reduce operational toil
Contribute to CI/CD pipelines that deliver code reliably from development to the fleet
Work with software teams to automate over-the-air (OTA) software and firmware updates across the fleet, with staged rollout, monitoring, and rollback
Build provisioning and configuration tooling to keep the fleet consistent and reproducible
3+ years of software engineering, with real ownership of internal tooling, infrastructure, or reliability systems
Strong proficiency in Python plus at least one systems language (Go, C++, or Rust)
Experience building and operating CI/CD pipelines and deployment automation
Experience with a major cloud (AWS or GCP), containers, and orchestration (Docker, Kubernetes)
Solid distributed-systems fundamentals across data transport, monitoring, and fault tolerance
Ability to debug across the full stack, from a device on the network to a service in the cloud
Background in robotics, autonomous vehicles, or other latency- or safety-critical domains
Experience with observability stacks (e.g., Prometheus/Grafana, Open Telemetry, Foxglove)
Experience with OTA or fleet deployment and safe-rollout patterns such as staged rollout and auto-rollback
Familiarity with ROS/ROS2, edge or embedded Linux, or low-latency data transport for real-time systems
Experience building tooling for an operations, on-call, or field team
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).