AI Software Engineer lll
Listed on 2026-09-05
-
Software Development
Cloud Engineer - Software, AI Engineer (Applied/Software), Backend Developer, DevOps
AI Software Engineer III
At Wyze, we make smart home technology accessible to everyone. We're known for disrupting markets with high-quality, affordable products - from cameras to lighting to sensors and more. We believe technology should simplify life, not complicate it. We're a fast-moving, customer-obsessed team driven by curiosity and powered by data. We are looking for a Software Engineer to help build and scale the infrastructure behind our production AI systems.
You will work on the platform that supports the full lifecycle of our AI models, from training and experimentation to deployment and production inference. This includes building scalable model training infrastructure, production AI services, cloud and compute infrastructure, deployment systems, observability, and developer tooling. This is primarily an infrastructure and systems engineering role rather than a model research role. You do not need to be an expert in LLM inference or GPU optimization when you join.
We are looking for a strong software engineer who can build reliable distributed systems, learn quickly, and solve infrastructure problems across different layers of the stack. You will work closely with AI scientists and other software engineers to make it easier and faster to train, deploy, operate, and iterate on AI models in production. The AI infrastructure landscape is evolving extremely quickly.
We value engineers who can evaluate new technologies pragmatically, move fast, adapt to changing requirements, and continuously improve how we build AI systems.
- Design, build, and operate infrastructure and backend services that power production AI features.
- Build and improve model training infrastructure, including systems that support training jobs, experimentation, compute management, data workflows, and model artifacts.
- Build infrastructure that supports the full model lifecycle, from training and experimentation through deployment and production serving.
- Improve the scalability, performance, reliability, and cost efficiency of our AI platform.
- Build and maintain cloud-based services and containerized workloads for AI and ML applications.
- Develop systems that make it easier for ML engineers to train, evaluate, deploy, and iterate on models.
- Build reusable platform capabilities that allow engineering teams to launch new AI-powered features quickly and safely.
- Improve deployment, rollout, monitoring, and operational workflows for AI models and services.
- Diagnose and resolve reliability and performance issues across application, infrastructure, compute, and ML system layers.
- Support large-scale AI workloads across text, image, video, and multimodal applications.
- Evaluate and adopt emerging AI infrastructure technologies when they provide meaningful improvements in productivity, performance, reliability, or cost.
- Improve software development and operational workflows through effective use of modern AI coding agents and agentic engineering tools.
- Work closely with AI scientists, and product teams to translate rapidly changing requirements into practical technical solutions.
- Own systems end-to-end, from architecture and implementation through deployment, monitoring, operation, and continuous improvement.
- 3+ years of professional software engineering experience building production backend systems, infrastructure, or distributed systems.
- Strong programming skills in Python, Java, Go, or a comparable backend or systems language.
- Strong understanding of distributed systems fundamentals, including concurrency, fault tolerance, messaging, back pressure, load balancing, and horizontal scaling.
- Production experience with Kubernetes or Docker-based containerized workloads.
- Experience operating production services on a major cloud platform such as AWS, GCP, or Azure.
- Experience designing and operating high-throughput or latency-sensitive production systems.
- Familiarity with No
SQL databases such as DynamoDB, Cassandra, or comparable distributed data stores, including common data modeling and scalability considerations. - Strong debugging and problem-solving skills across application, infrastructure, and networking…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).