Infrastructure Automation Engineer
Los Angeles, Los Angeles County, California, 90079, USA
Listed on 2026-08-03
-
IT/Tech
Systems Engineer, IT Infrastructure, Network Engineer
THUMBWAR is a technical solutions company that combines the talents of our diverse staff, our inventory of equipment, and the experience of over fifteen years in business to produce unique solutions to solve our clients needs.
We work with broadcast networks to design and execute live broadcast workflows, with an emphasis on remote work and hybrid workflows.
We work with content providers to enable workflows that leverage content aggregation to create engaging programming for platforms such as web and mobile during marquee events.
We work with large companies and institutions where we have integrated large-scale storage and archive environments and remain trusted partners after the initial installation for on-going support and maintenance. We do this by providing consulting, on-going support, and turnkey IT infrastructure design and deployment.
Our team, our clients, and our projects may be diverse, but at the core is a cohesive understanding of our customers’ needs and how THUMBWAR’s unique perspective and skillset can address them.
About the role
This is not a corporate help-desk or a static office IT admin role. At THUMBWAR, our hybrid infrastructure kit is our product. We operate in the high-stakes, zero-downtime world of live sports broadcasting and premium media production, where our customers rely on us to deliver critical broadcast workflows for air.
We plan and test rigorously before a kit ever leaves the shop — together, in a shared shop and engineering space where we sketch, argue, and build side by side. Every argument gets heard and thought through before a decision is made, but once we've decided, we move as one unified front. We don't relitigate a call after the fact, even if it wasn't the one you argued for.
The work itself means stitching together disparate networking, broadcast media, and cloud technologies. You won't walk in knowing every one of them cold. What you will bring is a disciplined approach to solving each one, pulling in subject-matter expertise from across THUMBWAR and our partner network, and codifying what you learn into tooling the whole team can reuse.
The work doesn't stop at the shop door. We are involved in some of the largest live sporting events in the world, and that's where we do our best learning — pulling real-time data from the infrastructure itself and real-time feedback from the customers and crews relying on it. Taking feedback well — without getting defensive — is part of the job, not a soft skill on the side.
We won't be perfect every show. What matters is that we take what we hear, organize it into real action items, communicate the plan, and make the next show better than the last.
If you're the person who reaches for a playbook instead of a notebook of tasks, who wants your fingerprints on both the rack and the repo, and who thinks a live broadcast is the best possible stress test for your automation — this is your next step.
Key Responsibilities
- Automate the Full Kit Lifecycle: Build and maintain Infrastructure-as-Code (IaC) pipelines using Terraform and Ansible to take our fly-pack and facility kits from initial build and configuration through deployment, validation, and teardown. Everything we build must be repeatable, versioned, and documented.
- Codify Virtualization & Compute: Manage and automate our hypervisor platforms, with a primary focus on Proxmox. Ensure provisioning, scaling, and performance tuning are driven entirely by pipelines and configuration management rather than manual GUI clicks.
- Build Show-Day Observability: Design and maintain active monitoring and alerting telemetry (network health, hypervisor performance, storage throughput) for live events so that edge cases and anomalies are surfaced and diagnosed before they ever threaten air.
- Physical Hardware Ownership: Rack, cable, stack, and troubleshoot physical server hardware, enterprise switches, and high-performance storage arrays. Automation starts with hardware you understand and respect at the bare-metal level.
- Support Zero-Failure Execution: Support live event schedules on-site or remotely. Use the automation tooling and pipelines you've built to proactively monitor, rapidly…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).