More jobs:
Development Engineer, Facility Software Automation
Job in
Austin, Travis County, Texas, 78716, USA
Listed on 2026-07-10
Listing for:
Fluidstack
Full Time
position Listed on 2026-07-10
Job specializations:
-
Software Development
DevOps, Cloud Engineer - Software
Job Description & How to Apply Below
Facility Software Automation Team
Examples of key problems the team is working on
- We do not just watch infrastructure run. We use the signals we collect to build automations that act on it -- compressing the time between a facility coming online and compute being in customers' hands. Detect, decide, act. If we fail to do any one of those three, we lose the trust of the customers who depend on us to keep the frontier moving.
- Bad telemetry is a trust failure at gigawatt scale. Every facility Fluidstack operates runs on the data we collect. A missed signal is a missed alert. A missed alert is a customer running frontier AI workloads on infrastructure that cannot see itself. At the scale we are building, that is not a monitoring gap – it is an outage.
- The platform has to scale as fast as the build does. Fluidstack delivers gigawatts of compute in six months -- a fraction of the industry's 18 to 24 month timeline -- and every new site adds more devices, more signals, and more automation decisions that depend on clean data. The telemetry platform either grows ahead of the business or it becomes the constraint that slows it down.
Scope
- Own the infrastructure underpinning the Facility Software Automation platform: cluster architecture, deployment configuration, and the operational reliability of every service in the stack.
- Own and manage containerization workflows and CI/CD pipelines for telemetry services, making deployments from commit to live in production fast, safe, and repeatable without manual intervention.
- Own the observability layer, Prometheus, alerting, and dashboards, that gives the platform team and Fluidstack operations the signal to catch problems in telemetry pipelines before they hit downstream teams.
- Drive infrastructure-as-code standards across the platform so every environment change is version‑controlled, reviewed, and auditable at the pace of a team shipping continuously.
- Work directly with engineers to take new telemetry services from first commit to production‑ready deployment, setting the bar for what production‑ready means on this platform.
- You have owned container‑based infrastructure in production: designed the architecture, debugged the failure modes that only appear under load, and carried the pager for it.
- You treat telemetry pipeline health and data integrity as non‑negotiable, since the automation that delivers compute to customers depends entirely on the reliability of the data flowing underneath it, and you build and operate accordingly.
- You have hands‑on experience with containerization technologies (Docker, Kubernetes, or equivalent) and have used them to build deployment workflows that engineering teams depend on without thinking about it.
- You have built CI/CD pipelines that engineers actually trust, where a merged PR reaches production safely without anyone watching over it, and you understand modern development workflows well enough to build them from scratch.
- You drive infrastructure‑as‑code as a functioning discipline: every environment change is version‑controlled, reviewed, and auditable, and you hold the platform to that standard.
- You have built observability stacks from scratch (Prometheus, Grafana, Open Telemetry) and you know the difference between a dashboard that looks good and one that actually catches problems before they become incidents.
- You have worked in environments where reliability matters because real operational decisions depend on the data flowing correctly, not just internal tooling.
- Bonus: Git Ops deployment tooling (ArgoCD, Flux, or equivalent). NATS and Click House operational experience, tuning, capacity planning, failure recovery. Compute or critical infrastructure telemetry environments. Experience deploying and operating services at scale across distributed sites.
Compensation
: $250,000 – $300,000 per year, depending on experience, skills, qualifications, and location. Offers equity in the form of stock options.
Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×