×
Register Here to Apply for Jobs or Post Jobs. X

Infrastructure Engineer

Job in Austin, Travis County, Texas, 78716, USA
Listing for: Pursuit Talent Advisory
Full Time position
Listed on 2026-09-09
Job specializations:
  • IT/Tech
    IT Infrastructure, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

We recently partnered with an early growth-stage company that builds AI systems for industrial environments. The platform runs on customer-owned hardware, on-site, without a cloud dependency  takes in live data from plant equipment and business systems, turns it into a structured view of the operation, and puts an AI layer on top of it for the people running the floor.

The team is small, flat, and AI-native. ICs own their work end to end: scope, build, ship, verify. Autonomy is the default.

THE ROLE

You own how the platform gets deployed, runs, and stays up, both on constrained on-premise hardware with unreliable connectivity and in the cloud environments used to develop and validate it. This is not a "keep the CI green" job: the deployment target is physical hardware at a customer site, and getting a release onto it repeatably is a genuine engineering problem.

You'll be one of the first few engineers on the team. Expect to write application code too.

WHY THIS IS INTERESTING

Most infrastructure roles are cloud roles. This one isn't. The software has to run on a box someone can walk up to and unplug, in a building with a spotty VPN, next to machines that cost more than the company. That constraint makes almost every decision (deployment, secrets, observability, rollback) more interesting than the cloud version of the same problem.

WHAT YOU'LL DO
  • Own the on-premise deployment path. Lightweight Kubernetes on single-node and small multi-node topologies, templated and layered manifests, staged bring-up of a full stack from bare hardware, and clean recovery after hard power loss.
  • Own Git Ops. Declarative reconciliation from a Git source of truth, CI-driven image pinning, and revert-based rollback, with a clear, enforced boundary around what the reconciler is and isn't allowed to manage.
  • Own the cloud dev and staging estate. Infrastructure as code for compute, managed databases, container registry, and storage. Enforce the promotion flow from branch to PR to static checks to a real staging environment before anything reaches a shared environment.
  • Own CI/CD. Multi-service image builds, manifest generation, migration gating, and episodic environments that come up and tear down without a human babysitting them.
  • Own the GPU inference layer. The company runs its own models on customer hardware rather than calling a provider API. You'll own that: driver and device-plugin prep, model and quantization choices against whatever hardware you're given, and serving configuration (batching, KV cache, context length, parallelism, memory utilization) tuned so a fixed box meets latency targets. Customers buy one appliance, so utilization is a hard budget, not a cost-optimization exercise.
  • Make the appliance operable. Secrets that don't live in Git, telemetry and model-call tracing a support engineer can actually read, and runbooks for the failure modes that will happen at 2 am on a factory floor.
  • Harden it. Industrial networks are a real trust boundary, and some components need elevated access to do their job. Keep the isolation between privileged components, application services, and the AI runtime intact as the system grows.
  • Reduce toil. If a deployment step needs a person, automate it or delegate it to an agent.
WHAT WE'RE LOOKING FOR

Required:
  • Strong Kubernetes fundamentals, not just kubectl apply, but Stateful Sets, storage, networking, ingress, and debugging a cluster that's misbehaving. Bare-metal or single-node experience (k3s, RKE2, MicroK8s) counts more here than managed EKS/AKS.
  • Infrastructure as code in production (Terraform or equivalent), with real state management and multi-environment discipline.
  • Docker beyond the basics: multi-service builds, image size and layer hygiene, registry workflows.
  • CI/CD ownership: you've built and maintained pipelines, not just consumed them.
  • Hosting AI workloads on GPUs. You've run model inference on your own hardware: an inference server (vLLM, TGI, TensorRT-LLM, or similar) behind a real workload, on GPUs you were responsible for. You can reason about VRAM budgets, batching and concurrency, KV cache, quantization tradeoffs, and where throughput vs. latency actually breaks. You know how…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary