More jobs:
Manager, Infrastructure Engineering
Job in
Austin, Travis County, Texas, 78716, USA
Listed on 2026-07-20
Listing for:
Jobtailor
Full Time
position Listed on 2026-07-20
Job specializations:
-
IT/Tech
SRE/Site Reliability
Job Description & How to Apply Below
Responsibilities
- Lead with platform and company outcomes over local optimization. You lean in wherever critical infrastructure risk or opportunity exists, regardless of org boundaries, and prioritize the success of the broader platform and product ecosystem.
- Create a safe, collaborative team environment. You name problems, invite open discussion, and help your team learn from incidents and change without blame.
- Raise the bar. You identify and influence improvements in reliability, performance, security, and engineering standards across services and environments.
- Provide clarity and empower autonomy. You ensure expectations, SLAs, and guardrails are clear so teams can move quickly and independently on a stable platform.
- Ensure our AWS environments are secure, cost‑efficient, and production‑ready, with strong foundations in IAM, networking, secrets management, encryption, and compliance‑aware controls.
- Lead the evolution of our CI/CD, observability, and platform tooling, making it easier and safer for product and data teams to ship, monitor, and operate services independently, and improving overall developer experience.
- Drive operational excellence across reliability, scalability, and performance, including capacity planning, incident management, on‑call health, and post‑incident learning culture.
- Define and track key platform health and cost metrics, such as availability, latency, error budgets, infrastructure cost per transaction, and change failure rates, using these insights to guide prioritization and continuous improvement.
- Partner closely with Security, Data Platform, and Product Engineering teams to ensure infrastructure decisions enable reliable experimentation, user‑centric product development, and scalable data and ML systems.
- Oversee large‑scale programs that span multiple teams, such as cloud modernization, multi‑region or cell‑based architecture, and foundational security or resilience upgrades.
- Represent Infrastructure in cross‑functional planning and technical forums, translating platform risks and opportunities into clear business tradeoffs and decisions.
- Have 3+ years of experience managing infrastructure, SRE, or platform engineering teams, and at least 3+ years of hands‑on experience building and operating production systems in the cloud.
- Are comfortable designing and reviewing solutions in AWS, including networking, container orchestration, storage, and IAM, and can make pragmatic tradeoffs between performance, reliability, and cost.
- Are eager to integrate generative AI into engineering workflows to improve delivery speed, quality, and developer experience (e.g., runbooks, incident analysis, platform automation).
- Have experience with modern Dev Ops practices and tooling (e.g., Terraform or other IaC, CI/CD systems, observability stacks, incident management) and know when to invest in automation and standardization to reduce toil.
- Thrive in environments where you’re asked to elevate systems, and bring clarity to ambiguity, especially in high‑growth, high‑change conditions.
- Communicate clearly across technical and non‑technical audiences, and can translate complex infrastructure and reliability tradeoffs into clear roadmaps and investments, advocating for the platform in a way that balances speed, quality, and long‑term business value.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×