Senior Manager, Software Engineering; Infrastructure
Job Description & How to Apply Below
What You’ll Be Doing
- Leadership & Team Development
- Lead and grow multiple teams across SRE, Cloud Infrastructure, and MLOps.
- Coach and develop engineering managers and senior individual contributors, fostering a culture of ownership and high craft.
- Build a "Platform-as-a-Product" mindset, ensuring that infrastructure and ML tooling serve as enablers for the rest of the engineering organization.
- Partner with Recruiting to attract and retain specialized talent in the cloud, reliability, and machine learning infrastructure space.
- Reliability & Operational Excellence
- Own the operational health of production systems, including availability, latency, and durability.
- Define and evolve SLIs, SLOs, and error budgets, moving the organization toward data‑driven reliability decisions.
- Lead incident response, driving blameless post‑mortems and systemic improvements to reduce "toil" and improve on‑call sustainability.
- Support ML‑specific reliability, ensuring that model inference pipelines and vector databases meet the same high standards as our core SaaS platform.
- Infrastructure & MLOps Strategy
- Evolve Loopio’s cloud architecture, overseeing capacity planning, disaster recovery, and business continuity.
- Drive the MLOps roadmap, establishing standards for model deployment, monitoring, and scaling (including LLM orchestration and RAG pipelines).
- Lead Cloud Fin Ops, ensuring our infrastructure and AI compute costs are visible, intentional, and optimized.
- Establish standards for infrastructure automation (IaC), configuration management, and secrets handling.
- Security & Cross‑Functional Leadership
- Partner with Security to ensure "secure‑by‑default" infrastructure and robust backup/recovery strategies.
- Communicate risks and trade‑offs clearly to senior leadership, acting as a calm, trusted voice during high‑severity events.
- Collaborate with Product Engineering to support the delivery of high‑impact AI features without sacrificing platform stability.
- 8+ years of experience in infrastructure, SRE, or cloud engineering roles, with 3+ years leading specialized engineering teams.
- Deep Cloud Proficiency:
Extensive experience with AWS (preferred) and modern infrastructure‑as‑code (Terraform). - Operational Grit:
Proven track record of leading teams through production incidents and complex architectural migrations. - MLOps Awareness:
Understanding of the unique infrastructure needs for machine learning, such as GPU orchestration, model serving, or data pipeline stability. - Systems Scaling & Observability:
Proven expertise in managing large‑scale containerized environments and leveraging observability stacks to ensure platform health. - Strategic Communication:
Ability to align technical roadmaps with business objectives and advocate for infrastructure investment. - Experience with Fin Ops or managing significant cloud budgets is a plus.
- Background in supporting AI agentic workflows or autonomous orchestration systems is a plus.
Loopio is an equal opportunity employer that is deeply committed to building equitable workplaces that are diverse and inclusive. We actively encourage candidates from all backgrounds and lifestyles to consider us as a future employer. If you require accommodations at any point during our virtual interview processes, please contact a member of our Talent Experience team at
#J-18808-LjbffrPosition Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×