Senior DevOps Engineer
Listed on 2026-09-04
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AWS
- We are looking for a Sr. Dev Ops Engineer II to help us build and scale the infrastructure that powers both our core platform and our rapidly growing agentic AI services
- You will be at the intersection of cloud infrastructure, AI operations, and platform engineering building the foundation that enables Hi Marley tooperatereliably at enterprise scale while deploying autonomous AI agents in regulated insurance workflows
- You’llalso be expected to raise the bar for the teams around you setting infrastructure standards, driving technical decisions in ambiguous situations, and helping less experienced engineers grow their operational instincts
- Design andoperatecloud infrastructure on AWS that supports both our core SaaS platform and our agentic AI services, ensuring reliability, scalability, and cost efficiency
- Build andmaintainAI/ML infrastructure and monitoring for LLM-powered agentic services
- Establish and enforce infrastructure-as-code standards using Terraform, defining the patterns other engineers follow for environment parity, drift detection, and automated compliance validation
- Implement observability beyond availability data integrity monitoring, SLO frameworks with error budgets, and automated regression detection for both platform and AI services
- Build deployment automation including pre-deployment verification, migration script validation, and codified rollback procedures toeliminatehuman-memory dependencies
- Support big data infrastructure: data pipelines, warehousing (Redshift), and analytics tooling that enables reporting, BI, and AI training workflows
- Implement security and compliance controls for AI workloadsoperatingin regulated carrier environments including audit logging, access governance, and configuration management
- Drive environment parity across all infrastructure with automated drift detection and remediation
- Improvedisaster recovery capabilities: documented and rehearsed DR procedures, defined RTO/RPO by service tier, and tested recovery runbooks
- Lead architecture reviews for new services, integrations, and AI agent deployments partnering with engineering, product, and security to ensure infrastructure decisions are sound before they ship
- Innovate on developer experience: reduce friction in testing environments, CI/CD pipelines, and local development workflows
- Act as a technical anchor for infrastructure decisions across teams providing clarity when requirements are ambiguous and helping the organization converge on consistent, scalable approaches
- A fun, lively startup culture
- Core values-based leadership
- A culture of employee engagement, diversity and inclusion
- Open vacation policy - we all work hard and take time for ourselves when we need it
- Ample opportunities to learn and take on new responsibilities in a fast-paced, growth-mode startup
- Full benefits package including parental leave, a matching 401k program, and medical, dental, vision, disability, and life insurance
- Generous stock options - we all get to own a piece of what we're building
You have strong infrastructure-as-code skills with Terraform and understand how to manage state, modules, and multi-environment configurations
Track recordof leading cross-team technical initiatives and mentoring engineers on infrastructure and operational best practices
You naturally step up to lead technical conversations, and people across teams seek you out when infrastructure decisions get complicated
You have experience with compliance-sensitive environments and understand why audit trails, access governance, and change management matter
You are comfortable operating in a fast-moving environment where AI capabilities are evolvingrapidlyand infrastructure decisions have regulatory implications
You communicate well with both engineering and non-technical stakeholders
Bachelor's degree in Computer Science, Engineering, or equivalent experience
You understand data infrastructure: pipelines, warehousing, ETL/ELT, and how to support analytics at scale
You have built and operated infrastructure fortraditional andAIorML workloads at a SaaS company
You have deep experience with AWS cloud services (ECS, Lambda, Sage Maker, Bedrock, S3, DynamoDB, Redshift,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).