×
Register Here to Apply for Jobs or Post Jobs. X

Engineering Manager, DevOps

Job in Boston, Suffolk County, Massachusetts, 02108, USA
Listing for: Venturefizz Product Management Community
Full Time position
Listed on 2026-08-12
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Project Manager, Systems Engineer
Job Description & How to Apply Below

Dev Ops Manager

Maven AGI is an enterprise AI platform founded in July 2023 by executives from Hub Spot, Google, and Stripe. We build conversational AI agents for autonomous customer support  team includes talent from Google, Meta, Amazon, Microsoft, and Stripe, with advisors from OpenAI, Google, Hub Spot, and Stripe.

We're looking for a Dev Ops Manager to lead and evolve the infrastructure powering Maven AGI's AI platform. You will manage and scale a high-performing infrastructure team while helping ensure our systems remain reliable, secure, and scalable across cloud and on-premises environments. This is a technical leadership role that combines people management, operational ownership, and strong infrastructure judgment.

You will partner closely with engineering leaders and technical leads to translate business and customer requirements into clear infrastructure priorities and execution plans. You will also work directly with enterprise customers to understand complex deployment requirements, particularly for private-cloud and on-premises environments, and coordinate stakeholders across the organization to deliver sustainable solutions.

Leadership and Management Responsibilities:

  • Manage, coach, and develop a team of Dev Ops and infrastructure engineers
  • Establish clear expectations, ownership, and accountability across the team
  • Partner with technical leads to align technical strategy, architecture, and execution
  • Hire and onboard engineers as the team grows
  • Lead performance management, career development, and regular feedback
  • Own team planning, prioritization, capacity management, and delivery
  • Balance reliability, security, customer commitments, and long-term platform investments
  • Communicate infrastructure risks, trade-offs, and progress to technical and non-technical stakeholders
  • Build strong partnerships across Engineering, Product, Security, and Customer Success
  • Improve operational processes while avoiding unnecessary overhead and reducing team toil

Technical and Operational Responsibilities:

  • Guide the design, implementation, and operation of cloud and on-premises infrastructure across Azure, AWS, and customer-managed environments
  • Oversee infrastructure-as-code practices using Pulumi, Bicep, Terraform, or similar tools
  • Own the reliability and operation of production Kubernetes environments, including deployments, scaling, monitoring, and incident response
  • Drive the development and improvement of CI/CD pipelines for a large-scale monorepo
  • Establish consistent observability practices across metrics, logs, traces, and alerting
  • Advance reliability practices, including SLOs, capacity planning, disaster recovery, and runbook development
  • Support and scale enterprise AI deployments, including GPU infrastructure, model-serving workloads, and high-concurrency systems
  • Partner with engineering teams to improve developer experience, platform usability, and deployment velocity
  • Strengthen secrets management, access controls, and infrastructure security
  • Evaluate and adopt tools that improve reliability, scalability, and operational efficiency
  • Participate in incident response and ensure incidents lead to durable improvements

Required Qualifications:

  • 7+ years of professional Dev Ops/SRE/Infrastructure experience
  • 3+ years of experience managing teams
  • Deep expertise with Kubernetes in production (AKS, EKS, or GKE)
  • Strong infrastructure-as-code skills (Pulumi, Terraform, or Bicep)
  • Experience operating CI/CD systems (Git Hub Actions, ArgoCD, or Jenkins)
  • Proficiency in at least one scripting/programming language (Python, Go, Type Script, or Bash)
  • Solid understanding of IaaS providers, networking, DNS, load balancing, and TLS
  • Experience with monitoring and observability stacks (Datadog, Prometheus, Grafana, or similar)
  • Experience with multi-cloud or hybrid (cloud + on-prem) deployments
  • Strong communication and cross-team collaboration skills
  • Organized, great attention to detail, comfortable operating in a ticketing environment
  • Thrives in fast-paced startup environments

Nice to have:

  • Experience with GPU infrastructure and ML/LLM serving workloads (vLLM, TEI)
  • Familiarity with Temporal or other workflow orchestration systems
  • Security and compliance…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary