×
Register Here to Apply for Jobs or Post Jobs. X

Manager of DevOps Engineering

Job in Jersey City, Hudson County, New Jersey, 07390, USA
Listing for: Arch Insurance Group Inc.
Full Time position
Listed on 2026-07-11
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 130000 - 223700 USD Yearly USD 130000.00 223700.00 YEAR
Job Description & How to Apply Below
Position: Manager of DevOps Engineering -

Core Responsibilities

With a company culture rooted in collaboration, expertise and innovation, we aim to promote progress and inspire our clients, employees, investors and communities to achieve their greatest potential. Our work is the catalyst that helps others achieve their goals. In short, We Enable Possibility.

The Manager of Dev Ops Engineering is responsible for leading the design, implementation, and operational excellence of enterprise-scale CI/CD systems, infrastructure automation, and engineering. This role requires deep expertise in modern Dev Ops tooling, distributed systems, and cloud-native architectures, while also providing technical leadership to a team of Dev Ops Engineers.

This individual will drive the evolution of our Dev Ops practices, ensuring automation-first delivery pipelines, hardened infrastructure, and highly available services that underpin mission-critical business applications.

Work Arrangement

Hybrid (3 days onsite per week) from our Jersey City, NJ or Raleigh, NC office.

Dev Ops Engineering & CI/CD
  • Own and scale enterprise-wide CI/CD pipelines using modern orchestration tools (e.g., Git Hub Actions, CI, ArgoCD).
  • Architect developer self-service platforms with Infrastructure-as-Code (IaC).
  • Implement role‑based access controls (RBAC) across Kubernetes, cloud IAM, and tool chains to ensure compliance and security.
  • Build extensible automation frameworks enabling teams to provision, deploy, and monitor workloads with minimal friction.
  • Evaluate and integrate next‑generation CI/CD features such as ephemeral environments, policy‑as‑code enforcement, and test environment provisioning on demand.
  • Establish and govern standardization of base container images, Helm charts, and deployment templates to promote consistency and reduce security drift across development teams.
Infrastructure & Cloud Automation
  • Manage cloud‑native infrastructure (Azure and AWS) with a focus on resiliency, scalability, and cost optimization as it pertains to product workloads.
  • Lead adoption of Kubernetes and container orchestration platforms with advanced configuration (e.g., service mesh, Cilium, Calico, OPA/Gatekeeper).
  • Standardize configuration management using Terraform, Terragrunt, or ArgoCD “Helm”, and integrate with CI/CD pipelines for immutable deployments.
  • Optimize cloud spend and resource utilization by implementing advanced autoscaling strategies, rightsizing recommendations, and reserved instance/savings plan management using Fin Ops best practices.
Reliability & Observability
  • Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Agreements (SLAs) for platforms and services.
  • Establish observability practices through metrics, distributed tracing, and logging using tools such as Prometheus, Grafana, ELK/EFK, and Dynatrace.
  • Drive proactive capacity management, chaos testing, and resilience engineering to validate system recovery under failure scenarios.
  • Advance the maturity of AIOps initiatives, leveraging machine learning techniques on telemetry data to predict and preempt potential service degradation.
Security & Compliance
  • Integrate Dev Sec Ops  practices into pipelines (e.g., Sys Dig, Artifactory X‑Ray, dependency scanning, container image hardening).
  • Enforce least‑privilege principles and manage secrets with tools like Hashi Corp Vault, AWS Secrets Manager, Azure Vault, or Kubernetes Secrets.
  • Ensure compliance with regulatory requirements and organization requirements.
  • Manage secrets rotation, key generation, and Public Key Infrastructure (PKI) at scale, ensuring cryptographic best practices are applied across all environments.
Disaster Recovery & Business Continuity
  • Architect and validate multi‑region, multi‑cloud disaster recovery strategies with automated failover testing.
  • Design recovery procedures to minimize RTO/RPO and validate through game‑day exercises.
  • Document and evangelize clear runbooks and incident response plans for all major infrastructure platforms, supporting a 24/7 on‑call rotation.
  • Develop and automate failover testing for all distributed systems, ensuring minimal impact during simulated regional outages.
Leadership & Strategy
  • Lead, mentor, and grow a…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary