Manager DevOps Engineer
Listed on 2026-09-02
-
Software Development
DevOps, Cloud Engineer - Software, Backend Developer, Software Architect
About the Role
Our Platform Engineering team builds andoperatesthe core backend systems that power
CoCounsel'sAI agent platform — the infrastructure that lets legal AI agents run reliably own the services that sit between product-facing chat/agent experiences and the AI runtime layer, the CI/CD and deployment infrastructure that ships them safely, and the internal developer tooling that other engineering teams build on. We work closely with product, applied-AI, and platform teams to ship AI-driven capabilities faster and more reliably across the business.
you will do
Own people-management responsibilities for a team of platform engineers — hiring, performance management, career development, and workload/priority balancing — whileremainingtechnically engaged in architecture and code-level decisions.
Set and communicate team-level technical strategy and roadmap in partnership with engineering leadership, translating business priorities into an execution plan.
Guide the design and delivery of backend and platform services that power AI agent workflows — request/event ingestion, agent orchestration,document and file handling — ensuring reliability and delivery speed for product teams building on the platform.
Steer the evolution of the CI/CD and progressive-delivery infrastructure that lets both customer-facing product systems and internal developer tooling ship safely and continuously — including release-decoupling work that separates code deployment from customer-facing exposure.
Ensure the team owns and evolves the cloud infrastructure the platform runs on — provisioning, capacity planning, and environment configuration — using Infrastructure-as-Code practices.
Champion sound backend architecture with a focus on scalability, reliability, and maintainability across a microservices/event-driven system, influencing long-term platform direction.
Drive cross-functional technical initiatives with product, applied-AI, and infrastructure teams, aligning priorities and de-risking complex, multi-system projects from design through production.
Raise the bar on engineering practices — code quality, automated testing, observability, CI/CD, and operational excellence — across the platform stack.
Champion observability across the platform — metrics, logging, and tracing — ensuring your team and downstream teams have the visibility needed tooperatecomplex systems with confidence.
Ensure strong operational posture for the team's services — monitoring, incident response, and root-cause analysis for production systems running on managed cloud AI runtimes (e.g.AWS Bedrock Agent Core ) and Kubernetes.
Ensure the teammaintainsa healthy on-call rotation supporting our customer-facing systems and critical internal tooling, and participate in it yourself — treating every incident as an opportunity to protect the customer experience.
Grow technical leadership within the team, fostering a culture of learning, ownership, and continuous improvement.
6+ years of professional software engineering experience, including prior technical leadership or people-management experience, witha track record designing, building, andoperatinglarge-scale backend systems in production.
Demonstrated experience leading a team directly — 1:1s, performance reviews, hiring, career development — while staying credible at the code/architecture level.
Strongproficiencyin Python (FastAPIor similar) or another backend language, with deep experience in distributed systems, microservices, and cloud-native development.
Hands-onexpertisewith relational databases (PostgreSQL), caching/messaging systems (Redis), API design, and AWS, including container orchestration with Kubernetes (EKS).
Experience with CI/CD pipeline engineering and progressive/safe delivery patterns, and release engineering (decoupling deployment from release exposure).
Infrastructure-as-Code practices for managing cloud infrastructure.
Experience with observability tooling (e.g.Open Telemetry, Datadog) for building and operating production systems.
Experience operating systems built on or integrating with LLM/Generative AI infrastructure is a strong plus.
Demonstrated strength in system…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).