DevOps Engineer
Listed on 2026-08-22
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Cybersecurity
Overview
The Dev Ops Engineer plays a pivotal role in driving operational efficiency and scalability across our development and IT teams. This role automates and streamlines the software development and deployment processes behind our next-generation tolling and intelligent transportation platform — a cloud-native, event-driven microservices system — ensuring faster and more reliable delivery of applications. By bridging the gap between development and operations, this role helps reduce bottlenecks and increase collaboration, ultimately accelerating the release cycle.
This role directly impacts the company by enhancing system reliability, reducing downtime, and ensuring our infrastructure can scale with growing business demands. This role is crucial for maintaining a robust and agile technology environment, allowing the company to deliver high-quality products to customers more efficiently and remain competitive in the market. This position participates in a 24/7 on-call rotation supporting revenue-critical systems.
ResponsibilitiesKey Responsibilities Infrastructure Management
- Design, deploy, and maintain scalable cloud infrastructure (AWS, Azure, Google Cloud).
- Manage cloud resources to ensure high availability, reliability, and security.
- Operate Kubernetes-based platforms end to end, including stateful workloads such as distributed SQL databases and event-streaming clusters (e.g., NATS, Kafka).
- Optimize cloud infrastructure for cost efficiency and performance.
- Build and maintain CI/CD pipelines, including code-first, containerized pipelines (e.g., Dagger, Git Hub Actions, Git Lab CI).
- Automate build, test, and deployment for a polyglot monorepo (e.g., Rust services and Type Script applications) with reproducible developer tool chains.
- Collaborate with development teams to integrate CI/CD best practices, including ephemeral local clusters (e.g., k3d, kind) that mirror production.
- Implement and manage infrastructure as code (IaC) with tools such as Terraform or Ansible, plus Helm charts and Git Ops workflows (e.g., Argo CD) for Kubernetes workloads.
- Automate provisioning, configuration, and scaling of infrastructure.
- Ensure repeatability and consistency in infrastructure setup and deployment.
- Set up and manage monitoring, logging, tracing, and alerting tools (e.g., Prometheus, Grafana, Open Telemetry, ELK Stack).
- Proactively monitor systems for performance issues and outages.
- Participate in a 24/7 on-call rotation and respond to incidents to minimize downtime for revenue‑critical services.
- Manage containerized applications using Docker and Kubernetes.
- Deploy and manage event-driven microservices and federated GraphQL API routers in a containerized environment.
- Ensure scalability and fault tolerance using orchestration tools.
- Work closely with software developers, system administrators, and QA engineers to troubleshoot and optimize system performance.
- Participate in code reviews and provide feedback on best practices for deployment and operations.
- Identify areas for process improvements and implement solutions to enhance productivity and system efficiency.
- Integrate security best practices into infrastructure design and deployments, including policy-as-code (e.g., OPA) and centralized secrets management (e.g., Vault‑compatible tooling).
- Manage secrets, permissions, and access control within cloud environments.
- Ensure compliance with relevant industry standards (e.g., SOC 2, ISO 27001, PCI DSS).
- Bachelor’s Degree in Computer Science, Information Technology, Software Engineering, or a related field.
- Equivalent practical experience may be considered in lieu of a formal degree.
- 3+ years of experience in a Dev Ops, Site Reliability Engineer (SRE), or similar role, with a strong focus on cloud infrastructure and automation.
- Proven experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Hands‑on experience with CI/CD tools (e.g., Dagger, Git Hub Actions, Git Lab CI, Jenkins) and automation frameworks…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).