Senior DevOps Engineer - Arena
Listed on 2026-09-15
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Senior Dev Ops Engineer
Location: San Mateo, CA (Hybrid) 3 days a week.
Department: Engineering / Dev Ops
Reports To: Zhaofeng Wei
About the RoleWe are seeking a Senior Dev Ops Engineer to help design, build, and operate the cloud infrastructure, CI/CD pipelines, and automation capabilities that power our SaaS platform. This role is ideal for an experienced engineer who enjoys solving complex infrastructure challenges, improving developer productivity, and ensuring platform reliability at scale.
You will work closely with software engineers, cloud operations teams, and security stakeholders to build resilient systems, streamline deployment processes, and maintain high availability across production and non-production environments.
What You'll Do Cloud Infrastructure & Automation- Design, deploy, and maintain cloud infrastructure primarily within AWS.
- Develop and manage Infrastructure as Code (IaC) solutions using Terraform.
- Automate repetitive operational tasks through scripting and tooling.
- Continuously improve platform scalability, reliability, and performance.
- Build, maintain, and optimize CI/CD pipelines using Jenkins, Git Hub Actions, and related technologies.
- Support software development teams by improving build systems, deployment processes, and release automation.
- Troubleshoot and resolve pipeline failures, deployment issues, and infrastructure bottlenecks.
- Partner with engineering teams to improve software delivery efficiency and release quality.
- Monitor production and non-production environments to ensure platform health and stability.
- Implement and maintain observability solutions, including logging, metrics, tracing, and alerting.
- Participate in incident response, root cause analysis, and post-incident reviews.
- Drive operational excellence through automation and continuous improvement initiatives.
- Work closely with globally distributed engineering teams across multiple time zones.
- Share knowledge and mentor junior engineers on Dev Ops best practices.
- Collaborate with security, compliance, and platform teams to maintain secure and reliable systems.
- Ability to commute to the office in San Mateo 3 days a week.
- 6+ years of experience in Dev Ops, Site Reliability Engineering (SRE), Cloud Engineering, or related disciplines.
- Proven experience supporting production cloud environments at scale.
- Experience working within Agile software development organizations.
- Strong experience with AWS cloud services.
- Hands‑on expertise with Terraform and Infrastructure as Code practices.
- Experience managing and optimizing CI/CD pipelines using Jenkins or similar tooling.
- Strong Linux systems administration knowledge.
- Proficiency in scripting using Python, Bash, or similar languages.
- Experience with containerization technologies such as Docker.
- Experience operating Kubernetes environments.
- Knowledge of monitoring and observability platforms such as Datadog, Prometheus, Grafana, or ELK.
- Strong troubleshooting skills and ability to resolve complex infrastructure issues.
- Excellent verbal and written communication skills.
- Ability to work independently while collaborating effectively across distributed teams.
- Experience with Git Hub Actions.
- Experience with service mesh technologies such as Istio or Linkerd.
- AWS, Kubernetes, or cloud-related certifications.
- Experience supporting compliance-sensitive environments.
- Background in software engineering or Site Reliability Engineering.
You’ll join a highly collaborative Dev Ops team responsible for the infrastructure and automation that supports a mission‑critical SaaS platform. This role offers the…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).