×
Register Here to Apply for Jobs or Post Jobs. X

Director of Production Engineering

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-08-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, AWS
Salary/Wage Range or Industry Benchmark: 220000 - 265000 USD Yearly USD 220000.00 265000.00 YEAR
Job Description & How to Apply Below

Director of Production Engineering Remote, United States About this Position

Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, Dev Ops & SRE to lead the teams responsible for the availability, scalability, and security of our production environment.

Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance. This is a hands‑on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross‑team execution.

Responsibilities
  • Hire and build a globally‑distributed Dev Ops/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.
  • Own the reliability and infrastructure roadmap for our AWS‑based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.
  • Lead the organization's security operations (Sec Ops) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.
  • Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.
  • Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean‑time‑to‑resolution.
  • Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and Dev Ops processes (e.g., automated investigation, remediation, and PR‑generation pipelines).
  • Drive Infrastructure‑as‑Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.
  • Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.
  • Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit‑readiness.
  • Lead and participate in the Incident Management on‑call rotation, working with SRE and development teams to meet and exceed availability goals.
  • Stay current on cloud, Dev Ops, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.
Required Qualifications
  • 8-12 years of experience in Dev Ops, Site Reliability Engineering, or production infrastructure roles, including people management experience.
  • Deep hands‑on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).
  • Demonstrated experience running security operations (Sec Ops) — vulnerability management, incident response, and remediation of production security issues.
  • 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.
  • Strong experience with Infrastructure‑as‑Code (e.g., Terraform, Cloud Formation) and CI/CD automation.
  • Proficiency in at least one of Go, Python, or Bash, with day‑to‑day use of Git and test automation pipelines.
  • Hands‑on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).
  • Proven track record partnering cross‑functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.
  • Demonstrated experience leading incident management and on‑call practices for high‑availability production systems.
  • Bachelor's degree in Computer Science, Engineering, or related field required;
    Master's degree preferred.
  • Experience building or scaling automated investigation and remediation pipelines for production error…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary