×
Register Here to Apply for Jobs or Post Jobs. X

DevOps Engineer; Infrastructure & Reliability

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: OpusClip
Full Time position
Listed on 2026-03-09
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing
Salary/Wage Range or Industry Benchmark: 80000 - 100000 USD Yearly USD 80000.00 100000.00 YEAR
Job Description & How to Apply Below
Position: DevOps Engineer (Infrastructure & Reliability)

🎨
Opus Clip is the world’s No.1 AI video agent, built for authenticity on social media.

We envision a world where everyone can authentically share their story through video, with no expertise needed. Within just 18 months of our launch, over 10 million creators and businesses have used Opus Clip to enhance their social presence.

We have raised $50 million in total funding and are fortunate to have some of the most supportive investors, including Soft Bank Vision Fund, DCM Ventures, Millennium New Horizons, Fellows Fund, AI Grant, Jason Lemkin (Saa Str), Samsung Next, GTMfund, Alumni Ventures, and many more.

Check out our latest coverage by Business Insider featuring our product and funding milestones, and our recognition as one of The Information’s 50 Most Promising Startups in 2024.

Headquartered in Palo Alto, we are a team of 100 passionate and experienced AI enthusiasts and video experts, driven by our core values:

  • Be a Champion Team
  • Prioritize Ruthlessly
  • Ship fast, Quality Follows
  • Obsess over customers

Be a part of this exciting journey with us!

The Mission

The Mission we are on is to find a hands‑on founding Dev Ops Engineer to own the reliability and scalability of the Agent Opus platform – a leading agentic video generation product. This person will stabilize our processing clusters, design the next phase of our agentic platform, and serve as the technical bridge between our cloud infrastructure and our 15M users.

You will engineer the infrastructure strategy that underpins our trust and reliability in the market. You will help setup on‑call rotation and own the full incident lifecycle from minimizing Time-to-Detect to tracking post mortem actions ensuring that our high‑velocity growth never compromises our performance.

Key Responsibilities
  • Infrastructure Architecture & Cluster Operations
    • Architect Dedicated Environments: Lead the design and implementation of high‑throughput, isolated processing environments and clusters for compliance needs.
    • Scale Production: Drive general improvements in our Temporal clusters and production Kubernetes environments.
    • Technical Execution: Be hands‑on with the stack to optimize resource allocation, reduce latency, and enforce isolation strategies for critical accounts.
  • Monitoring, Alerting & Detectability
    • Beat the Customer to the Alert: Overhaul our Datadog observability suite to aggressively reduce Time-to-Detect (TTD). You ensure we identify latency spikes and stalled projects before users do.
    • Threshold Tuning: tune alert thresholds to eliminate noise and focus on “symptom‑based” alerts that reflect the actual user experience.
    • External SLO Ownership: Define and report on Service Level Objectives (SLOs), acting as the internal guarantor that we are meeting the targets we sold.
  • Incident Command & “Extreme Ownership”
    • First Responder & Mitigation: Serve as the first line of defense during outages. You will own immediate mitigation, including cluster debugging and manual scaling intervention if Horizontal Pod Autoscalers (HPA) fail or lag.
    • Drive Recovery Metrics: You are accountable for shortening Time-to-Mitigation (TTM) and Time-to-Recover (TTR). Your priority is to stop the bleeding first, then fix the wound.
    • Root Cause Analysis: Lead the post‑mortem process to determine Time-to-Root Cause and implement systemic fixes. You will translate these technical findings into clear updates for Customer Experience (CX) and Leadership.
    • Accountability: Work collaboratively with Engineering Owners to track improvements against the reliability roadmap. You are responsible for flagging risks early and resetting expectations on platform performance when necessary.
    • Cross‑Functional Bridge: Serve as the primary technical voice to the Customer Experience (CX), Marketing, Sales, and Leadership teams. You will translate technical constraints and roadmaps into clear updates for stakeholders.
Qualifications
  • Production K8s & Temporal: Expert‑level ability to debug Kubernetes internals (HPA logic, node scaling) and operate stateful workflow engines (Temporal) at scale.
  • Incident Command: Proven track record as a primary first responder, demonstrating the ability to aggressively reduce Time-to-Mitigation…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)

Job Posting Language
Employment Category
Education (minimum level)
Filters
Education Level
Experience Level (years)
Posted in last:
Salary