More jobs:
Site-Reliability Engineer, Operations
Job in
Irving, Dallas County, Texas, 75084, USA
Listed on 2026-08-30
Listing for:
Scorpion Therapeutics
Full Time
position Listed on 2026-08-30
Job specializations:
-
IT/Tech
SRE/Site Reliability, IT Infrastructure
Job Description & How to Apply Below
Responsibilities:
- Serve as primary responder in scheduled on-call for production incidents across SOX- and FDA-regulated clinical applications; triage, elevate and remediate per runbooks/incident playbooks.
- Execute approved production runbooks and maintenance scripts; document privileged actions for SOX ITGCs and FDA audit-trail requirements.
- Implement/maintain application-layer instrumentation (dashboards, alert thresholds, SLO monitors) and monitor SLO/error-budget consumption;
** escalate
* * when budgets are at risk. - Support releases/deployments: verify deployment health, monitor post-deploy behavior, and perform rollback when required.
- Participate in post-incident reviews: reconstruct timelines, identify contributing factors, and drive remediation action items to closure.
- Maintain audit-ready access logs, privileged-action records, and role-change documentation.
- Automate recurring operational work; track manual-intervention rate and reduce it by retiring runbook steps into automation.
- Author/update runbooks and known-issue documentation; improve operate discipline.
- Collaborate with application engineering to validate production fixes; contribute to on-call tooling and alert-quality improvements.
- Use AI coding assistants/agentic workflows for monitoring, triage, runbooks, and automation scripting (review as quality gate).
- BS in CS/IS/SE or related, or equivalent experience.
- 5+ years in SRE/production ops/Dev Ops/platform engineering (production support).
- 2+ years on-call primary responder experience for Tier 1/business-critical systems.
- Experience executing production runbooks/maintenance/change procedures with privileged-action controls.
- Experience with production monitoring/alerting dashboards in an observability platform.
- Experience in post-incident review/blameless retrospectives.
- Hands-on use of AI coding assistants for operational automation.
- Familiarity with SLO/error budgets and alerting design.
- Experience reducing toil via automation.
- Experience in SOX-controlled IT or CLIA/CAP-regulated environments.
- Cloud-native observability at the application instrumentation layer; telemetry/tracing/APM.
- Domain experience in clinical diagnostics/lab information systems/digital health.
- Familiarity with SAST/DAST and secure CI/CD; deployment pipelines/container orchestration/release automation.
- Relevant cloud/security certifications.
- Ability to sit/stand/work at a computer for extended periods.
- On-call includes required after-hours response (evenings/weekends); periodic travel may be required.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×