Site Reliability Engineer (SRE) Sylmar CA - CA - California
Listed on 2026-08-10
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Performance & Reliability Engineering troubleshooting latency, throughput, instability, slow login, and production issues.
Distributed Systems designing and optimizing highly available, scalable, fault-tolerant systems.
Kubernetes pods, resources, scaling, container behavior, computestorage.
Linux Administration & Performance Tuning CPU, memory, IO, networking, ulimits, OS internals.
Java Troubleshooting analyzing Java code and performance-critical application paths.
PostgreSQL query optimization, indexing, connection pooling, and database performance.
Cloud Azure Azure infrastructure, networking, and application operations.
Capacity Planning & Scalability workload modeling, benchmarking, forecasting, and scale analysis.
Observability & Monitoring metrics, logs, traces, monitoring, and root-cause analysis.
Load Testing & Benchmarking performance testing and evidence-based capacity decisions.
Messaging Systems especially ActiveMQ and middleware.
Automation & Scripting Python, Bash, Power Shell, deployment automation and pipelines.
CICD & Dev Ops deployment pipelines, automation, and operational practices.
LLM AI practical experience applying LLMs to engineering or operational workflows.
Networking & Infrastructure load balancers, traffic routing, cloud networking, and infrastructure.
SRE Platform Engineering reliability, availability, scalability, and production operations.
Root-Cause Analysis deep end-to-end troubleshooting using metrics and production data.
Communication & Leadership collaborating with developers and explaining complex technical findings to stakeholders.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).