Site Reliability Engineer
Job in
Seattle, King County, Washington, 98127, USA
Listed on 2026-08-17
Listing for:
Optomi
Full Time
position Listed on 2026-08-17
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Optomi, in partnership with a leading client in the Entertainment industry, is seeking an experienced SRE for their, Orlando, Seattle, New York, OR Burbank location. The right candidate will have Kubernetes expertise, with experience in Monitoring, capacity planning, and SLA / SLO's. The SRE will support multiple critical platforms, including a new AI platform, and an existing automation platform.
Responsibilities- Support and evolve critical platforms, including an AI platform, an existing automation platform, and a new data and observability platform.
- Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.
- Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.
- Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.
- Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/Open Tofu, Ansible, and Git Lab CI/CD.
- Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.
- Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.
- Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.
- Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows.
- 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.
- Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.
- Strong Linux administration and distributed systems experience.
- Experience with cloud platforms, particularly AWS, and container technologies such as Docker.
- Strong infrastructure-as-code experience with Terraform, Open Tofu, or similar tools.
- Hands-on experience with CI/CD pipelines and source control systems such as Git Lab or Git Hub Actions.
- Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or Cloud Watch.
- Strong understanding of networking fundamentals, APIs, and infrastructure automation.
- Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.
- Experience operating in large enterprise environments with cross-team collaboration and competing priorities.
- Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.
- Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.
- Experience with AI-assisted development tools such as Cursor, Claude Code, or Git Hub Copilot is preferred.
- Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×