Lead Site Reliability Engineer
Listed on 2026-07-06
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Lead Site Reliability Engineer (Generative AI Platform)
Optomi, in partnership with a leading entertainment company, is seeking a Lead Site Reliability Engineer (Generative AI Platform) to join their team in Orlando, FL. This is a long-term hybrid opportunity (4 days onsite) supporting the organization's enterprise Generative AI platform by driving cloud infrastructure strategy, automation, platform reliability, and scalability across multi-cloud environments. The Lead Site Reliability Engineer will play a key technical leadership role supporting the enterprise Generative AI platform (JedAI).
This engineer will help architect, build, and operate highly available cloud infrastructure across GCP, AWS, and Azure while leading Kubernetes operations, Infrastructure-as-Code initiatives, observability, and CI/CD modernization. The ideal candidate is a hands-on technical leader with deep SRE expertise, strong automation skills, and experience supporting large-scale distributed systems in production.
- Building infrastructure that powers enterprise Generative AI applications!
- Working with modern cloud-native technologies across GCP, AWS, and Azure!
- Leading platform reliability and automation initiatives in a highly collaborative environment!
- Influencing long-term cloud architecture and engineering best practices!
- Mentoring engineers while remaining hands-on with cutting-edge technologies!
- Solving complex distributed systems challenges at enterprise scale!
- 7+ years of experience in Site Reliability Engineering, Dev Ops, or Platform Engineering
- Extensive production experience administering Kubernetes clusters and Helm Charts
- Expert-level Terraform and Infrastructure-as-Code experience
- Strong CI/CD experience with Harness, Git Hub Actions, Git Lab, Jenkins, or Azure Dev Ops
- Multi-cloud experience with GCP (preferred), AWS, and Azure
- Strong Python and Bash scripting skills
- Experience supporting PostgreSQL, Redis, Kafka, MongoDB, and Vault in production
- Hands-on experience with observability tools including Splunk, Prometheus, Open Telemetry, or App Dynamics
- Strong troubleshooting skills supporting highly available distributed systems
- Technical leadership, mentoring, and Agile/Scrum experience
- Experience supporting AI, Machine Learning, or Generative AI platforms
- Lead the design, implementation, and reliability of the enterprise Generative AI platform
- Build, maintain, and optimize Kubernetes environments using Infrastructure-as-Code and Terraform
- Develop and enhance CI/CD pipelines, deployment automation, and cloud-native platform services
- Improve platform scalability, observability, monitoring, and incident response while maintaining 99.99% uptime
- Support production databases, messaging systems, and cloud infrastructure across GCP, AWS, and Azure
- Mentor engineers, establish SRE best practices, and partner with architecture, security, product, and engineering teams to deliver highly available AI solutions
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).