×
Register Here to Apply for Jobs or Post Jobs. X

Lead Site Reliability Engineer

Job in Orlando, Orange County, Florida, 32885, USA
Listing for: KellyMitchell Group
Full Time position
Listed on 2026-07-22
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 70 - 80 USD Hourly USD 70.00 80.00 HOUR
Job Description & How to Apply Below

Our client, a leading entertainment and technology provider, is seeking a Lead Site Reliability Engineer to join their team! This position is a 96-week contract and can be based in Orlando, Florida;
Burbank, California; or Seattle, Washington. Candidates must be able to work onsite as required. This role will support a large-scale Generative AI platform and cloud infrastructure environment responsible for enabling AI capabilities across enterprise applications and experiences.

This position is W-2 only. We do not work with third-party firms or C2C arrangements for this role.

Core Responsibilities
  • Lead the design, implementation, and maintenance of cloud infrastructure supporting Generative AI and data service workloads across GCP, AWS, and Azure
  • Manage Kubernetes clusters, scaling strategies, node pools, resource allocation, autoscaling, and platform reliability initiatives
  • Build and maintain Infrastructure as Code solutions using Terraform and automated deployment pipelines through Harness
  • Drive platform reliability initiatives to support 99.99% uptime service level objectives across distributed, multi-region environments
  • Implement observability, monitoring, alerting, and tracing solutions using Splunk, Open Telemetry, Prometheus, and App Dynamics
  • Partner with architecture, capacity planning, and engineering teams to ensure scalability, performance, and operational excellence
  • Mentor Site Reliability Engineers and Dev Ops specialists while providing technical leadership and infrastructure strategy guidance
Required Skills/Experience (Must-Haves)
  • 7+ years of Site Reliability Engineering, Dev Ops, Platform Engineering, or related experience
  • Extensive experience managing Kubernetes clusters in production environments, including scaling, security, and high-availability architectures
  • Strong hands‑on experience with both AWS and GCP cloud platforms within large-scale enterprise environments
  • Advanced experience with Terraform and Harness for infrastructure automation and deployment management
  • Experience supporting highly available, multi-region, distributed systems with large-scale traffic and capacity planning requirements
  • Strong scripting and automation experience using Python, Bash, and YAML
  • Experience supporting AI platform infrastructure, AI model deployment at scale, vector databases, or Generative AI platform environments
  • Experience with AI-powered development tools, AI agents, or AI-assisted engineering workflows
  • Experience deploying custom AI or machine learning models in cloud environments
Preferred Skills/Experience (Nice-to-Haves)
  • Master's degree in Computer Science, Information Systems, or a related field
  • Experience with vector databases such as Pinecone or Qdrant
  • Experience with PostgreSQL, Redis, Kafka, MongoDB, and Vault in production environments
  • Experience with enterprise observability tools including Splunk, Open Telemetry, Prometheus, and App Dynamics
  • Experience working in hybrid cloud or multi-cloud environments spanning AWS, GCP, and Azure
  • Technical leadership and mentorship within complex cloud infrastructure environments
  • Strong troubleshooting and problem‑solving skills across distributed systems
  • Ability to communicate highly technical concepts to both technical and non-technical stakeholders
  • Strategic thinking around scalability, reliability, and capacity planning
  • Strong collaboration and interpersonal skills within cross-functional teams
  • Location:

    Orlando, FL;
    Burbank, CA; or Seattle, WA
  • Onsite as required
  • Multi-cloud enterprise environment supporting Generative AI platforms and large-scale distributed systems
  • Pay Range:
    The approximate pay range for this position is between $70.00 and $80.00 per hour. Please note that the pay range provided is a good faith estimate. Final compensation may vary based on factors including but not limited to background, knowledge, skills, and location. We comply with local wage minimums.
  • Employee-Owned Profit Sharing (ESOP)
  • 401K offered
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary