×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: The Mice Groups, Inc.
Full Time position
Listed on 2026-07-17
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

We’re partnering with an innovative leader in AI infrastructure and semiconductor technology that’s building the next generation of AI computing platforms. They’re looking for an experienced Site Reliability Engineer to help build, automate, and operate the critical infrastructure powering AI development.

This is a hands‑on role for someone who enjoys owning production infrastructure—not just supporting deployments. You’ll work across Linux, Kubernetes, cloud platforms, and bare‑metal environments while improving reliability, scalability, and automation for mission‑critical systems.

What You’ll Be Doing
  • Own and operate production infrastructure across on-premises, colocation, and cloud environments
  • Manage Linux servers, Kubernetes clusters, networking, and storage
  • Build Infrastructure as Code using Terraform and/or Ansible
  • Develop automation with Python and/or Bash to improve operational efficiency
  • Design and maintain monitoring, alerting, and observability solutions
  • Participate in incident response, troubleshooting, and root cause analysis
  • Partner with engineering teams to deliver highly reliable infrastructure
What We’re Looking For
  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or Systems Administration
  • Hands‑on Kubernetes operations experience
  • Experience with Terraform and/or Ansible
  • Scripting experience with Python and/or Bash
  • Experience with Prometheus, Grafana, or Datadog
  • Strong troubleshooting and production incident response experience
Nice to Have
  • AI infrastructure or GPU cluster experience
  • High-performance computing (HPC)
  • Infini Band or RoCE networking
  • Semiconductor or hardware industry experience
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary