×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer; AI Acceleration- Hybrid

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: Calance
Full Time position
Listed on 2026-07-10
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 140000 - 220000 USD Yearly USD 140000.00 220000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (AI Acceleration)- Hybrid

Work Authorization:
Must be authorized to work in the U.S. without current or future sponsorship

You will be a core member of the SRE team, responsible for the reliability, automation, and observability of the infrastructure that the company runs on. You will work across colocation, on-premises lab environments, and cloud platforms — and you will own your systems end-to-end, from initial provisioning through live incident response.

You will partner with hardware and software development teams to support their workload needs, including CI/CD pipelines and automation layer and the associated CI/CD pipelines and automation layer for software tooling. You will also support customer-facing environments where partners collaborate on hardware and software deployments.

What You Will Do Infrastructure Operations
  • Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
  • Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up.
  • Conduct capacity planning and hardware lifecycle management for assigned infrastructure domains; track and report cloud spend for your domains to support Fin Ops and workload placement decisions.
Automation & Infrastructure as Code
  • Own IaC and configuration management (Terraform, Ansible) for your infrastructure domains — all provisioning and changes through code, not manual steps.
  • Build, deploy and document automation to eliminate toil: host lifecycle management, fleet health checks, auto-remediation workflows, and self-service tooling for engineering teams.
  • Develop networking automations for cluster interconnects, VLAN management, and lab network configurations.
  • Contribute to shared IaC modules and automation libraries used across the global SRE and data center services teams.
Observability & Incident Response
  • Design and maintain monitoring dashboards, alerting, and SLIs (Prometheus/Grafana, Data Dog) for your infrastructure domains — ensuring signal quality, actionable alerts, and contributing to AIOps-driven detection workflows that reduce time to detect and respond.
  • Participate in on-call rotation; triage and resolve incidents from bare metal to application layer using structured, AI-assisted workflows — distinguishing infrastructure faults from software or hardware product issues and escalating with clear context.
  • Produce high-quality RCA reports for P0/P1 incidents with root cause analysis and tracked action items.
  • Detect performance issues, recommend solutions, and implement fixes that permanently improve system reliability.
Customer & Platform Development Services
  • Support and operate platform services used by both internal teams and external customers for hardware and software deployment collaboration with d-Matrix.
  • Ensure QoS and uptime commitments for customer-facing environments; elevate reliability risks proactively.
  • Document platform configurations, access procedures, and operational runbooks for customer environments.
  • Maintain high-quality runbooks, architecture diagrams, and troubleshooting guides — documentation is part of the job, not an afterthought.
  • Partner with the Dev Ops team to ensure infrastructure reliability supports CI/CD pipeline performance and developer experience.
  • Serve as a technical resource for engineering teams — sharing operational knowledge and raising infrastructure risks early.
What You Will Bring Required
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics.
  • Hands-on experience with colocation or on-premises server infrastructure — physical hardware, rack networking, and bare-metal provisioning.
  • Hands-on experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary