×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer - Infrastructure and Agentic Automation

Job in Santa Clara, Santa Clara County, California, 95050, USA
Listing for: CYNET SYSTEMS
Full Time position
Listed on 2026-09-24
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 55 - 60 USD Hourly USD 55.00 60.00 HOUR
Job Description & How to Apply Below
Pay Range: $55.00hr - $60.00hr Job Overview:
Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the successful candidate will help design, scale, and secure enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence. The candidate will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines.

This role is ideal for a professional passionate about reducing toil through code, leveraging modern AI agent frameworks like Model Context Protocol (MCP), and ensuring 24/7 system reliability across massive fleet environments.

Key Responsibilities:

Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management. Drive configuration management across the environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments. Design and maintain secure, scalable CI/CD pipelines (Git Lab CI/CD) and Git Ops workflows for automated system configuration, package rollout, and patch management.

Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows. Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies. Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.

Qualifications &

Required Skills:

5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments. Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning. Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture. Proven track record of building automated CI/CD pipelines and Git Ops workflows using modern version control (Git).

Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks. Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability. Experience managing and securing Windows infrastructure alongside Linux. Proficiency in languages such as Python, Go, Power Shell, or Bash for automation and tooling development. Benefits:
Our Benefits Include:
Medical, Dental, and Vision Insurance 401(k) Retirement Plan Health Savings Account (HSA) Disability Insurance (Short-Term and Long-Term) Life and AD&D Insurance Paid Sick Leave (where required by applicable state or local law) Supplemental Insurance Plans Identity Theft Protection Pet Insurance Employee Wellness Programs Employee Assistance Program (EAP) Career Growth and Professional Development Opportunities

Disclaimer: Benefits eligibility, accrual rates, and usage limits may vary based on employment status, length of service, and work location. Paid Sick Leave is provided in strict accordance with applicable state and municipal mandates. Cynet Systems Inc. reserves the right to modify, amend, or terminate any benefit plans at any time in accordance with applicable laws. About Cynet Systems Founded in 2010 and headquartered in the Washington, DC metro area, Cynet Systems Inc.

is a leading technology staffing and workforce solutions company serving Fortune 500 companies, government agencies, and enterprise organizations across the United States and Canada. We deliver agile, scalable talent solutions across IT, engineering, life sciences, clinical, and professional staffing, powered by a high-performing recruitment engine operating across North America and Asia. As a nationally and locally certified Minority Business Enterprise (MBE), Cynet Systems is committed to helping organizations build high-performing teams while empowering professionals to grow rewarding careers.

Our organization is certified to ISO 9001,…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary