Site Reliability Engineer
Listed on 2026-08-20
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Unix/Linux
Site Reliability Engineer
Looking for an experienced Site Reliability Engineer who can strengthen our team's platform engineering and operational capabilities. You will play a key role in supporting Kubernetes infrastructure managed through Rancher, improving system reliability and automation, and advancing infrastructure-as-code and Git Ops practices across our environment.
Key Responsibilities- Platform Operations:
Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency. - Cluster Management:
Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher. - Linux Systems Administration:
Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support. - Infrastructure as Code:
Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches. - Git Ops and Deployment Automation:
Support and improve Git Ops workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner. - Collaboration:
Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions. - Continuous Improvement:
Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.
- 5+ years of professional experience in site reliability engineering, platform engineering, Dev Ops, or systems engineering roles.
- Hands-on experience operating and supporting Kubernetes platforms in production environments.
- Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
- Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
- Experience with Rancher for Kubernetes cluster management and platform operations.
- Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management.
- Demonstrated success working in Agile teams (Scrum, Kanban).
- Kubernetes:
Cluster operations, upgrades, networking, storage, troubleshooting, and workload support. - Platform Management:
Rancher or similar Kubernetes management platforms. - Linux:
Advanced administration of Linux/Unix systems. - Infrastructure as Code:
Strong IaC experience;
Cluster API (CAPI) preferred. - Git Ops/CI-CD:
ArgoCD, Git version control, and deployment automation practices. - Scripting/Automation:
Bash, Python, or similar scripting languages for automation and operational tooling.
- Experience with hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, GCP).
- Experience with Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
- Familiarity with security best practices for Kubernetes and Linux platforms.
- Experience supporting scientific research environments, high-performance computing, or computational science workflows.
- Knowledge of CI/CD pipeline development and platform automation patterns.
- BS in Computer Science, Software Engineering, Information Technology, or related field preferred; or equivalent professional experience.
US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand and total workforce solutions. To know more about US Tech Solutions, please visit US Tech Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
AIStatement
By applying, you acknowledge that AI-assisted tools may be used during hiring.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).