Job Description & How to Apply Below
Join the forefront of AI transformation as our Site Reliability Engineer, focusing on automation and reliability across diverse infrastructure domains. Collaborate with engineering teams to tackle challenges head-on.
This position emphasizes 5+ years of experience in SRE, where you will manage both on-premises and cloud environments.
Your role will ensure high uptime and reliability for customer-facing platform services while automating infrastructure operations. Skills in Python, Bash, and IaC tools will be essential.
Key Responsibilities:
• Ensure availability of Colo server fleets and cloud environments
• Conduct capacity planning for infrastructure needs
• Automate provisioning and deployment processes
• Develop incident response and monitoring frameworks
• Document all infrastructure operations and configurations
Requirements:
• Bachelor’s or Master’s in Computer Science or related field
• 5+ years' experience in SRE or infrastructure management
• In-depth Linux systems knowledge
• Experience with infrastructure automation tools
• Proven incident response capabilities
Contribute to cutting-edge AI solutions by elevating infrastructure reliability and automation.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×