More jobs:
AI Inference Systems Reliability Lead
Job Description & How to Apply Below
This hands-on role requires a seasoned engineer with over 7 years in reliability engineering for large-scale distributed systems. You will be responsible for defining service-level objectives, leading incident management, and developing advanced reliability mechanisms. If you have strong programming skills in popular languages and a deep understanding of SLOs, this position is for you.
Key Responsibilities:
• Define reliability goals and align engineering efforts
• Design and develop fault detection and recovery systems
• Lead incident management and root-cause analysis
• Architect reliable systems with observability features
• Create chaos testing and load simulation tools
Requirements:
• Bachelor's or master's degree in a relevant field
• 7+ years of experience in distributed systems
• Strong programming experience in Python, C++, or Go
• Knowledge of reliability architectures and practices
• Excellent communication and leadership abilities
Join Cerebras to leverage your expertise in ensuring reliable AI systems and driving significant technological advancements.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×