Site Reliability Engineer Position
Job Description & How to Apply Below
Join Cohere as a Site Reliability Engineer to build advanced AI platforms. Utilize your expertise in Kubernetes and machine learning in a remote-friendly environment.
As part of the Model Serving team at Cohere, you'll develop scalable infrastructures that support complex NLP applications. With over five years of engineering experience, this role emphasizes your skills in deploying high availability distributed systems on cloud platforms. Your input will directly impact product deployments while enhancing collaboration with developers.
Key Responsibilities:
• Manage and deploy AI model services efficiently
• Implement automated systems for performance observability
• Participate in on-call rotations to ensure service levels
• Actively gather and implement feedback from peers
• Drive knowledge sharing practices within the team
Requirements:
• Over 5 years in large-scale systems engineering
• Deep experience with Kubernetes and GPU workloads
• Knowledge of multi-cloud infrastructures including AWS
• Proficiency in Linux systems and troubleshooting
• Familiarity with programming in languages like C++
Leverage your technical skills to enhance Cohere's AI capabilities, driving innovation in language models.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×