More jobs:
Site Reliability Engineer at Cohere
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-08-06
Listing for:
Visa Hunt
Full Time
position Listed on 2026-08-06
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, Network Engineer
Job Description & How to Apply Below
As part of the Model Serving team at Cohere, you will develop and operate a platform for large language model API endpoints. This role requires collaborating with various teams to ensure efficient deployment and optimized performance in high-availability environments. You’ll engage with customers to tailor specific deployments that meet their needs, marking a significant impact in AI technology.
Key Responsibilities:
• Build self-service systems for managing services
• Automate Kubernetes deployments for language models
• Ensure environment observability and resilience
• Participate in on-call rotation for SLO adherence
• Foster relationships with internal teams for feedback integration
Requirements:
• 5+ years in production infrastructure engineering
• Experience with Kubernetes and large distributed systems
• Background in GCP, Azure, AWS, or similar
• Excellent troubleshooting and collaboration skills
• Familiarity with GPUs and distributed system technologies
Drive infrastructure excellence and build impactful AI systems as a Site Reliability Engineer at Cohere.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×