Job Description & How to Apply Below
Location: Quebec City
Elevate your career as a Cloud Site Reliability Engineer with NVIDIA. Ensure our Digital Marketing Services run reliably and efficiently while leveraging AWS Infrastructure.
NVIDIA seeks a skilled Site Reliability Engineer to enhance the performance and reliability of our Digital Marketing ecosystem. This role, requiring 3+ years of experience, focuses on automation, monitoring, and alerting within AWS Infrastructure. You will be responsible for deploying new applications, managing incidents, and improving deployment pipelines in a high-paced environment.
Key Responsibilities:
• Identify and categorize user-reported problems swiftly
• On-board new applications and AI/ML services
• Implement monitoring and alerting systems
• Automate daily tasks and ML CI/CD deployment
• Provide on-call support and resolve urgent incidents
Requirements:
• MS or BS in Computer Science/Engineering or equivalent experience
• 3+ years experience in live-site production environments
• Proficient in Python or Java for model serving
• Experience with Kubernetes and MLOps frameworks
• Strong problem-solving skills with automation focus
Leverage your expertise in SRE and automation to drive impactful solutions for NVIDIA’s innovative services.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×