Cloud Operations Engineer
Job in
Vancouver, BC, Canada
Listing for:
MongoDB
Full Time
position
Listed on 2026-01-02
Job specializations:
-
IT/Tech
Systems Engineer
-
Engineering
Systems Engineer
Job Description & How to Apply Below
Mongo
DB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of Mongo
DB. The Cloud Operations Engineering team at Mongo
DB is a worldwide team responsible for the consistent operational success of every Mongo
DB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of Mongo
DB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team.
Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.
At Mongo
DB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems.
This role can be based out of our West Coast offices or remotely in the United States.
This role specifically follows a weekend support model (Wednesday-Sunday, Saturday-Wednesday, or Friday-Tuesday, with the next two days as the week-off) and requires adherence to West Coast Hours (9am to 5pm PST). If you're passionate about being an Cloud Operations Engineer, working with our largest enterprise customers, and are open to flexible, weekend-oriented scheduling, we encourage you to apply.
Responsibilities
Successfully coordinate and collaborate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer baseHelp scale the worldwide Cloud Operations Engineering team with the strategic implementation and refinement of new processes and toolsAssist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidentsMonitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolutionAutomate routine monitoring and troubleshooting tasksDiagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolutionAssist in performing root cause analysis after incident recovered; identifying any breakdowns in processes or workflows that contributed to the event and what changes need to be made to prevent similar eventsContribute to documentation of corner case scenarios, troubleshooting workflows and SOPs.Work alongside our product management, cloud engineering and support organizations by identifying areas for improvement in the management applications powering the Atlas infrastructureInform executive leadership and escalation management personnel of major outagesCoordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (proactively from automated monitoring or through reactive alerts via our Technical Services team)Requirements
Experience with being an on call Dev Ops, SRE, or Cloud Operations engineer (at least 2 years)Expertise with Linux system administration, configuration, troubleshootingExperience in monitoring, system performance data collection and analysis, and reportingKnowledge of database operations and conceptsExpertise with networking technologies like DNS, TCP/IP, etc.Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)Knowledgeable about a wide range of web and internet technologiesCapability to write small programs/scripts to solve both short-term systems problemsA CS/CE degree or equivalent experienceAt least 1 of the following programming languages:
Java, Go, Python, JavascriptA keen interest in learning new thingsNic…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here: