Platform Engineer; Remote
Toronto, Ontario, C6A, Canada
Listed on 2026-09-04
-
IT/Tech
SRE/Site Reliability, IT Infrastructure, Systems Administrator, Cloud Computing: Infrastructure & Operations
Salary Range: $ To $ Annually
As a
Platform Engineer
at Think On, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span Zabbix and Prometheus/Grafana for monitoring and dashboards, Opsgenie for alert routing and on-call management, and infrastructure-as-code tooling (Ansible, Terraform, Git Lab CI/CD) to deploy and manage it all. You'll work with in a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and Dev Ops teams.
Think On is a remote-first organization. At this time, we are welcoming candidates from Canada for this position.
Please note that this listing is for a current vacancy at Think On.
We are looking for qualified candidates who are eager to contribute and grow with us.
You Will:Monitoring & Observability
- Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
- Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
- Configure and manage Ops genie for alert routing, escalation policies, and on-call schedules.
- Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
- Integrate monitoring systems with ticketing, communication, and incident management tools.
- Maintain and optimize monitoring infrastructure — database tuning, storage management, high availability.
- Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
- Use Terraform for provisioning infrastructure resources where applicable.
- Build and maintain Git Lab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
- Follow infrastructure-as-code practices — version-controlled, peer-reviewed, reproducible.
- Act as a first responder for monitoring-related incidents and alerts.
- Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
- Escalate to the appropriate teams (network, security, infrastructure) when needed.
- Document incidents, root causes, and resolutions. Contribute to post-incident reviews.
- Provide technical support to internal teams on monitoring tools and dashboards.
- Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
- Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting).
- Support compliance requirements (ISO 27001, SOC
2) by maintaining audit logging, access controls, and security configurations for monitoring systems.
- Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
- Eligible to obtain Secret Level Clearance within your first 3 months of employment. This requires the successful candidate to be a Canadian Citizen and have 10+ years of verifiable police background history.
- Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations).
- Hands-on experience with Zabbix(or comparable: Nagios, Icinga,Checkmk).
- Working knowledge of Prometheus and Grafana— writing exporters, building dashboards, PromQL.
- Experience with alert management and on-call tooling(Ops genie, Pager Duty, or similar).
- Comfort with Linux systems administration (this is a Linux-heavy environment).
- Proficiency in scripting and automation— Bash and Python at minimum.
- Experience with at least one IaCtool (Ansible, Terraform).
- Familiarity with CI/CD pipelines(Git Lab CI, Git Hub Actions, Jenkins, or similar).
- Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
- Strong diagnostic and troubleshooting skills — you can work through a problem methodically.
- Clear written and verbal communication —you'll document your work and explain technical issues to varied audiences.
- Attention to detail — monitoring generates a lot of data, and you need to separate signal from noise.
- Comfort working independently in a remote environment while collaborating across teams.
- Ability to stay…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: