×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Infrastructure Engineer

Job in Cambridge, Ontario, Canada
Listing for: VGS
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AWS
Salary/Wage Range or Industry Benchmark: 145000 - 185000 CAD Yearly CAD 145000.00 185000.00 YEAR
Job Description & How to Apply Below

VGS is the world's leader in payment tokenization. Large banks, aspiring fintechs, and growing merchants embed our universal token vault into their technology stack to manage the complexities of payment data tokenization across processors and networks, open banking, card issuance, omnichannel loyalty, PCI compliance, payment orchestration, and more. We empower our clients and partners by tokenizing sensitive payment data, limiting compliance scope, and consolidating payments to unlock revenue and business opportunities.

VGS provides processor-agnostic tokenization solutions via secure universal token vaults, iframes, mobile SDKs, tokenization proxies, APIs, and data orchestration tooling to support payment acceptance, card issuance, PII and bank account tokenization, and other payments value-added services. Some of the use cases we enable include multi-processor Network Tokenization, Account Updater, payment orchestration, secure settlement file processing, 3DS, and Risk provider connectivity.

We’re looking for a well‑versed, passionate Engineer who wants to play a key role in site reliability engineering and cloud operations of our global cloud infrastructure.

We’re seeking individuals with creative problem‑solving, enthusiasm for new technologies, and a desire to contribute to our product. You will likely be successful in this role if you identify with the following traits: attention to detail, problem solver, customer‑oriented, versatile, resilient, and confident.

What you will be doing at VGS (Responsibilities)
  • Architect and maintain scalable, reliable infrastructure: Design and optimize infrastructure for high availability, fault tolerance, and performance across distributed systems.
  • Lead incident management and root cause analysis: Own incident response processes, ensure swift resolution of issues, and drive post‑incident improvements to prevent recurrences.
  • Service monitoring and automation: Build and maintain automated monitoring, alerting, and healing systems that improve system health, reduce manual intervention, and minimize downtime.
  • Performance tuning and capacity planning: Identify bottlenecks and optimization opportunities, and implement scaling strategies to handle traffic spikes and growing workloads efficiently.
  • Collaborate with cross‑functional teams: Work closely with software engineers, product teams, and Dev Ops to enhance system reliability and delivery pipelines.
  • Improve operational processes: Champion continuous improvement initiatives in deployment, scaling, and performance testing, while advocating for the adoption of SRE best practices across the organization.
  • Mentorship and leadership: Provide technical mentorship to junior engineers, contribute to strategic decisions around infrastructure, and ensure best practices are implemented at scale.
  • Be proactive and innovative: We rely on your feedback to build a world‑class product.
  • Be a part of a team that believes in the core values of transparency, collaboration, grit, and humility; in going above and beyond what is required to do the right thing for our customers and the company; and in having fun while doing all this!
What we are looking for from you (Requirements)
  • Proven experience in Infrastructure/SRE roles, with a track record of managing production systems in complex, large‑scale environments.
  • Strong proficiency in AWS, including infrastructure‑as‑code (Terraform, Cloud Formation, etc.).
  • Solid understanding of cloud‑native architecture, Linux Systems, microservices, infrastructure‑as‑code (Terraform, Cloud Formation, CDK), CI/CD (CircleCI, Git Hub Actions, Argo), Git Ops, Authentication and Authorization, APIs and API Gateway, Docker, Kubernetes (EKS), Kafka (MSK), Java, Spring Framework, Python, and AWS services.
  • Strong plus if you are a database wiz.
  • Expertise in monitoring and observability tools like Prometheus, Grafana, Open Telemetry, New Relic, or similar tools to measure system health and performance.
  • Programming and scripting experience in languages such as Python, Go, Bash, or other relevant languages used in automating infrastructure.
  • Solid understanding of networking, security, and load balancing in cloud‑native…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary