Senior Platform Engineer
Listed on 2026-07-25
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Let’sTango!
Where Innovation Meets Impact .
At Tango Analytics, we’re all about helping businesses make smarter decisions through powerful technology, insightful data, and a whole lot of collaboration. Whether you are a creative thinker, a strategic planner, a tech wizard, or a customer champion, there's a place for you on our team. We believe work should be meaningful and fun — so if you're ready to make a difference while enjoying the journey, come join us and let's Tango!
We are looking for an Senior Platform Engineer to join our dynamic and growing Platform Engineering team.
Role SummaryWe are building a platform engineering function from the ground up and this role is at the center of it. As a Senior Platform Engineer you will be a founding team member with a clear, three-part mission: fully rearchitect and codify ourcloud estate on AWS and Azure, stand up a world-class SRE and observability practice, and build an Internal Developer Platform with golden paths that make shipping software fast, safe, and self-service.
This is a high-ownership individual contributor role. You will work directly with the Platform Engineering Manager to translate strategy into production infrastructure, drive adoption across engineering teams, and set the technical barfor how the platform is built and operated.
If you have strong opinions about infrastructure-as-code, love building things other engineers depend on every day, and want genuine end-to-end ownership — this role is for you.
Key Responsibilities- Migrate all existing AWS and Azure infrastructure to Open Tofu/Terraform and Ansible; establish module standards, remote state, and Git Ops-based plan/apply pipelines — no unmanaged resources
- Audit the cloud estate against the AWS and Azure Well-Architected Frameworks; produce a remediation backlog and drive it to completion across networking, IAM, landing zones, account structure, and cost governance
- Implement policy-as-code (OPA/Conftest, AWS SCPs, Azure Policy) to enforce security, tagging, and complianceguardrails at the platform layer — governance embedded, not bolted on
- Build and maintain reusable Terraform modules for compute (EKS, AKS, EC2), networking, storage, databases, and identity as shared building blocks for all engineering teams
- Define Fin Ops standards: tagging taxonomy, cost allocation dashboards, rightsizing recommendations, and reservedcapacity planning across both clouds
- Design and implement the full observability stack: metrics (Prometheus/Datadog), logs (Loki/Open Search), traces(Tempo/Datadog APM), and dashboards (Grafana) — instrumented end-to-end via Open Telemetry
- Define SLIs and SLOs for all platform shared services and critical applications; build error budget dashboards andburn-rate alerting — alert on symptoms, not raw metrics
- Establish the SRE practice from scratch: incident runbooks, post-incident review templates, and at least one chaosengineering exercise (AWS FIS or equivalent)
- Partner with engineering teams to instrument their services, define meaningful alerts, and build operationaldashboards — reliability is a shared responsibility, not a platform team tax
- Build capacity planning models for compute and storage so engineering leadership can make data-driven scalingdecisions
- Deploy and operate a developer portal (Backstage, Git Hub or equivalent) as the single front door: service catalog,scaffolding templates, runbooks, API docs, and on-call ownership all in one place
- Build and maintain golden paths for the highest-frequency developer workflows: new service creation, Kubernetesdeployment, database provisioning, secrets management, and CI/CD pipeline setup - opinionated defaults withescape hatches for legitimate edge cases
- Own the CI/CD platform layer: standardized pipeline templates (Git Hub Actions, Git Lab CI), reusable workflowlibraries, container image build and scan pipelines, and environment promotion workflows with security scanning(SAST, Snyk) built in by default
- Own Kubernetes platform operations: EKS and/or AKS cluster lifecycle, Helm chart standards, admission controllers,RBAC, network policies, and service mesh (Istio or Linkerd)
- Build the self-service…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).