Senior Platform Operations Engineer, Infrastucture
Listed on 2026-08-22
-
IT/Tech
Unix/Linux, Cloud Computing: Infrastructure & Operations, IT Infrastructure, Systems Engineer
At Viome, we are driven by a singular mission: to help people live a healthy, disease-free life. This mission guides our actions, fuels our passion, and shapes the impact we aim to have in the world. Our core values - Be Bold, Be Collaborative, Be Frugal, and Grow Continuously - underpin our approach to achieving this goal. If you are motivated by the idea of working in an environment that prioritizes bold innovation, teamwork, efficient resource use, and continuous learning, all towards promoting health and preventing disease, we warmly invite you to apply.
Join us in our journey to transform lives and create a healthier future for all.
We are looking for a Senior Platform Operations Engineer, Infrastructure to take ownership of Viome’s Azure-based service platform with a clear mandate: make it dramatically simpler through robust Unix-based systems engineering.
This is a consolidation role for a seasoned systems administrator. The work involves migrating a suite of abstracted services onto a deliberately stable target platform:
Linux VMs, systemd-supervised services, Apache, HAProxy, and Nginx routing. You must possess strong Unix knowledge to operate and eventually transform the current stack safely. We are looking for a traditional operations background where success is measured by the removal of unnecessary layers and a genuine bias toward simplicity and host-level stability.
- Plan and execute zero-downtime migrations of services from AKS/containers to VM-based hosting: systemd unit authoring and service supervision (restart policy, resource limits, sandboxing), Apache, HAProxy and/or nginx as reverse proxy and TLS terminator, certificate automation.
- Build and own the deployment scripts for the target platform: build → test → archive → ship → symlink-flip → health-check → rollback, scripted in bash.
- Preserve two non-negotiable in variants while simplifying everything else: immutable, commit-traceable artifacts and scripted rollback.
- Inventory and retire stale infrastructure: dormant deployments, unused DNS records, orphaned firewall rules, unpinned image tags.
- Maintain host hygiene for consolidated services: OS patching discipline, runtime vendoring, log rotation, centralized log aggregation.
- Familiarity with Uptime Kuma, Nagios or similar.
- Operate and migrate data stores:
PostgreSQL and/or MySQL.
- Operate and troubleshoot workloads during the transition, focusing on the networking layer, load balancing, and core platform services.
- Administer the hub network:
Azure Firewall rules, VPN gateways, VNet peering, public and private DNS zones and reason about a packet’s full path from public IP to service. - Support the existing release process and network and infrastructure operations until each service is migrated to the new VM-based standard.
- Keep the observability stack healthy (Open Telemetry, ELK, Grafana, uptime and cost monitoring) and carry its essentials forward to the simplified platform.
- Coordinate cross-cloud dependencies with AWS: DNS/edge routing, queue consumers, and egress IP allow lists.
- L2+ operations support; manage runbooks for external L0, L1 support.
- Own the external integrations most at risk during migration: e-commerce and subscription platforms, messaging/notification providers, and clinical/health-data partners — webhook delivery, signature verification, idempotency, and retry semantics.
- Raise the security baseline as you consolidate: secrets management, webhook authentication, least-privilege network access, and data-retention hygiene. Findings from an internal review are ready for you to remediate; the instinct to spot and close this class of issue — and not create more — is part of the job.
- 7+ years operating production Unix/Linux systems, with deep knowledge of systemd, process supervision;
Apache, HAProxy - Strong shell plus one scripting language (bash/Python/PHP) with a track record of building deploy and rollback tooling, not just using it.
- Demonstrated reverse-engineering ability: taking ownership of an undocumented production service…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).