Principal Platform Engineer
Listed on 2026-07-20
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Location: New York
Staff / Principal Platform Engineer
Location: New York City | Hybrid
Department: AI Platform & Infrastructure Team
Reports to: Vangie Shue - Principal Engineering Manager
About App GateApp Gate secures and protects an organization’s most valuable assets with its high performance Zero Trust Network Access (ZTNA) solution and Cyber Advisory Services. App Gate ZTNA is the only direct-routed Zero Trust solution built for peak performance, superior protection and seamless interoperability. App Gate Cyber Advisory Services harden your security posture and ensure business continuity. App Gate safeguards Fortune 500 enterprises and government agencies worldwide.
Learn more at
As we expand our platform, we are standing up a new AI Platform & Infrastructure team: the engine room of App Gate’s AI strategy. This team owns the infrastructure layer that every next-generation security capability is built on, from network observability to AI-driven threat detection and the secure operation of emerging Agentic AI systems.
We’re looking for a Staff or Principal Platform Engineer to build and operate the foundational platform behind App Gate’s AI products. You combine deep Dev Ops and cloud infrastructure expertise with hands-on experience operationalizing AI/ML systems, and you treat observability as a first-class engineering discipline. This is a rare opportunity to join a small, private, high-impact company where your work directly shapes the architecture, reliability and core platform that defines the future of security.
You’ll own the platform spanning APIs, cloud and self-managed solutions and AI/ML infrastructure, and you’ll make it fast, reliable and observable s is a high-leverage, hands-on role for a senior engineer who sets technical direction and still ships.
Key Responsibilities- Build the Platform: design, build and operate the cloud infrastructure, services and pipelines that App Gate’s AI and cloud products run on. Strong experience with self-managed technologies (kafka, elasticsearch) and Kubernetes are a must.
- Infrastructure as Code & Deployment Orchestration: Terraform and Helm for cloud provisioning, service deployment and configuration management.
- Implement Observability: instrument APIs, cloud services and AI/ML infrastructure with metrics, logging, tracing and alerting, and define SLOs and operational health metrics that teams trust.
- Data Platform: real-time and batch data ingestion pipelines, feature stores and data quality.
- Integrations: third-party connectors, APIs and platform integrations.
- Operationalize AI/ML: build model serving and inference pipelines, experiment tracking and the MLOps tooling for deployment, versioning, drift monitoring and lifecycle management.
- Engineer for reliability & automation: apply SRE practices to reduce toil, improve resilience and keep latency and uptime within target across the platform. Automate everything - deliver infrastructure-as-code, CI/CD and self-service tooling so product teams ship safely and quickly.
- Set technical direction: define platform standards, architecture and best practices, and raise the engineering bar through design reviews and mentorship.
- Collaborate cross-functionally: partner with data scientists, product teams and leadership to align platform investment with App Gate’s strategic vision.
- Experience: extensive platform, infrastructure or SRE engineering experience, with a track record of operating production systems ff-level candidates typically bring 8+ years and Principal-level candidates 12+ years, though we hire on demonstrated impact.
- Dev Ops depth: strong command of infrastructure-as-code (Terraform or equivalent), CI/CD, containers and orchestration (Docker, Kubernetes), and cloud platforms (AWS).
- Observability expertise: hands-on experience implementing observability across APIs, cloud services and distributed systems using tools such as Prometheus, Grafana, Open Telemetry, the ELK stack or comparable, including SLO and error-budget practice.
- Data platform skills: familiarity with real-time and batch ingestion pipelines, feature stores and data quality at production scale.
- Engineering craft: fluenc…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).