Cloud Engineer
Listed on 2026-07-13
-
IT/Tech
Cloud Computing: Infrastructure & Operations, IT Infrastructure, Azure
Job Type Full-Time
OverviewThe Accelerator seeks a Cloud Engineer to design, build, and operate the secure cloud infrastructure that powers large-scale academic research on the information environment. Working as part of a small, high‑trust cross‑functional team, this individual will contribute across the full stack — from infrastructure and Dev Ops to backend services and data pipelines — and will have meaningful ownership over the technical systems that enable researchers at Princeton and across a global consortium to do their work.
This is a role for a senior, self‑directing engineer who is equally comfortable designing architecture and writing code, and who takes satisfaction in building systems that are reliable, secure, and well‑understood by the people who depend on them. The right candidate brings deep cloud expertise alongside strong software engineering fundamentals — someone who can own infrastructure end to end and contribute meaningfully to application development.
ResponsibilitiesCloud Infrastructure
- Design, deploy, and maintain cloud infrastructure on Azure, with responsibility for performance, cost‑effectiveness, and reliability across research and production environments.
- Architect and manage Databricks work spaces, including compute cluster configuration, access controls, and cost optimization for large‑scale data processing workflows.
- Manage Azure networking, storage, identity (Azure AD / Entra ), and resource governance across multiple environments.
- Implement infrastructure‑as‑code using Terraform and/or Bicep; maintain version‑controlled, reproducible infrastructure definitions including modules, remote state management, and PR‑based workflow.
- Deploy, operate, and maintain AKS clusters running containerized workloads — including containerized data crawlers — managing deploys, scaling, health monitoring, patching, and upgrades.
- Administer Azure Blob Storage, including lifecycle policies, redundancy configuration, and access tier management.
- Manage Azure networking and security, including Private Link, network rules, RBAC, and secrets hygiene across environments.
- Own Azure cost management: budget alerts, cost/cluster policies, anomaly detection and response, and Fin Ops practices to keep infrastructure spend predictable and efficient.
- Design, build, and maintain backend services, APIs, and data pipelines using Python and/or Type Script/Node.js.
- Develop and maintain CI/CD pipelines using Git Hub Actions, ensuring reliable and automated delivery of infrastructure and application changes.
- Build and maintain internal tooling that improves the experience and efficiency of the research and operations teams.
- Contribute to frontend integrations where needed; comfortable working across the stack on a small team.
- Develop and support data pipelines for ingesting, transforming, and serving large‑scale behavioral and social media datasets to researchers.
- Implement and maintain infrastructure for machine learning workflows, including model serving, experiment tracking, and compute resource management.
- Support integration with ML frameworks and tools (e.g., MLflow, Hugging Face, or equivalent) within the managed environment.
- Implement and maintain security controls across all systems, including encryption at rest and in transit, identity and access management, network segmentation, and secrets management.
- Design and operate environments meeting IRB, data governance, and institutional compliance requirements; ensure adherence to standards equivalent to SOC 2, HIPAA, or ISO 27001 as applicable.
- Conduct regular security reviews, vulnerability assessments, and penetration test coordination; manage remediation tracking.
- Implement audit logging, access controls, and data handling procedures for sensitive research data in compliance with IRB protocols and data use agreements.
- Operate, patch, and upgrade the self‑hosted observability stack — Grafana (dashboards), Loki (log aggregation), and Prometheus (metrics) — including security patching and version upgrades; implement and maintain alerting,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).