Principal Azure Capacity Manager
Job in
New York City, Richmond County, New York, USA
Listed on 2026-08-18
Listing for:
Veterans Sourcing Group
Full Time
position Listed on 2026-08-18
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Principal Azure Capacity Manager
Duration: 12+ Months (Possible extension)
Location:
New York, NY 10286 Onsite Role (4 days a week)
Responsibilities:
- Principal Azure Capacity Manager (Consultant) to lead capacity planning and optimization for an Azure public cloud project operating to High requirements.
- This role ensures adequate, resilient capacity and buffer across compute, storage, network, and platform services; supports Site Reliability Engineers (SREs) with performance and reliability goals; and drives evidence-based compliance with program High control expectations.
- Own the end-to-end Capacity Management operating model for Azure services in scope of the High program—planning, modeling, forecasting, monitoring, tuning, and governance.
- Ensure sufficient capacity and engineered buffer to meet service-level objectives (SLOs), recovery objectives (RTO/RPO), and regulatory/contractual requirements, with particular focus on U.S.
-only region restrictions and continuous monitoring. - Partner closely with SREs to operationalize capacity practices through IaC, gated change control, performance baselines, autoscaling policies, and resilience patterns.
- Contribute to documentation and evidence (e.g., SSP updates, control narratives, POA&M items, continuous monitoring artifacts).
- Capacity Planning & Optimization:
Build and maintain service-level capacity models, App Services, databases, storage, messaging, networking, Key Vault/HSM, and other Azure/PaaS components. - Secure System/Service Acquisition & Region Restrictions:
Ensure external services supporting capacity (e.g., third-party telemetry or scaling tools) conform to required requirements with documented oversight and continuous monitoring - Resilience, DR, and Performance Engineering:
Perform criticality analysis to prioritize capacity for high-critical components; align hardening, monitoring, backup/DR, and buffer policies to criticality tiers. - Metrics & Reporting:
Define and publish capacity KPIs: utilization, saturation, headroom %, runway weeks, scaling efficacy, quota consumption, DR readiness, cost-to-performance efficiency.
Education/
Experience:
- Bachelor's degree in computer science or related discipline; advanced degree preferred.
- 10–12+ years in infrastructure capacity/performance engineering across compute, storage, network, and platform services; financial services experience is a plus.
- Demonstrated experience operating in regulated environments; familiarity with FedRAMP High concepts and evidence requirements.
- Strong data analysis skills; capable of translating telemetry and forecasts into clear decisions and stakeholder communications.
- Experience coordinating cross-functional engineering teams and aligning delivery across multiple platforms and tools.
- Familiarity with Azure services and concepts (e.g., Entra , managed identities, Azure SQL/MI, storage, networking, policies, RBAC) from a PM perspective.
- Azure capacity ecosystem:
Monitor/Log Analytics/Metrics, Advisor, Cost Management, Reservations/Savings Plans, quotas/limits management. - Compute/container scaling: AKS, VMSS, App Service; HPA/VPA, autoscaling policies; performance testing (k6/JMeter); observability (Prometheus/Grafana).
- Storage and database performance: tiering, IOPS/throughput planning, caching, indexing, and connection management.
- Networking and security capacity:
Azure Firewall, NSGs, private endpoints, Bastion; throughput/latency planning and allow-listing discipline. - Cryptography services:
Key Vault, managed HSM; FIPS-validated modules; key lifecycle capacity considerations. - IaC and config management:
Terraform/Bicep/ARM;
Ansible/Chef; integration with gated CI/CD. - Governance:
Azure Policy/Blueprints/Initiatives for configuration baselines and region restrictions; SSP and evidence artifact production.
Preferred:
- Financial services experience
- Building capacity models for multi-region architectures with strict U.S.
-only constraints. - DR planning and execution with validated failover capacity and documented evidence.
- POA&M management and continuous monitoring submissions in a FedRAMP context.
- Collaboration with SREs on SLI/SLOs, error budgets, and reliability patterns.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×