×
Register Here to Apply for Jobs or Post Jobs. X

Senior Azure Cloud Infrastructure Engineer; Healthcare AI Platform

Job in Washington, District of Columbia, 20022, USA
Listing for: CIVIE
Full Time position
Listed on 2026-09-30
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Azure, Disaster Recovery IT, Cybersecurity
Salary/Wage Range or Industry Benchmark: 120000 - 190000 USD Yearly USD 120000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: Senior Azure Cloud Infrastructure Engineer (Healthcare AI Platform)

Overview

We’relooking fora Senior Azure Cloud Infrastructure Engineer to design, build, andoperatea highly resilient, secure, and cost-efficient cloud platform supporting advanced AI workloads in a healthcare environment.

This roleis responsible for mission-critical infrastructure powering our proprietary foundational AI model, including GPU-based compute, while meeting strict requirements for compliance, data protection, and high availability. You will play a key role in ensuring our systems are fault-tolerant, auditable, and continuouslyoptimizedfor both performance and cost.

What You’ll Do
  • Architect and manage highly available, fault-tolerant systems on Microsoft Azure with multi-region redundancy and disaster recovery
  • Design infrastructure with strict adherence to healthcare compliance standards (e.g., HIPAA, HITRUST, SOC
    2)
  • Provision andoptimizeGPU-based environments for AI/ML workloads, including large-scale model training and inference
  • Build secure, zero-trust architectures (private networking, encryption, identity isolation, least privilege access)
  • Implement backup, failover, and business continuity strategies with clearly defined RTO/RPO targets
  • Continuously reduce infrastructure costs through intelligent scaling, reserved capacity, spot instances, and workload optimization
  • Develop Infrastructure as Code (Terraform, Bicep, ARM) for repeatable, auditable deployments
  • Partner with AI/ML teams to product ionize and scale foundational models reliably
  • Establish observability across systems (logging, monitoring, alerting) with proactive incident response
  • Conduct architecture reviews, risk assessments, and security audits
Required Experience
  • 5–8+ years of hands-on experience with Microsoft Azure cloud infrastructure
  • Proven experience designing high-availability and disaster recovery systems in regulated environments
  • Strong background in healthcare or other compliance-heavy industries
  • Deep expertise in:
    • Azure Virtual Machines, VM Scale Sets, and GPU compute
    • Azure networking (VNets, Private Link, Express Route, firewalls)
    • Storage solutions (Blob, Files, managed disks with redundancy options)
  • Experience implementing compliance frameworks such as HIPAA or SOC 2
  • Strong knowledge of identity and access control (RBAC, Azure AD, managed identities)
  • Experience with Kubernetes (AKS) and containerized workloads
  • Proficiency in scripting (Python, Bash, Power Shell)
Preferred Qualifications
  • Experience with Azure AI ecosystem (Azure Machine Learning, Azure AI Foundry, Cognitive Services)
  • Familiarity with distributed training, model parallelism, and GPU orchestration
  • Experience implementingMLOpspipelines in regulated environments
  • Azure certifications (Solutions Architect Expert, Security Engineer Associate, Dev Ops Engineer Expert)
  • Experience with zero-downtime deployments and blue/green or canary strategies
Infrastructure Expectations
  • Multi-region architecture with automated failover
  • End-to-end encryption (data at rest and in transit)
  • Segmented environments (dev/staging/prod) with strict isolation
  • Real-time monitoring and alerting with defined SLAs
  • Automated backup and recovery with regular testing
  • Cost visibility and governance across all resources
What Success Looks Like
  • Near-zero downtime systems with tested failover capabilities
  • Full compliance readiness with audit trails and documentation
  • Efficient GPUutilizationsupporting AI workloads at scale
  • Measurable reduction in cloudspendwithout compromising reliability or security
  • Seamless collaboration between infrastructure and AI teams
Why This Role Matters

You will be building the backbone of the next-generation healthcare AI platform - where reliability, security, and performance directly impact real-world outcomes. This is not just infrastructure; it is critical systems engineering at the…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary