Azure Cloud Engineer
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-09-30
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software), Azure
We're rebuilding vertical software with AI — from the inside.
Embrace owns the software running inside 16% of the Fortune 500, 45+ state agencies, and 450+ banks and credit unions. We acquire entrenched vertical software businesses and rebuild them around AI — products, operations, go-to-market, all of it. Our Venture Lab launches new AI-native products into those same markets, using the distribution and customer relationships our portfolio already owns.
You'll ship AI into production against real workflows, real customers, and a P&L you can see move within a quarter.
We hire people who want scope, speed, and ownership, and who are tired of working on AI that never reaches a customer. If you want to spend the next five years shipping into software that already runs the economy, talk to us.
Job DescriptionThis is a remote position.
Embrace Technology Group is the unified engineering organization across the Embrace portfolio, encompassing our Venture AI Labs. We build and modernize software products across six regulated industry verticals, and we are reshaping how that work gets done — AI-first, forward-deployed, and outcome-driven. Our engineers ship real products to real customers, fast.
We are looking for a Cloud Ops Engineer to operate and continuously improve the reliability, security, scalability, observability, and cost efficiency of our Azure-hosted SaaS products. Our products run across dev, QA, staging, and production environments, with infrastructure managed in Terraform and CI/CD automated through Git Hub Actions.
You will partner with engineering teams to ensure our SaaS platforms and AI-enabled solutions are deployed consistently, monitored effectively, secured properly, and operated reliably in production.
Environment and Technology Context- Microsoft Azure-hosted SaaS products across dev, QA, staging, and production.
- Terraform for infrastructure as code and repeatable provisioning.
- Git Hub Actions for application and infrastructure CI/CD.
- AI-enabled capabilities: STT workloads, LLM integrations, AI service endpoints, quotas, usage and latency monitoring, and cost controls.
- Manage and support Azure infrastructure across dev, QA, staging, and production.
- Maintain operational health of Static Web Apps, Container Apps, PostgreSQL, Storage Accounts, SignalR, Service Bus, Azure AI Foundry, and Azure Arc.
- Ensure resources are provisioned, monitored, maintained, and retired per company standards.
- Support environment setup for new products, customers, and integrations.
- Identify and resolve infrastructure issues affecting performance, reliability, availability, or security.
- Build and maintain Terraform modules and environment configurations.
- Ensure infrastructure changes are version-controlled, peer-reviewed, tested, and approved.
- Manage Terraform state, work spaces, variables, secrets, and deployment workflows.
- Detect and resolve drift between Terraform and deployed Azure resources.
- Standardize naming, tagging, resource group structure, environment isolation, and module patterns.
- Support scalable provisioning of new SaaS environments using reusable templates.
- Build, maintain, and troubleshoot Git Hub Actions workflows for application and infrastructure deployments.
- Support CI/CD pipelines across multiple SaaS products and environments.
- Implement promotion flows from dev to QA to staging to production.
- Add deployment safeguards: environment protection rules, approvals, rollback procedures, validation checks, release gates, and audit trails.
- Manage pipeline secrets, service principals, managed identities, and deployment credentials.
- Improve build and deployment reliability and traceability.
- Operate and monitor Azure AI services, including Azure AI Foundry and Speech-to-Text workloads.
- Support production operations for LLM integrations and AI-enabled product features.
- Monitor AI service availability, latency, quota usage, token consumption, API failures, throttling, and cost.
- Help define operational standards for AI workloads: access control, logging, alerting, failover, usage governance, and provider disruption handling.
- Partner with engineering to troubleshoot AI service issues, integration failures, degraded model responses, or provider-side disruptions.
- Support secure handling of AI secrets, endpoints, keys, managed identities, and private network access.
- Implement and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).