Principal Cloud Engineer
Listed on 2026-08-15
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Sr. Staff Cloud Engineer
At Bloom Energy, our vision for a world powered by clean, reliable, and affordable energy is more than just a dream—we're making it reality. For over two decades, we've been at the forefront of the global energy transition, pioneering solutions that empower critical industries to thrive in a rapidly digitizing, energy-intensive world. From revolutionizing power for AI-driven data centers to ensuring resilience for hospitals, electric grids, manufacturing facilities, and utilities, our solid oxide fuel cell (SOFC) and solid oxide electrolyzer (SOEC) technologies are redefining what's possible by delivering energy abundance for all.
With more than 30,000 fuel cell modules deployed worldwide, we are the trusted partner for Fortune 100 companies and innovators alike. Our cutting-edge solutions enable unparalleled "time-to-power" capabilities, reliability, and sustainability, ensuring our customers remain ahead in a world where soaring energy demand and intensifying energy scarcity are rapidly becoming the new norm.
We are looking for a Sr. Staff Cloud Engineer to join our Cloud Engineering and Infrastructure team in one of today's most important technology areas. This role will be responsible for designing, building, automating, and operating scalable cloud and hybrid infrastructure supporting enterprise applications, production workloads, AI/data platforms, and modern Dev Ops ecosystems.
This role will report to Cloud Engineering leadership and will be based in San Jose, CA.
Role and Responsibilities- Design, implement, and maintain scalable cloud and hybrid infrastructure supporting enterprise applications, production workloads, AI/data platforms, and CI/CD ecosystems.
- Architect highly available, resilient, secure, and cost-optimized solutions across cloud platforms, with deep focus on AWS.
- Lead adoption and operationalization of Infrastructure-as-Code using Terraform, Cloud Formation, and related automation frameworks.
- Develop cloud platform standards, reusable infrastructure modules, engineering patterns, and best practices to improve consistency, scalability, and operational efficiency.
- Build and mature observability capabilities using metrics, logs, traces, dashboards, and automated alerting to improve operational visibility and incident response.
- Drive adoption of SRE practices including service level objectives, error budgets, incident management, operational reviews, capacity planning, and reliability improvement.
- Lead production readiness reviews, root cause analysis, performance optimization, resiliency planning, and operational risk assessments for critical systems.
- Design, build, and optimize CI/CD pipelines using modern Dev Ops, automation, and Git Ops practices.
- Integrate security, compliance, and operational controls into infrastructure provisioning and deployment workflows.
- Implement automated remediation, rollback, guardrails, and self-healing infrastructure patterns to improve reliability and reduce operational risk.
- Establish and enforce operational standards for monitoring, patching, change management, disaster recovery, and production support.
- Partner closely with Security, Compliance, Infrastructure, Application, and Engineering teams to align cloud platforms with enterprise security, hardening, governance, and regulatory requirements.
- Evaluate emerging cloud, automation, AI/ML infrastructure, and platform engineering capabilities to support Bloom Energy's modernization and scalability goals.
- Mentor cloud, Dev Ops, and infrastructure engineers while promoting engineering excellence, ownership, documentation, and continuous improvement.
- Participate in on-call rotations and incident escalation processes for critical production systems.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field. Master's degree preferred.
- At least 10 years of experience in cloud, infrastructure, platform engineering, or Dev Ops, including 3 or more years in a senior, staff, principal, or technical leadership role.
- Deep hands-on expertise in AWS cloud services, cloud architecture, cloud networking, security, resiliency, and production…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).