DevOps Engineer II
Listed on 2026-07-14
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Our core values — Stewardship, Character, Collaborate, Learn, Disrupt — are the lens through which we evaluate every business decision. As a dynamic, growing company that offers extremely competitive compensation and benefits, our employees are our most valued assets and the foundation of Expand's performance among our E&P competitors.
We seek applicants from all backgrounds to ensure we get the best, most creative talent on our team. We realize that, historically, underrepresented groups feel the need to be 100% qualified in order to apply. If you meet any combination of our requirements, we encourage you to apply. We strive to hire people from a wide variety of backgrounds, not just because it’s the right thing to do, but because it makes our company stronger.
Job SummaryThe position is a technical role responsible for ensuring the reliability, scalability, and operational excellence of enterprise platforms and applications. This role partners closely with developers, data engineers, Salesforce administrators, and platform owners to design resilient systems, improve deployment practices, and automate operations across cloud, integration, and low-code environments.
This position will contribute to Dev Ops and cloud engineering practices by supporting automation efforts, enhancing system observability, and assisting with incident response and continuous improvement initiatives. This role combines hands‑on technical skills with a focus on system performance, reliability, and operational efficiency across a diverse enterprise platform landscape.
Job Duties & Responsibilities- Partner with application, data, Salesforce, and platform engineering teams to design and support reliable, scalable, and secure systems
- Maintain deep expertise on supported platforms and services, including Salesforce, Mule Soft, Azure, M365, Informatica, and Automic, to proactively identify reliability risks and improvement opportunities
- Design and implement infrastructure and platform automation using infrastructure-as-code and configuration management practices
- Build, maintain, and optimize CI/CD pipelines across cloud and SaaS platforms
- Apply and promote Dev Ops and SRE best practices, including automation standards, deployment strategies, and cloud architecture patterns
- Participate in incident response efforts, including troubleshooting, root cause analysis, and post-incident reviews to drive systemic improvements
- Develop and maintain modern observability solutions using tools such as Splunk for log aggregation, correlation, and alerting across distributed and integrated systems
- Assist with the setup, administration, and operational support of AI development platforms and tools (e.g., Microsoft Foundry, Git Hub Copilot, Claude Code), including environment provisioning, CI/CD enablement, and support for AI-assisted and low-code development workflows
- Identify and eliminate manual operational processes through scripting, automation, and platform engineering solutions
- Collaborate with engineering teams to improve application performance, fault tolerance, and resiliency across integration middleware (Mule Soft) and data pipeline (Informatica) layers
- Contribute to capacity planning, cost optimization, and system scalability initiatives
- Evaluate and recommend tools, technologies, and approaches to improve platform reliability and developer productivity
- Document architecture, operational processes, and reliability standards
- Strong experience with cloud platforms, primarily Microsoft Azure, with working knowledge of AWS
- Experience deploying and supporting infrastructure using infrastructure-as-code tools (e.g., Terraform, ARM, Bicep)
- Experience supporting and administering enterprise SaaS and PaaS platforms, including configuration, environment management, and operational support
- Strong foundational knowledge of operating systems (Linux and Windows Server)
- Understanding of distributed systems, microservices architectures, and resiliency patterns
- Working knowledge of identity and access management concepts, including integration with identity providers (e.g., Azure…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).