Cloud Operations Engineer
Listed on 2026-08-21
-
IT/Tech
AWS, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Cloud Operations Engineer (AWS or Azure) - Cheltenham (3 days Office Based) About Finova
Finova is the UK’s largest financial services technology provider, supporting one in every five mortgages nationwide. Our agile, cloud-native solutions enable over 60 banks, building societies, specialist lenders, equity release providers and a network of 2,400+ brokers to stay ahead in a competitive market.
Built on open architecture and backed by deep industry expertise, our platform is designed to scale. Each year, we process over £50 billion in loans, manage nearly £50 billion in savings, and support the digital servicing of more than 650,000 UK borrower accounts.
Be part of a team that’s driving innovation, enabling growth and shaping the future of UK lending.
For LendersFinova offers a flexible, modular technology suite designed to help lenders move faster, scale efficiently and deliver standout digital experiences.
Financial Institutions use Finova to launch products faster, process applications up to 50% more efficiently and reduce operational costs — all while staying fully compliant in a fast-moving market.
About the Role:We are looking for an experienced, hands‑on Cloud Operations Engineer to join the team responsible for the availability, security, resilience, and performance of our AWS‑hosted infrastructure.
This is a strictly hands‑on, operational role focused on keeping our production environments stable, secure, and performant. You will operate as a core technical practitioner, balancing day‑to‑day BAU operations and infrastructure changes with continuous improvements to our resilience, automation, monitoring, and security posture. If you thrive in a busy environment, enjoy collaborating with development and security teams, and have a strong bias toward operational discipline, we want to hear from you.
About you:- AWS Expertise:
Proven hands‑on experience managing AWS‑hosted production environments, particularly compute (EC2, ECS), Application Load Balancers (ALB), and networking (VPCs, Security Groups, routing). - Core Infrastructure
Skills:
Strong technical competency in Windows Server administration and foundational networking/load‑balancing principles. - SQL Server
Experience:
Operational support knowledge of Microsoft SQL Server (
Note:
This is not a DBA role, but requires comfort with basic triage and monitoring). - Automation & CI/CD:
Hands‑on experience executing, supporting, and troubleshooting CI/CD pipelines and infrastructure automation issues. - Observability Tooling:
Proficiency with monitoring systems like Datadog, Amazon Cloud Watch, AWS X‑Ray, or similar tools. - Process‑Driven Mindset: A highly disciplined approach to change management, access control, and operational governance, ideally within a regulated environment.
- Infrastructure Management:
Operate and support our AWS production and non‑production environments, focusing heavily on compute services (EC2, ECS), VPC Networking, and storage. - Incident Response & Triage:
Act as the 1st and 2nd line of defense for platform/application issues. Participate in an out‑of‑hours on‑call rotation for P1/P2 incidents and scheduled deployments. - Change & Release Support:
Execute approved infrastructure changes following strict security controls. Support delivery teams with automated deployments and troubleshoot failed or degraded CI/CD releases. - Database Support:
Provide operational support for Microsoft SQL Server (health checks, job monitoring, basic triage, backup verification) and collaborate with DBA teams for deeper optimizations. - Monitoring & Observability:
Own and optimize monitoring tools (Datadog, Cloud Watch, AWS X‑Ray) to catch genuine service‑impacting issues while reducing alert fatigue. - Security & Compliance:
Assess and remediate vulnerabilities, assist with compliance audits (ISO, regulatory assurance), and contribute to incident post‑mortems. - Disaster Recovery:
Support resilience activities, including failover testing, game days, and the validation of recovery runbooks.
- Hybrid working We operate on a hybrid model that is primarily office‑based, requiring three days in the office each week, with the flexibility to…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: