Site Reliability Engineer
Listed on 2026-07-19
-
IT/Tech
SRE/Site Reliability
Best Egg is a market-leading, tech-enabled financial platform helping people build financial confidence through a variety of installment lending solutions and financial health tools. We aim to help customers make smart financial decisions and stay on track, so they can be money confident no matter what life throws at them.
We offer top-tier benefits and growth opportunities in a culture built on our core values:
Put People First – We foster an inclusive, flexible, and fun workplace.
Create Clarity – Open communication drives trust and results.
Get Things Done – We focus, prioritize, and deliver with excellence.
Deliver with Heart – We lead with kindness, humility, and strong teamwork.
Listen to Our Customers – Their needs drive our innovation.
Barclays has entered into an agreement to acquire Best Egg with closing expected to take place in Q2 2026. This acquisition will give us the resources and capital to continue on our mission and drive our strategy forward. With an aligned culture, lower cost of funds, and increased employee growth opportunities across a global brand, we are excited about the future of the Best Egg brand under the Barclays umbrella.
We are looking for collaborative, innovative team players who like to solve problems. There will also be immense opportunities for those willing to dive in. If you're inspired by growth and want to make a real difference, Best Egg is the place for you.
We’re proud to be an equal opportunity employer committed to building a diverse, inclusive team.
The JobAs Site Reliability Engineer, you will serve as a reliability subject matter expert who leads major incident recovery, drives observability and reliability improvements, mentors associate engineers, reduces operational toil, influences technical decisions, and improves resiliency standards. This role requires depth across production systems, telemetry, batch operations, automation, and incident response. You will be expected to guide technical direction for reliability improvements and help teams prevent recurring failures.
Employees joining Best Egg's Information Technology organization can expect a culture centered on Continuous Delivery, Total Quality Management, Knowledge Sharing, Personal and Career Advancement, Empowerment, Innovation, and Collective Ownership.
Duties & Responsibilities- Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
- Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
- Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
- Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
- Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
- Improve JAMS, Go Anywhere, Datadog, xMatters, and service support practices through automation and standards.
- Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.
- Hands‑on familiarity with production support, monitoring, alerting, and incident response practices.
- Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
- Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
- Exposure to AWS or cloud operations and scripting with Python, Power Shell, Bash, or similar tools.
- Clear communication skills during incidents, service requests, and post‑incident follow‑through.
- Strong experience leading production incident recovery and cross‑system reliability investigations.
- Ability to mentor engineers and influence technical decisions without direct authority.
- Datadog, AWS, ITIL, Linux, or automation certification.
- Experience with JAMS, Go Anywhere, xMatters, Service Now/Jira, or CI/CD environments.
- Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
- Familiarity with financial services controls, secure file transfer, or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).