More jobs:
Site Reliability Engineering Lead
Job in
Woking, Surrey County, GU22, England, UK
Listed on 2026-07-21
Listing for:
Exceptional Dental
Full Time
position Listed on 2026-07-21
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Background
At Motion Applied – Connected Intelligence (CI), we create cutting edge wireless connectivity solutions that are transforming customer experience across the transport industry. We create solutions that drive efficiency and cost-effectiveness for customers, delivering unrivalled internet connectivity services.
Purpose of the RoleThis is an opportunity to join our Software Engineering community as a Site Reliability Engineer (SRE) leading initiatives that create and improve the CI Platform as a Service (Paas) offering to the product delivery teams.
- Enabling product development teams to deliver software products on immutable infrastructure.
- Developing and facilitating production and development infrastructure and associated tooling.
- Integrating third party managed services used for delivery and development lifecycle.
- Lead on implementation of the CI cloud governance policy and be a key participant in maintaining and updating the policy in accordance with technology shifts, customer feedback and product development needs.
- Be an open‑minded technologist who values a collaborative work environment and is willing to learn and explore as the fast‑paced industry evolves and changes.
- Key contributor to the Roadmap for the CI Platform.
- Development and maintenance of the CI Platform.
- Collaborate and consult with software engineers and data scientists to help design and implement robust and scalable software products.
- Knowledge sharing and education of team members to enable our Dev Ops culture.
- Proactively monitor costs and security posture of the CI Platform and products running on it.
- Define and implement tooling to continually improve our software development, release and maintenance processes.
- Develop product features with product delivery teams, building upon the CI Platform offering.
- Supporting our live systems, including identifying and implementing improvements to products, tools and processes to improve the on‑going reliability of our solutions.
- Design and implement infrastructure for multi‑region and multi‑tenant products and platform.
- Design and implement monitoring infrastructure for real‑time data streams.
- Design network and access to allow software engineers and data scientists to access services in AWS while keeping the services and data safe and secure.
- Enable and collaborate with teams to automate the entire delivery of a product. From a single web application to the configuration of a cloud account.
- Design and implement security and access management so that users and roles have access only to resources they need within the AWS account and attaching IoT devices.
- Identify root cause of live issues, to help both recover any immediate situation and design/implement improvements for future reliability.
- Working with delivery teams deploying software on the cloud.
- Strategic technical leadership, in particular related to Site Reliability Engineering.
- Evidence of tailored and contextual communication in all directions to realise value as feedback and enquiry.
- Hands on experience in delivering production quality services.
- Experience in supporting live production systems.
- We are looking for an applicant who has: a Bachelor’s degree in computer science, similar technical field of study, or equivalent practical experience.
- 5 + years of hands‑on experience with large cloud providers such as AWS, Azure, GCP, or OCI.
- Leading technical initiatives including roadmap input, cross‑team collaboration and mentoring.
- Proficiency in Infrastructure as Code (IaC) tools such as Terraform.
- 3 + years of hands‑on experience with containerization and orchestration tools like Docker and Kubernetes.
- Experience working in and developing for a Linux environment.
- Excellent programming, debugging, and optimisation skills and at least one strong purpose programming language (Go or Python preferred).
- Solid understanding of Dev Ops practices, CI/CD pipelines and version control e.g. Git.
- Knowledge of observability and monitoring tools e.g. Prometheus, Grafana, ELK etc.
- Experience in troubleshooting incidents and live environments.
- Ability to write and speak in English fluently.
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×