Senior Site Reliability Engineer
Listed on 2026-09-05
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Senior Site Reliability Engineer (SRE)-GCP/Kubernetes
Aboutthe RoleWeareseekinganexperiencedandhighlymotivated Senior Site Reliability Engineer (SRE) tojoinoursmall,agileengineeringteam.
Thisroleofferstheuniqueopportunitytodrivethereliability,scalability,andperformanceofourcoreplatformwithahighdegreeofautonomyandownership.
Thesuccessfulcandidatewillsplittheirtimebetweenprovidingexpertoperationalsupportforourcriticalsystemsandleadingexcitingnewinfrastructureprojects.
Ourmindsetistogettherightpersonnotthepersonwiththeskillsthatmatchourstack.
Itisimportanttobeabletoforeseeproblemsbeforetheyshowupandcreatesolutionsthatmitigatethem.
Ifyouenjoyachallengingenvironment,implementing“infrastructureascode”principles,anddirectlyseeingtheimpactofyourwork,thisistheplaceforyou.
- Design&Build:
Architect,deploy,andmaintainhighlyscalableandreliableinfrastructureon
Google Cloud Platform (GCP) using
Kubernetesand
Infrastructure-as-Codetools. - Automation:
Championautomationacrosstheentiresoftwaredevelopmentlifecycle(SDLC),utilizing
IaC,Pythonand Bashtoreducetoilandimproveoperationalefficiency. - Infrastructure-as-Code(IaC):
Ownandevolveourdeclarativeinfrastructureusing Terraformforcloudresourcesand Helmfor Kubernetesapplicationdeployment . - Monitoring&Observability:
Implementandmanagerobustmonitoring,alerting,andloggingsolutionstoensureclearsystemvisibilityandproactiveissueidentification. - Reliability&Performance:
Define,measure,and enforce
Service Level Objectives (SLOs) and Service Level Indicators (SLIs).Participateinon-call rotation(if applicable) andleadpost-incidentreviewstodrivecontinuousimprovement. - Collaboration:
Workcloselywithsoftwaredevelopmentteamstoprovideexpertguidanceondeploymentstrategies,scalabilityconcerns,andcloud-nativebestpractices. - Ownership:
Takefullownershipofprojectsfrominceptionthroughtoproductionoperation,includingdocumentationandknowledgetransfer.
- Cloud Platform:4+yearsofhands-onexperiencewith
Google Cloud Platform (GCP)(orsimilarcloudinfrastructure). - Container Orchestration:
Expert-levelproficiencyinmanaging,scaling,andtroubleshootingproduction
Kubernetesenvironments. - Infrastructure-as-Code:
Deepexpertisein Terraformformanagingcloudand Kubernetesresources . - Deployment:
Strongexperiencewith Helmforpackaginganddeployingapplicationson Kubernetes . - Scripting/Programming:
Proficientinatleastonemajorprogramminglanguage,preferably
Python,forautomationandtooldevelopment.
- CI/CD:
ExperiencesettingupandmaintainingmodernCI/CD pipelines. - Observability:
Practicalexperienceimplementingandmanagingmonitoringandloggingtools. - Networking:
SolidunderstandingofTCP/IP,load balancing,DNS,andcloud-nativenetworkingwithin
Kubernetes. - Operating Systems:
Strongcommand-lineskillsandexperiencewith
Linuxsystems.
- High Ownership:
Demonstratedabilitytoownaproblemend-to-end,frominvestigationtoresolutionandpreventativemeasures. - Small Team Mentality :
Happytobeageneralistandswitchcontextquicklybetweensupporttickets,operationaltoilreduction,andlong-termprojectwork. - Adaptability:
Aproventrackrecordofrapidlylearningandapplyingnewtechnologiesandtools.
Equivalentexperiencewithotherclouds(AWS/Azure) orsimilartoolsishighlyvalued. - Communication:
Excellentverbalandwrittencommunicationskillsfordocumentationandinteractingwithnon-technical stakeholders.
- Familiarity with Service Meshtechnologies (e.g.,Istio).
- Experienceinsecuritybestpracticeswithincloudandcontainerenvironments(e.g.,hardening,secrets management).
- CertificationsinGCPorKubernetes(e.g.,CKAD,CKA,Professional Cloud Dev Ops Engineer ).
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: