Job Description & How to Apply Below
About the Role
We are looking for a Senior Site Reliability Engineer to join our existing SRE team and build tools and automation that improve production reliability and day-to-day operations. As a senior individual contributor, you will turn initial operational needs into clear requirements, practical system designs, working proofs of concept, and maintainable production solutions.
You will work closely with SRE, Engineering, and Production Support teams, quickly learn our products and support model, and contribute to daily SRE activities. Building new tools and automation is the primary focus of this role, alongside participation in the weekly support rotation.
LocationHybrid | Bandung, Indonesia (at least 2-3 days per week in the office)
What You’ll be Doing Tooling and Automation- Design, build, and maintain internal tools, services, and scripts using Python, Go, and shell scripting to reduce repetitive work and improve operational efficiency.
- Gather and clarify raw requirements from SRE and Engineering teams; define the problem, scope, acceptance criteria, and expected operational outcomes.
- Create system designs covering components, integrations, data flows, security, scalability, and failure handling; explain design choices and trade-offs.
- Build and demonstrate proofs of concept, incorporate feedback, and own delivery through testing, rollout, documentation, and ongoing maintenance.
- Develop and improve CI/CD pipelines for reliable builds, automated testing, deployment, and rollback; automate infrastructure and configuration workflows.
- Quickly learn product architecture, service dependencies, infrastructure, deployment processes, and the production support model.
- Work alongside the existing SRE team on daily server and service operations, including monitoring, troubleshooting, maintenance, and controlled changes.
- Participate in the weekly rota for SRE operations and on-call support, including out-of-hours incident response when scheduled; follow escalation and handover procedures.
- Investigate production incidents with Engineering and Production Support, contribute to root cause analysis, and turn recurring issues into lasting automation improvements.
- Work closely with SRE, Engineering, and Production Support to prioritize operational needs and deliver tools that fit existing workflows.
- Write tested, maintainable code, participate in code and design reviews, and apply secure development practices.
- Maintain system designs, runbooks, operating procedures, and user documentation; demonstrate solutions and support their adoption across teams.
- Improve monitoring, alerting, and service performance, and evaluate delivered automation by its reliability and reduction in manual effort.
- At least 5 years of relevant experience in Site Reliability Engineering, Dev Ops, platform engineering, systems engineering, or production software engineering with operational responsibilities.
- Strong hands-on programming ability in Python or Go, practical shell scripting skills such as Bash, and willingness to work across the team’s languages and tooling.
- Demonstrated experience building and maintaining operational tools or automation beyond one-off scripts, including APIs, error handling, testing, and version control.
- Hands-on experience designing and maintaining CI/CD automation using tools such as Jenkins, Bitbucket Pipelines, or equivalent platforms.
- Ability to independently clarify incomplete requirements, design a suitable solution, and build a proof of concept that demonstrates its value and feasibility.
- Experience with cloud infrastructure such as AWS or GCP, Linux server administration, networking fundamentals, and troubleshooting…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×