Sr. Platform Engineer
Listed on 2026-08-07
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description
- Design, build, and operate a secure and reliable internal platform for running business applications and machine learning workloads at scale.
- Develop and maintain self-service capabilities (golden paths) for provisioning infrastructure, deploying workloads, and managing the application lifecycle through standardized APIs, templates, and automation.
- Build and evolve reusable platform primitives such as K8s clusters, ingress, service networking, secrets management, policy controls, and identity integration.
- Maintain cloud and platform infrastructure using Infrastructure as Code (e.g., Terraform, Pulumi, Crossplane) in a secure, scalable, and reusable manner, including modular design, versioning, policy guardrails, automated validation, and safe rollout practices.
- Implement and maintain platform-level CI/CD patterns that support multiple teams while enforcing secure and compliant SDLC practices.
- Provide opinionated reference architectures and reusable building blocks (e.g., Helm charts, Argo CD apps, Terraform modules, scaffolding tools) that enable consistent delivery.
- Implement guardrails for security and compliance (policy-as-code, least privilege, workload identity, image provenance, runtime controls) while maintaining developer velocity.
- Own platform observability by establishing monitoring, logging, tracing, alerting, and SLO practices to keep the developer platform stable and measurable.
- Drive platform reliability and operational excellence through incident response, root cause analysis, postmortems, and continuous improvement to reduce toil.
- Support daily operations, monitoring, and security functions of the platform stack, including routine maintenance, access and identity workflows, vulnerability remediation, and operational support for internal platform services.
- Partner closely with application teams and stakeholders to understand friction points, prioritize the platform roadmap, and deliver measurable improvements in developer experience and time-to-production.
- Lead through influence and consensus, providing technical guidance, reviews, and mentorship to peers and junior engineers.
- Own your professional development, continuously learning new technologies and domain context to operate as a subject matter expert (SME).
Qualifications- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Hands‑on experience designing and operating platform services using Infrastructure as Code (Terraform, Pulumi, Crossplane, or CSP-native tooling) with a modular, reusable approach.
- Strong experience with K8s or Open Shift and core platform components, including networking, ingress, service discovery, storage, RBAC, admission control, and multi‑tenancy.
- Experience implementing and standardizing CI/CD for multi‑team environments, ideally with Git Hub Actions and Argo CD, including release strategies and deployment automation.
- Experience with at least one major Cloud Service Provider such as Azure, GCP, or AWS.
- Strong networking fundamentals, including DNS, TLS, load balancing, routing, and private connectivity.
- Strong security foundations across identity, application, data, network, and supply‑chain security, including scanning, signing, SBOM, secrets handling, and policy enforcement.
- Strong software engineering fundamentals and experience building internal tooling and automation, ideally in Go, Python, or Type Script.
- Experience with containers and container tooling, including Docker, registries, and image build pipelines.
- Excellent knowledge of Linux operating systems and operational troubleshooting.
- Proven experience leading incident response, including triage, mitigation, coordination across teams, and driving post‑incident improvements.
- Deep knowledge of observability tooling and practices, such as Prometheus, Grafana, and Datadog.
- Strong problem‑solving and analytical abilities, with excellent communication and cross‑functional collaboration skills.
About Munich ReTogether, we engage with everything we have and are, to help humankind act braver and better.
As the world’s leading reinsurance company with more than…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).