Platform Engineer
Listed on 2026-07-31
-
IT/Tech
SRE/Site Reliability, IT Infrastructure
Who We Are: Inspired by a shared love of fashion, sneakers and the surrounding culture, END. was founded in Newcastle upon Tyne back in 2005. Today, END. serves over 2 million customers worldwide through a seamless blend of online and physical retail. Our industry‑leading stores in Newcastle, Glasgow, Manchester, London, and Milan reflect our commitment to innovation, influence, and inspiration. We offer a carefully curated selection of menswear, womenswear, sneakers, homeware, and lifestyle products to a global audience.
The RoleOwn the reliability, security, performance, and day‑to‑day operational excellence of our Shopify Plus platform and its critical integration ecosystem (Shopify Patchworks D365). You’ll lead observability, incident response, stability and problem management, and the operational governance that keeps trading safe and predictable. The role also takes BAU ownership of vendor‑delivered CI/CD pipelines—maintaining release controls, access and configuration, and continuous improvement—while operating effectively within Shopify’s SaaS constraints.
Here’sa breakdown of what you’ll be doing
- CI/CD maintenance & governance: maintain vendor‑delivered pipelines for themes/apps/integration services; manage access, configurations, release controls, and quality gates; ensure pipelines remain reliable, secure, and well‑documented.
- Release & change governance: own the change calendar, release readiness checks, environment parity checks, and exception processes during freeze/peak windows.
- Shopify Plus security posture: understand Shopify‑managed protections and operate the process to engage Shopify Plus Support when required.
- Observability & monitoring: define and maintain dashboards/alerts across storefront experience (RUM/synthetics where available) and integration services; ensure actionable alerting, clear ownership, and ongoing tuning.
- Service ownership & SLOs: define ownership boundaries, SLIs/SLOs, alert thresholds, and regular reporting for platform and integration services.
- Incident management, stability & problem management: lead triage, stakeholder communications, mitigation, RCA and preventive actions; maintain a stability backlog and drive repeat‑incident themes to closure to improve MTTR and reduce recurrence.
- Integration runtime operations: operational ownership of Shopify APIs/webhooks and Patchworks/D365 flows—rate‑limit safe patterns, idempotent processing, retries/DLQs, replay tooling, reconciliation, and data correctness checks.
- Data reconciliation & control reporting: implement operational controls for key flows (orders, inventory, pricing, fulfilment) including mismatch detection, audit trails, and repeatable replay/backfill procedures.
- D365 environment management: oversee Dev/Sandbox, UAT and Production environments (access, refresh cadence, configuration hygiene, and release readiness).
- D365 upgrades & release support: manage D365 version upgrades by testing in lower environments first; coordinate regression validation and go/no‑go decisioning.
- D365 licence, storage & cost observability: manage user/licence reviews and access governance; monitor and optimise storage/capacity; track and report cost/consumption drivers and optimisation actions.
- Security & compliance controls: secrets management, least privilege, auditability, dependency hygiene, and secure configuration for apps/services; coordinate security reviews and remediation plans.
- Peak readiness & resilience: trading‑event readiness plans, incident playbooks, change risk controls/freeze coordination, validation checklists and business continuity procedures for integration failures.
- Vendor/platform escalation: manage Shopify Plus / Patchworks / D365 support escalations with evidence packs and track actions to closure.
- Access governance & audit: run periodic access reviews across Shopify apps, Patchworks, D365 and monitoring tooling; maintain joiner/mover/leaver processes and audit evidence.
- Platform documentation & guardrails: maintain runbooks, operational standards and "how‑to‑operate" documentation; ensure onboarding for Engineering/Support is clear and current.
- Platform/SRE/operations experience…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: