Platform Reliability Engineer
Listed on 2026-08-28
-
IT/Tech
Systems Administrator, IT Infrastructure, Unix/Linux, Database Administrator
Overview
Serves as the primary technical owner and operational steward of the organization's database, middleware, and web platform ecosystem. This role carries a dual mandate: ensuring the reliability, security, and continuity of production systems that underpin both internal operations and customer-facing applications today, while actively partnering with IT leadership to rationalize, modernize, and reduce the legacy footprint over time. The ideal candidate brings the operational discipline to maintain complex, interdependent systems with minimal oversight and the forward‑looking mindset to recognize when a platform should be evolved, replaced, or retired rather than simply preserved.
Looking ahead, this role will play a foundational part in building and sustaining the platform environment that powers KMBS's enterprise AI capabilities including self-hosted large language model environments and the data pipelines that connect them to our core business systems.
- Administer and monitor enterprise database, middleware, and web platform systems daily; triage error logs and alert queues as first priority
- Manage and maintain scheduled automation including SQL Agent jobs, Linux cron jobs, and payroll load processing
- Deploy and maintain applications across Web Sphere (WAS), Open Liberty, and Apache/PHP platform environments
- Perform OS-level and application-level patching across Linux (RHEL) and Windows Server environments
- Own end-to-end SSL/TLS certificate lifecycle across web servers, Java keystores, and application servers
- Manage DNS records and domain registrations across multiple registrars
- Support vulnerability management via Tenable and Bitsight reporting and drive remediation of findings within defined SLA windows
- Serve as primary on-call escalation for production outages across web, middleware, and database platform tiers including after-hours response
- Diagnose and resolve failures systematically across the full platform stack (firewall web server app server database)
- Tune and optimize database queries, batch processes, and scheduled jobs for performance and reliability across SQL Server and other administered database platforms
- Maintain and support integrations across the middleware platform stack including ETL pipelines, API connections, and data flows between enterprise systems
- Support the operational platform layer underlying Mule Soft integration deployments including server health, connectivity, and middleware dependencies between enterprise systems
- Evaluate and lead the migration of legacy scheduling and automation toward enterprise platforms, reducing manual operational dependencies over time
- Identify and recommend platform components that are candidates for modernisation, managed service transition, or retirement in partnership with IT leadership
- Introduce and advance configuration-as-code practices to replace manual configuration with repeatable, version‑controlled platform operations
- Support the design, deployment, and ongoing operations of an enterprise AI platform environment including self-hosted large language model infrastructure, vector databases, and RAG pipeline components
- Participate in functional and technical design reviews to ensure platform decisions align with performance, security, and scalability requirements
- Maintain runbooks, architecture diagrams, credential inventories, and operational playbooks for all owned systems, ensuring no critical system or process is dependent on a single individual's undocumented knowledge
- Document scheduled jobs, system dependencies, and non-standard configurations
Core Requirements -
- Proven production experience administering Microsoft SQL Server environments including SSIS, SQL Agent scheduling, linked servers, and T‑SQL development (stored procedures, functions, views, indexes)
- Demonstrated ability to tune and optimize SQL queries, SSIS packages, and batch processes for performance and reliability
- Hands‑on Linux/RHEL systems administration including shell scripting, cron job management, patching, and Apache/PHP stack operations
- End-to-end SSL/TLS certificate lifecycle management across web servers and Java keystores
- DNS…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).