More jobs:
Senior Kubernetes Platform Systems Engineer
Job in
Laurel, Anne Arundel County, Maryland, 20724, USA
Listed on 2026-08-13
Listing for:
Endepth Solutions
Full Time
position Listed on 2026-08-13
Job specializations:
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Administrator
Job Description & How to Apply Below
We’re seeking a Senior Kubernetes Platform Systems Engineer to support our U.S. Government program(s) in Laurel, MD
. The Senior Kubernetes Platform Systems Engineer will be responsible for administering, automating, securing, monitoring and maintaining Kubernetes platform infrastructure across development, staging and operational environments while supporting mission-critical systems.
Responsibilities:
- Administer, maintain, and optimize a highly available Kubernetes platform supporting mission-critical applications across Development, Staging, and Production environments.
- Partner with software developers, systems engineers, and government stakeholders to translate operational requirements into scalable, reliable platform solutions that support mission success.
- Design, build, and maintain hardened container images and deployment templates that ensure consistency, repeatability, and security across enterprise environments.
- Develop, automate, and continuously improve CI/CD pipelines to streamline application delivery, reduce deployment time, and improve overall platform reliability.
- Perform routine vulnerability scanning, system patching, security hardening, and compliance activities to maintain enterprise security standards and accreditation requirements.
- Plan, coordinate, and execute platform upgrades, software releases, and infrastructure enhancements while minimizing operational impact and maintaining service availability.
- Continuously monitor the health, performance, and availability of Kubernetes clusters and supporting infrastructure, proactively identifying and resolving issues before they impact customers.
- Troubleshoot complex infrastructure, application, and platform issues by performing root cause analysis and implementing long-term corrective actions.
- Configure, maintain, and enhance enterprise monitoring, alerting, health check, and logging solutions to improve operational awareness and system performance.
- Establish, manage, and maintain representative development and test environments that accurately mirror production configurations for validation, testing, and release activities.
- Evaluate test coverage, identify operational gaps, and collaborate with engineering teams to improve system quality, reliability, and deployment confidence.
- Coordinate planned maintenance windows and system outages while communicating schedules, risks, and impacts to internal teams, operations personnel, and external mission partners.
- Develop, maintain, and continuously improve technical documentation including Standard Operating Procedures (SOPs), Administrator Guides, User Guides, Knowledge Base articles, and operational runbooks.
- Research emerging technologies, Kubernetes best practices, automation tools, and platform enhancements, providing recommendations that improve operational efficiency, scalability, and security.
- Generate operational metrics, service reports, and performance benchmarks to support leadership visibility, capacity planning, and continuous process improvement.
- Provide Tier II/Tier III operational support by investigating customer support requests, troubleshooting complex technical issues, managing certificate requests and renewals, and ensuring timely issue resolution.
- Provision and onboard new customer projects, configuring Kubernetes resources, platform services, and supporting infrastructure to enable secure, reliable, and scalable deployments.
Qualifications:
- Bachelor's degree in Systems Engineering, Computer Science, Information Systems, or a related discipline is desired. An additional five (5) years of experience may be substituted for the degree.
- Twenty (20) years of Systems Engineering experience in programs and contracts of similar scope
- Demonstrated experience administering and supporting Red Hat Enterprise Linux (RHEL) environments, including system provisioning, storage and network interface management, OS hardening, vulnerability remediation, patch management, and performance optimization.
- Hands-on experience automating infrastructure and application deployments using Ansible or comparable Infrastructure-as-Code (IaC) and configuration management tools to improve operational efficiency…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×