Engineer - Contingent
Listed on 2026-07-01
-
IT/Tech
SRE/Site Reliability
Principal Engineer – Platform Engineering and Production Support
Job Description:
Principal Engineer – Platform Engineering and Production Support
This role supports a critical Platform Engineering team responsible for stabilizing, scaling, and operating applications as they move closer to production release.
The team plays a key role in post‑deployment, ensuring reliability, performance, and operational excellence across a portfolio of applications.
This is not traditional infrastructure support—it is application‑focused production engineering, requiring deep technical expertise, proactive issue prevention, and strong ownership of application health in cloud environments.
Role SummaryClient seeks a Principal Engineer to backfill a key contractor position within the client Platform Engineering team. The candidate must be Day 1 ready, capable of operating in fast‑paced, production‑critical environments, and able to seamlessly balance multiple priorities.
The ideal candidate is a strong Dev Ops and Site Reliability Engineering (SRE) professional with hands‑on expertise in observability, incident management, and cloud platforms (Open Shift). They will play a leading role in supporting production systems, preventing outages, and improving system reliability through automation and intelligent monitoring.
Key Responsibilities- Lead production support efforts across a portfolio of 20+ applications, ensuring stability, performance, and rapid issue resolution.
- Design and build advanced monitoring, alerting, and observability dashboards using tools such as Splunk, Grafana, App Dynamics, and Prometheus.
- Proactively identify risks through gap analysis, anomaly detection, and predictive alerting, preventing production incidents before they occur.
- Troubleshoot complex production issues across distributed microservices environments, reducing MTTR through deep technical expertise.
- Drive adoption of modern SRE practices, including automation, AIOps, and intelligent monitoring solutions.
- Support applications running on Open Shift and cloud‑native platforms, focusing on reliability and scalability.
- Collaborate closely with development teams during release cycles, providing production‑readiness guidance and operational support.
- Participate in 24x7 on‑call rotation, demonstrating urgency and ownership during incidents.
- Mentor and guide engineers, helping elevate team capabilities in SRE, Dev Ops, and platform engineering practices.
- Act as a trusted technical leader, switching priorities quickly and managing competing demands in a high‑pressure environment.
- A genuine, hands‑on engineer who can operate across multiple roles (SRE, Dev Ops, Production Support).
- Strong ability to shift priorities quickly and respond with urgency in critical situations.
- Deep understanding of application support in cloud environments, especially Open Shift.
- Experience in the financial services industry strongly preferred.
- Prior development experience is a plus, particularly in Java‑based ecosystems.
- 10+ years of platform and production support.
- 5 years of Red Hat Linux, Open Shift, Kubernetes, Java, microservices, Spring Boot, Python experience.
- 5 years of observability dashboard creation experience – Grafana, Splunk, SPLOC, App Dynamics.
- 5 years of observability alerts and incident handling – AIOps, Service Now, Big Panda, etc.
- 4 years of React.js, Apache, Kafka, relational database experience.
- 4 years of distributed systems, microservices architectures, and cloud‑native platforms experience.
- 7+ years of engineering experience, or equivalent experience evidenced through prior work, consulting, training, military, or education.
Irving, TX;
Charlotte, NC;
Minneapolis, MN.
- Standard expectation: 8‑hour workday.
- Monthly on‑call rotation (supported by offshore team; extended late hours are unlikely).
- Typical hours fall within 8:00 AM to 8:00 PM, with remaining time potentially on‑call.
- Generally, not expected to exceed 40 hours/week.
- Start date:
ASAP.
12‑month contract with potential for extension and/or conversion (not guaranteed at this time).
CompensationPay range: $85‑$95 per hour. Specific compensation will be determined by scope, complexity, location, and candidate qualifications.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).