×
Register Here to Apply for Jobs or Post Jobs. X

Senior Reliability Engineer

Job in Malvern, Chester County, Pennsylvania, 19355, USA
Listing for: RIT Solutions
Full Time position
Listed on 2026-07-01
Job specializations:
  • IT/Tech
    Systems Engineer
Job Description & How to Apply Below
Location: Malvern

Senior Reliability Engineer

Hybrid - Malvern, PA

Job Description

As a Senior Reliability Engineer, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.

Core Responsibilities

Team is focused on automating incident response and infrastructure management. While Java and Python receive a stronger emphasis, candidates with solid programming fundamentals in any language and the ability to adapt will be considered. Experience with AWS and event-driven architectures is also valuable. From a technical standpoint, familiarity with observability concepts (e.g., distributed tracing) and tools like Prometheus or Grafana is beneficial, though not mandatory.

More important is an understanding of the underlying principles, such as instrumentation and monitoring strategies.

  • Improve resiliency engineering practices across platforms and applications, including resilient application design patterns, system observability and deployment strategies
  • Incident detection, troubleshooting, and resolution.
  • Develop automation for incident response and infrastructure management
  • Develop and support Open Telemetry integrations for multiple application platforms (browser, ECS, lambda, etc) and languages (JavaScript, Java)
  • Contribute to architectural decisions and support implementation of solutions.
Skills and Qualifications

* Deep knowledge of Java or Java script. Practical experience developing and operating software in distributed systems environments.
* Problem-solving and analytical thinking: ability to diagnose complex issues and propose efficient solutions. Strong debugging and optimization skills for performance and scalability.
* Cloud platforms:
Hands-on experience with AWS services and cloud infrastructure
* System architecture and design: ability to design scalable, secure, and maintainable systems.
* Working knowledge of Python (or similar scripting language).
* Strong knowledge of resiliency engineering techniques for both platforms and applications.
* Experience troubleshooting complex production issues and implementing effective mitigations.
* Familiarity with Open Telemetry specification and core APIs. From a screening perspective, we recommend focusing on:

  • How candidates approach software releases and validate functionality
  • Their understanding of system dependencies and fault tolerance
  • Experience with diagnosing and resolving production issues
  • Their ability to reflect on past incidents and identify improvements
  • Evidence of systems thinking and architectural awareness
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary