×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer -Jersey , NJ & Dallas, TX

Job in Dallas, Dallas County, Texas, 75215, USA
Listing for: Stradit LLC
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Role: Site Reliability Engineer

Experience: Min 10 Years

Locations: Jersey City, NJ & Dallas, TX

Work mode: Hybrid

Employment: W2

The Application Support Engineering role advances

Site Reliability Engineering (SRE) practicesforapplications running inproduction. Therole scope includes

Application Support Engineers, SDETs, Software Engineers, and SREsfocused on improving reliability, observability, recovery, automation, and operational readiness.

This role brings production supportexpertiseand engineering discipline earlier in the lifecycle to influence design,validate reliability requirements, reduce production risk, and drive measurable operational improvement.

Your Primary Responsibilities
  • Participate in design reviews, sprint zero, and delivery planning todefine andvalidatereliabilityrequirements,including resiliency, observability, fault tolerance,performance, scalability,holiday and special-day processing,and disaster recovery.
  • Collaborate with Major Release Management to ensure each release meets SRE standards for observability, resiliency,and reliability requirements, support readiness, and knowledge base coverage.
  • Define and improve monitoring,observability, dashboards, telemetry coverage, and alertstrategyto strengthen outage detection, reduce noise, improve signal quality, andaccelerateincident response.
  • Assistin major incident response and root cause analysis byidentifyingobservability gaps, improvingtelemetryandknowledge articles, and driving actions that reduce repeat incidents.
  • Drive automation, intelligent tooling, and AI-assisted remediation to reduce manual toil, improve consistency, accelerate recovery, andscaleoperationalsupport.
  • Serve as the operational readiness authority before production releases byvalidatingreliability requirements,assessing support readiness, surfacingproductionrisks, and confirming release supportability.
  • Lead capacity, performance, workload trend, and resiliencyanalysisto ensure applications scale reliably under normal, peak, and stress conditions.
  • Establish and track reliability metrics such as availability, incident volume,MTTx, alert quality, automation coverage,reliability requirement compliance, change failure rate, and repeat incident reduction.
  • Participate in application reliability governance and service reviews by presenting incident trends,compliance metrics, operational risks, improvement actions, and readiness gaps.
  • Prepare executive reporting on reliability posture, release readiness, observability maturity, alert quality, incident trends, automation progress,risks, and improvement outcomes.
  • PromoteSREpracticesthrough mentoring,standards adoption, best-practice sharing, and approved AI toolsthat improve knowledge, observability, performance, security, and maintainability.
Qualifications
  • Minimum of10+years of related technical experience across application support engineering, software engineering,site reliability engineering,production support,or application operations.
  • Bachelor’s degree preferred or equivalent practical experience.
  • Experience supporting business-critical applications in production environments.
  • SRE,observability,automation,or ITIL certifications are a plus.
Talents Needed for Success
  • Proven experience in one or more in-scope roles including

    Application Support Engineer, SDET, Software Engineer, or SRE, with responsibility for improvingreliabilitypractices,application validation, observability coverage, automation frameworks, and reliability standards.
  • Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation.
  • Programming or scriptingproficiencyin one or more languages such as Python, Java, Go, Power Shell, orsimilar for automation,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary