×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Fort Lauderdale, Broward County, Florida, 33322, USA
Listing for: American Express
Full Time position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Site Reliability Engineer I

Sunrise, FL, United States

(Hybrid)

** Job Description*
* Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.

** Responsibilities*
* + Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms

+ Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives

+ Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics

+ Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)

+ Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure

+ Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures

+ Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability

+ Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE‑owned work items

+ Drive reliability improvements through sprint‑based commitments, including automation, operational fixes, and platform enhancements

+ Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business‑critical services

+ Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements

+ Support ITIL‑based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post‑implementation validation

+ Monitor network performance and troubleshoot connectivity, latency, and access‑related issues impacting platform traffic

+ Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates

+ Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies

+ Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability

+ Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions

+ Participate and lead the Development change review and change validation processes

+ Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues

+ Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers

** Qualifications*
* ** Education

Qualifications:

*
* + Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.

+ Bachelor's degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus

+ Strong knowledge of operating systems and application runtimes such as Java and .NET

+ Knowledge of distributed systems and service‑based architectures from an operations and reliability perspective

+ Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana

+ Knowledge of observability practices including logging, monitoring, tracing, and performance…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary