×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer​/Production Support

Job in New York, New York County, New York, 10261, USA
Listing for: Smart IT Frame LLC
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Administrator
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below
Location: New York

Job Title:

Production Support/Site Reliability Engineer

Location:

New York City, NY Key Responsibilities

  • Monitor production applications, infrastructure, services, and databases to ensure high availability and performance.
  • Provide L2/L3 production support and troubleshoot application, infrastructure, and deployment issues.
  • Analyze alerts, logs, metrics, and traces to identify and resolve production incidents.
  • Participate in 24x7 on‑call / shift‑based support, including incident response and escalation when required.
  • Perform Root Cause Analysis (RCA) for major incidents and implement preventive actions.
  • Manage incidents, problems, and service requests according to defined SLAs.
  • Work with development teams to troubleshoot application defects and production issues.
  • Support application deployments, releases, rollbacks, and production changes.
  • Develop automation/scripts to eliminate repetitive manual operational activities.
  • Implement and maintain monitoring, alerting, dashboards, and observability solutions.
  • Identify performance, capacity, and reliability bottlenecks and recommend improvements.
  • Participate in disaster recovery, backup, failover, and business continuity activities.
  • Maintain operational documentation, runbooks, SOPs, and troubleshooting guides.
  • Continuously improve system reliability, scalability, resilience, and operational efficiency.
Required Technical Skills
  • Strong experience in Linux/Unix administration and troubleshooting.
  • Good knowledge of AWS / Azure / GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, or App Dynamics.
  • Hands‑on experience with Docker and Kubernetes.
  • Good understanding of CI/CD pipelines using Jenkins, Git Lab CI, Git Hub Actions, Azure Dev Ops, or similar tools.
  • Scripting/programming experience in Python, Shell/Bash, or Power Shell.
  • Strong knowledge of Git and source‑control practices.
  • Experience troubleshooting REST APIs, microservices, web applications, and distributed systems.
  • Good understanding of networking concepts such as DNS, HTTP/HTTPS, TCP/IP, load balancers, and firewalls.
  • Working knowledge of SQL and databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
  • Experience with incident management and ticketing tools such as Service Now, Jira, or Pager Duty.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary