More jobs:
Production Support Engineer Site Reliability Engineer SRE
Job in
New York City, Richmond County, New York, USA
Listed on 2026-08-16
Listing for:
Futran Tech Solutions Pvt. Ltd.
Full Time
position Listed on 2026-08-16
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer
Job Description & How to Apply Below
Production Support Engineer Site Reliability Engineer SRE
Location:
New York
Duration: FTE
Experience:
5 Years
Domain:
Banking Financial Services Capital Markets Fin Tech Digital Assets
Role Overview:
We are seeking a Production Support Engineer SRE with experience in cloud native platforms Kubernetes application support observability automation and platform reliability. The ideal candidate will support mission critical applications ensuring high availability performance scalability and operational excellence across cloud environments.
Key Responsibilities:
- Provide L2L3 production support and incident management for business critical applications
- Perform troubleshooting root cause analysis RCA and implement preventive solutions
- Support and optimize cloud environments across Azure and AWS
- Manage and troubleshoot Kubernetes AKS Docker VMs and distributed systems
- Build and maintain monitoring solutions using Splunk Prometheus Grafana Azure Monitor and Cloud Watch
- Develop automation scripts using Python Bash or Power Shell to improve operational efficiency
- Support CICD pipelines and release management processes
- Troubleshoot application API database and infrastructure issues
- Ensure compliance through IAM management certificate renewals vulnerability remediation and security best practices
- Contribute to AIdriven operational automation and intelligent monitoring initiatives
Required Skills:
- Cloud Azure AKS Azure Monitor Cloud Watch Terraform
- Containers Kubernetes Docker Helm
- Monitoring Splunk Prometheus Grafana New Relic
- Dev Ops Jenkins Azure Dev Ops Git Hub Actions Git Lab CICD
- ProgrammingScripting Python Bash Power Shell
- Databases PostgreSQL SQL Server MySQL
- Service Management Service Now Jira ITIL Processes
- Operating Systems Linux RHEL Ubuntu
Preferred
Skills:
- Domain experience MLOps MLflow AIOps and AILLMbased operational automation experience
- Infrastructure as Code IaC using Terraform
- Experience in regulated financial services environments
Key
Competencies:
- Strong troubleshooting and analytical skills
- Incident management and RCA expertise
- Reliability engineering mindset
- Automationfirst approach
- Stakeholder management and communication
- Cross functional collaboration and customer focus
Skills Mandatory
Skills:
CI/CD Architecture, Python, SITE
24X7
Good to Have
Skills:
LLMOps
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×