×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Buffalo, Erie County, New York, 14266, USA
Listing for: BCforward
Full Time position
Listed on 2026-07-28
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 100 - 110 USD Hourly USD 100.00 110.00 HOUR
Job Description & How to Apply Below

Job Title:

Lead Site Reliability Engineer Duration:
Temp - 12 months Pay Range: $100/hr $110/hr (W2) Job  About BCforward

BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity.

Job Description

We are seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex reliability initiatives.

Responsibilities:
  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure aligned to enterprise standards and SRE best practices.
  • Lead initiatives to improve reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor SLOs, SLIs, and error budgets for critical services.
  • Develop observability strategies using Dynatrace, Open Telemetry, distributed tracing, metrics, logs, dashboards, and alerting.
  • Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to identify performance bottlenecks, reliability risks, and capacity constraints proactively.
  • Lead incident response for high-severity events and coordinate cross-functional restoration and communications.
  • Perform and facilitate RCAs with corrective and preventive actions tracked to completion.
  • Automate repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to embed reliability and observability across the SDLC.
  • Design, develop, and execute automated regression testing to validate stability and performance after changes.
  • Review test coverage and reliability validation to ensure comprehensive risk mitigation.
  • Create, maintain, and improve Terraform-based IaC for provisioning, configuration, and standardization.
  • Support and optimize Microsoft Azure environments, including App Services, resource management, scaling, and deployment automation.
  • Use Azure Monitor, Application Insights, and Log Analytics to improve visibility and reliability.
  • Drive performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness.
  • Establish operational readiness standards and enforce requirements before production deployments.
  • Review architectures and recommend improvements for resiliency, efficiency, and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization across production environments.
  • Develop and maintain runbooks, incident playbooks, knowledge articles, and SOPs.
  • Partner with engineering, infrastructure, cybersecurity, architecture, and support teams on cross-functional improvements.
  • Communicate system health, reliability trends, risks, and remediation to technical and business stakeholders.
  • Present initiatives, metrics, and recommendations at reviews, forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational practices.
  • Adhere to risk and regulatory standards and identify issues requiring escalation.
  • Promote a culture of belonging consistent with company values and maintain internal control standards.
Required Skills &

Qualifications:
  • Associate’s degree with 7+ years in SRE, Cloud, Systems, Infrastructure Engineering, Dev Ops, or Application Support. Bachelor’s degree with 5+ years. Or 9+ years combined education and experience with 5+ years in a technology engineering role.
  • Hands‑on observability and monitoring…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary