More jobs:
SRE Engineer Java Buffalo, NYC
Job in
Buffalo, Erie County, New York, 14266, USA
Listed on 2026-07-07
Listing for:
Tech Mirrors
Contract
position Listed on 2026-07-07
Job specializations:
-
IT/Tech
SRE/Site Reliability
Job Description & How to Apply Below
SRE Engineer with Java
Contract
Location:
Onsite in Buffalo NYC
- Define, measure, and enforce SLOs, SLAs, and error budgets across all critical services.
- Own incident management end-to-end: detection, triage, escalation, mitigation, and post-mortems.
- Build and maintain runbooks, playbooks, and on‑call rotation schedules.
- Conduct blameless post‑incident reviews and drive systemic improvements to prevent recurrence.
- Design and maintain Java‑based microservices, reliability tooling, and automation frameworks.
- Collaborate with development teams to embed SRE best practices during the SDLC (code reviews, CI/CD gates).
- Implement chaos engineering experiments to proactively surface weaknesses in production systems.
- Develop self‑healing automation scripts and toil‑reduction tooling in Java and Python/Shell.
- Design and own the full observability strategy: structured logging, distributed tracing, and real‑time metrics.
- Build, maintain, and optimize Kibana dashboards, index patterns, and Elasticsearch pipelines for log analytics.
- Configure and manage Dynatrace monitoring — including One Agent deployment, Davis AI problem detection, service flows, and synthetic monitoring.
- Create alerting rules and anomaly detection policies in Dynatrace aligned to SLO thresholds.
- Correlate signals across Kibana and Dynatrace to enable rapid root cause analysis during incidents.
- Manage and improve CI/CD pipelines (Jenkins, Git Hub Actions, or similar) for reliability‑focused deployments.
- Support Kubernetes/Docker‑based infrastructure; optimize resource utilization and autoscaling policies.
- Collaborate with Security and Compliance on hardening production environments and managing vulnerabilities.
- Drive capacity planning, load testing, and performance benchmarking initiatives.
Skills & Qualifications Core Technical Requirements
- 10+ years of hands‑on SRE, Dev Ops, or backend engineering experience in production environments.
- Strong Java development proficiency — Spring Boot, microservices architecture, REST APIs, JVM tuning.
- Demonstrable expertise with Kibana — dashboard creation, KQL queries, index lifecycle management (ILM), Beats/Logstash pipelines.
- Hands‑on Dynatrace experience — One Agent, Smartscape topology, Davis AI, SLO configuration, Synthetic monitoring, DQL/USQL.
- Solid understanding of distributed systems, fault tolerance patterns, and CAP theorem.
- Proficiency in at least one scripting language:
Python, Bash, or Groovy. - Experience with containerization (Docker) and orchestration (Kubernetes / Open Shift).
- Familiarity with cloud platforms — AWS, GCP, or Azure — and IaC tools such as Terraform or Ansible.
- Experience defining and managing SLIs, SLOs, SLAs, and error budgets.
- Proficiency with incident management workflows and tools (Pager Duty, Ops Genie, or similar).
- Knowledge of capacity planning, load testing tools (JMeter, Gatling, k6), and performance optimization.
- Strong understanding of networking fundamentals: DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×