×
Register Here to Apply for Jobs or Post Jobs. X

SRE Engineer – Data Analytics

Job in Washington, District of Columbia, 20001, USA
Listing for: Software Technology, Inc.
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below

Site Reliability Engineer (SRE) for Data Analytics

The Site Reliability Engineer (SRE) for Data Analytics is a critical mid-level role focused on applying robust SRE and Dev Ops principles to ensure the stability, performance, and scalability of our client's core data platforms. This role will drive operational excellence by automating CI/CD pipelines and infrastructure (IaC), leveraging advanced observability tools like Dynatrace, and leading incident response for key systems including Databricks, Informatica, and Power BI.

The successful candidate will have 2-4 years of experience, a passion for automation, strong cloud skills (AWS/Azure), and a dedicated focus on maintaining high service reliability (SLIs/SLOs) for a critical Data & Analytics ecosystem in a fast-paced environment in the DC area.

Responsibilities

Deployment & Automation

  • Implement and maintain CI/CD pipelines using tools such as Git Hub Actions, AWS Code Pipeline, and Jenkins.
  • Automate infrastructure provisioning and management using Infrastructure-as-Code (IaC) with Terraform, Cloud Formation, or AWS CDK.
  • Develop robust automation scripts and self-service tooling to minimize toil and enhance operational efficiency.

Capacity, Performance & Cost Optimization

  • Lead and implement operational cost optimization initiatives across cloud infrastructure and data platforms.
  • Configure, maintain, and tune auto-scaling policies and performance thresholds.
  • Develop and execute Resiliency Test plans and provide critical support for Performance testing efforts.

Incident Management & SRE Principles

  • Serve as a production on-call responder, employing strong troubleshooting skills to quickly resolve complex incidents.
  • Proficiently utilize ITIL framework concepts and ITSM tools (e.g., Service Now) for incident and change management.
  • Develop high-quality Root Cause Analysis (RCA) documentation and Knowledge articles to prevent future recurrence.
  • Implement and enforce SRE principles, including the definition and tracking of Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

Observability & Monitoring

  • Manage and leverage advanced observability platforms (Dynatrace preferred, App Dynamics, ELK, etc.).
  • Implement distributed tracing with accurate context propagation across data services and applications.
  • Optimize monitoring queries, and configure actionable dashboards, alerts, and anomaly detectors using tools like Dynatrace and Kibana.

Data Analytics Platform Reliability

  • Ensure the reliability, performance tuning, and access control for Databricks cluster management and data pipelines.
  • Maintain Informatica workflow orchestration, connector reliability, and error handling for critical data flows.
  • Manage Power BI gateway health, access control, and ensure reliable data refresh processes.

Security & Compliance

  • Manage service accounts, access permissions, and roles following the principle of least privilege.
  • Create, deploy, and manage digital certificates and TLS/SSL configurations.
  • Execute effective remediation tasks and respond to security incidents as part of the operational team.
Qualifications

Education & Experience

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 2 to 4 years of hands-on experience in a Dev Ops, Site Reliability Engineering (SRE), or Cloud Infrastructure role.
  • Practical, working experience with major cloud platforms, specifically AWS and Azure.

Technical Skills

  • Mid-level proficiency in Python or other scripting languages (e.g., Bash, Go) for automation tasks.
  • Mid-level proficiency with Configuration Management tools, including Ansible.
  • Strong knowledge of containerization technologies (Docker, Kubernetes/ECS).
  • Solid understanding of Linux systems and networking fundamentals (TCP/IP, DNS, Load Balancing).
  • Working knowledge of relational, cloud-native (e.g., AWS RDS), and No

    SQL database technologies.
  • Direct hands-on experience supporting and maintaining data platforms like Databricks, Informatica, or Power BI is highly desirable.

Professional Attributes

  • Excellent written and verbal communication skills, with a proven ability to document complex systems.
  • Demonstrated ability to work independently, manage shifting priorities, and drive initiatives to completion.
  • Availability for on-call duties and to work outside of standard business hours as required to support a 24/7 production environment.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary