×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Alpharetta, Fulton County, Georgia, 30022, USA
Listing for: Incident IQ
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below

Site Reliability Engineer

Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room.

We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda.

This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

This role is hands-on from day one. Your initial focus will be:

  • SLI/SLO Definition & Grafana Implementation:
    Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Incident Management:
    Stand up our incident management practice (tooling such as Pager Duty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Observability Stack Ownership:
    Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Team Enablement:
    Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Toil Reduction:
    Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Chaos & Performance Engineering:
    Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:

  • Education & Systems Foundations:
    Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
  • AI-Accelerated Execution (core requirement):
    You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it. This is not a bonus skill here; it's how we expect this role to operate.
  • Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, Dev Ops, or production engineering role. We care far more about what you've actually shipped than the number itself.
  • SLI/SLO Methodology:
    Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
  • Observability Tooling:
    Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with Open Telemetry and a…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary