×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Kansas City, Jackson County, Missouri, 64101, USA
Listing for: Ad Astra
Full Time position
Listed on 2026-07-31
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 110000 - 150000 USD Yearly USD 110000.00 150000.00 YEAR
Job Description & How to Apply Below

Competitive Compensation & Benefits Package
* 401(k) with Profit Sharing
* Flexible Time Off
* Office Dog!!

ABOUT US

By combining our unparalleled domain expertise with leading-edge technology, Ad Astra is helping higher education in its mission to advance timely student completions. We are building a cloud-based software platform that will provide the foundation for our next generation of industry-leading solutions and analytics. Simply put, we're helping students graduate faster.

OUR CORE VALUES
  • We recognize talent. We recognize and appreciate the unique God-given talents that our people bring to Ad Astra. Aligning these individual gifts with our work sets team members up to succeed.
  • We’re unpretentious. There’s no room for ego. We admit our imperfections and have the humility to know what we don’t know.
  • We’re passionate. We aren’t satisfied with the status quo. We’re on a mission together to protect the value of degree completion and to transform the higher education industry.
  • We’re pioneering. We’re pioneering and aren’t afraid of failing—in fact, we celebrate it. We love it when our people boldly experiment with innovative solutions.
  • We love fun. The health of our relationships is strengthened by working with people who stretch our thinking—and by enjoying the lighter side of life together. We don’t take ourselves too seriously, but we do take fun seriously.
  • We have grit. Beyond talent and intelligence, our people have stick-to-itiveness. We push through challenges to make goals a reality.
POSITION SUMMARY

The Site Reliability Engineer (SRE) will ensure the performance, reliability, and scalability of our systems as we continue to grow. This role bridges the gap between software development and operations, applying software engineering principles to automate, optimize, and enhance the reliability of our infrastructure and production systems. Your role includes identifying recurring failure patterns, implementing automated solutions, and continuously improving platform performance.

Leveraging your intellectual curiosity and expertise in operations and development, you will also play a pivotal role in monitoring security and reliability threats, while actively advocating effective solutions.

This role spans a genuinely wide range of work, from deep automation and greenfield infrastructure projects to legacy system support and cross-team collaboration. You'll thrive here if you enjoy variety and can move between priorities without missing a beat, and if you're motivated by helping shape reliability practices as we grow toward higher availability targets.

ESSENTIAL FUNCTIONS/CORE RESPONSIBILITIES
  • Write automation and production code to improve system reliability and performance.
  • Design, build, and maintain highly available, scalable systems across cloud environments (e.g., AWS, Azure, or GCP).
  • Own reliability and scalability considerations unique to our multi-tenant SaaS platform, including tenant isolation and blast radius containment.
  • Maintain and extend logging, monitoring, and alerting systems to enhance observability and proactive incident response.
  • Bridge development and operations by automating workflows, deployments, and infrastructure provisioning.
  • Proactively monitor and respond to alerts and incidents, ensuring system uptime and performance.
  • Support security and compliance initiatives.
  • Eliminate manual toil across legacy systems (currently being migrated away from), client onboarding, and data integration, with the autonomy to build lasting automation.
  • Collaborate with engineering, product, and operations teams to capacity plan, drive cloud cost reduction, and enhance the overall reliability and efficiency of our products
  • Support production systems, including participation in on-call rotations and performing limited after-hours maintenance.
  • Lead and contribute to post-incident reviews, driving root cause analysis and long-term solutions.
  • Document reliability patterns, runbooks, and learnings to build operational maturity
  • Other duties as assigned
POSITION REQUIREMENTS
  • Bachelor’s degree in Computer Science, Engineering, or related field preferred; equivalent experience in supporting distributed software…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary