More jobs:
Senior Software Engineer; SRE
Job in
Miami, Miami-Dade County, Florida, 33222, USA
Listed on 2026-07-18
Listing for:
eMed
Full Time
position Listed on 2026-07-18
Job specializations:
-
IT/Tech
SRE/Site Reliability, AWS
Job Description & How to Apply Below
As a Senior Software Engineer in SRE at eMed, you will play a key role in ensuring our platform is highly available, secure, and performant. You’ll lead reliability engineering efforts across production systems, drive operational excellence, and collaborate closely with application and infrastructure teams to design resilient services. This role suits an engineer with a software mindset and deep operational experience, who thrives on improving systems through automation and proactive engineering.
2.2WHAT YOU WILL WORK ON
- Design and implement robust monitoring, alerting, and observability systems across all services and infrastructure
- Lead reliability reviews, incident response, and post-incident analysis—focusing on prevention, learning, and long-term improvements
- Improve service scalability, fault tolerance, and performance through architectural input and systems optimisation
- Build and maintain automation for infrastructure management using Terraform, and delivery pipelines using Git Hub Actions
- Partner with software engineers to improve the operational readiness and resilience of services, including capacity planning and runbooks
- Lead initiatives to reduce operational toil through tooling, automation, and process improvement
- Manage and optimise our production Kubernetes and AWS environments with a focus on reliability, security, and cost-effectiveness
- Contribute to security hardening efforts, including network controls, secrets management, and compliance readiness
- Participate in and lead in-person stand‑ups, incident reviews, and cross‑team planning sessions
- Share knowledge and mentor engineers on best practices in observability, incident response, and operational engineering
WHAT WE’RE LOOKING FOR:
Technical Skills (Essential)
- Strong experience operating Kubernetes and cloud‑native infrastructure (preferably EKS on AWS) in production environments
- Proficiency in AWS services, including networking, compute, IAM, and logging/monitoring tools (e.g., Cloud Watch, ELB, VPC)
- Skilled in Terraform and Infrastructure as Code practices
- Deep understanding of observability tooling (metrics, logs, tracing) and incident management workflows
- Strong coding skills for building tools, scripts, and automation
- Ability to troubleshoot complex infrastructure issues and lead delivery of reliable cloud solutions
- Experience implementing SLAs, SLOs, and error budgets to guide operational priorities
- Background in healthcare or other regulated industries with security and compliance requirements
- Previous involvement in platform security reviews
- Retirement Plan (401k with Company Match)
- Life Insurance (Basic, Voluntary & AD&D)
- Paid Time Off
- Short Term & Long Term Disability
- Training & Development
- Catered Breakfast and Lunch 5 days a Week
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×