DevOps Engineer
Listed on 2026-07-24
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS, IT Infrastructure
Doctronic is the first AI legally authorized to practice medicine, processing millions of consultations monthly with 99%+ treatment plan accuracy validated by board‑certified clinicians.
About the RoleWe’re looking for a Dev Ops Engineer to own our infrastructure. We are HIPAA‑compliant and SOC 2 Type II certified, and you will maintain and strengthen that foundation as we scale to serve millions of patients and enterprise partners.
This role is critical to our mission. When healthcare consultations depend on your infrastructure, reliability isn’t just best practice—it’s a sacred responsibility. You’ll combine hands‑on technical work with strategic infrastructure leadership to keep Doctronic as the most trusted AI diagnostic platform.
What You'll Build- Design, deploy, and maintain AWS infrastructure using ECS and core services (EC2, IAM, VPC, ALB/NLB, Cloud Watch, ECR, S3, RDS, Glue).
- Operate and scale production Kubernetes clusters with Helm‑based application deployments.
- Implement Git Ops workflows with Argo CD for secure, automated, and auditable releases.
- Provision and manage cloud infrastructure with Terraform and automate operational workflows.
- Build and optimize CI/CD pipelines with Git Hub Actions, Git Lab CI, or similar platforms.
- Deploy, maintain, and optimize infrastructure for AI/ML services and large language models (LLMs) with a focus on performance, reliability, and cost efficiency.
- Collaborate with engineering teams to deliver production‑ready AI platforms.
- Design and manage cloud networking (VPCs, VPNs, Load Balancers, DNS, routing) and secure connectivity.
- Integrate SIEM solutions and implement security best practices, identity management, and least‑privilege access across cloud environments.
- Build monitoring, logging, and alerting solutions using Prometheus, Grafana, and Cloud Watch.
- Improve platform reliability through proactive monitoring, incident response, and performance optimization.
- Build and optimize containerized workloads using Docker and administer Linux‑based production environments.
- Automate operational tasks and infrastructure management using Bash and Python.
- 5+ years of experience as a Dev Ops Engineer or Site Reliability Engineer.
- Strong hands‑on experience with AWS, Kubernetes, Terraform, Docker, Helm, and Argo CD.
- Proven experience designing and operating highly available, cloud‑native production infrastructure.
- Solid understanding of Linux administration, networking, cloud security, and Infrastructure as Code principles.
- Experience building and maintaining CI/CD pipelines and deployment automation.
- Familiarity with monitoring, logging, and observability platforms.
- Experience supporting AI/ML workloads or modern distributed systems is a strong advantage.
- Strong problem‑solving skills with the ability to troubleshoot complex production environments.
- Comfortable working in cross‑functional teams and collaborating closely with software engineers, security teams, and product stakeholders.
- Passionate about automation, reliability, scalability, and operational excellence.
- Experience with GPU infrastructure and AI/ML model deployment.
- Familiarity with Dev Sec Ops practices, vulnerability management, and compliance frameworks.
- Experience with multi‑cluster Kubernetes environments.
- Knowledge of performance tuning, autoscaling, and cost optimization in AWS.
- Experience supporting high‑traffic, mission‑critical production systems.
- Understanding of disaster recovery, backup strategies, and business continuity planning.
- Experience working in fast‑paced startup or scale‑up environments.
New York City | On‑site
Equity OpportunitiesShare in Doctronic’s growth as we transform healthcare with AI.
BenefitsWe offer comprehensive health, dental, and vision coverage, plus mental health support and flexible time off.
Compensation Range: $180K – $240K
Reports ToDirector of Engineering
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).