Senior Cloud/AWS Infrastructure Engineer
Listed on 2026-10-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, Systems Engineer, Cybersecurity
Senior Cloud / AWS Infrastructure Engineer
IT / Infrastructure | Full-Time, Salaried, Individual Contributor | Reports to Director of IT
Company (Reflex Media)
Location Remote, anywhere in the USA.
Compensation $130,000 to $180,000 base.
Schedule Three evenings per week (5:30 to 7:30 PM US Pacific) to overlap with our Singapore Dev and Dev Ops partners. Flexed time available during the day.
Travel Occasional travel to Las Vegas for planning, incident postmortems, and team weeks.
About the Roleis the world's largest premium dating platform, founded and led by an MIT alumnus and headquartered in Las Vegas. We are hiring a Senior Cloud / AWS Infrastructure Engineer to build, operate, and secure the AWS environment that runs our products.
The environment is AWS Organizations with Control Tower, accounts vended through Account Factory for Terraform, and governance applied through global and per-account customization layers. Infrastructure is Terraform, delivered through a versioned in-house module library. Workloads run across ECS Fargate, Aurora MySQL and PostgreSQL, Elasti Cache, EFS, Lambda, Open Search, Redshift, MWAA, and Bedrock, with Cloudflare at the edge.
Disaster recovery. We are building cross-region DR for a flagship platform, essentially from the ground up. Today the platform is regionally concentrated. You will define the DR plan against real RPO and RTO targets, design and build the replication and failover path, and prove it with real exercises.
Security engineering. Identity hardening, replacing static credentials with federated and role-based access, extending preventive and detective controls, and turning findings into engineering fixes.
Who Should Apply- Senior Cloud or Platform Engineers at consumer subscription, marketplace, or SaaS companies who have run AWS Organizations at scale.
- Site Reliability Engineers with a heavy IaC and security bias who have taken a regionally concentrated platform to cross-region DR.
- Dev Ops Engineers from a Terraform-first shop who own module libraries other teams depend on.
- Cloud Security Engineers who built the guardrails, the SCPs, and the federated access patterns their org runs on today.
- Extend the Terraform module library and the AFT customization layers (VPC and IPAM allocation, private hosted zones, SSM access, backup vaults, account baselines) so new accounts land secure and consistent by default.
- Design VPC architecture, cross-account networking, and shared services in a hub-and-spoke configuration model.
- Migrate legacy brand infrastructure into the current account structure, state backend model, and module standards.
- Design and build cross-region recovery for a flagship platform:
Aurora replication topology, cross-region backup copy, S3 and EFS replication, container image distribution, DNS failover, and region-parity IaC. - Establish RPO and RTO targets with the business, map dependencies into recovery tiers, and define restoration order.
- Standardize backup policy and retention across the estate, including organization-level backup policy and centralized vaults.
- Run scheduled DR tests and game days, and write runbooks that make recovery repeatable by anyone on call.
- Design least-privilege IAM policy and IAM Identity Center permission sets. Reduce standing access and direct assignments.
- Replace long-lived IAM access keys with federated SSO, OIDC, and assumed-role access, including in CI/CD.
- Author and maintain SCPs and organization controls, plus automated remediation where prevention is not possible.
- Extend detective controls: AWS Config rules and conformance packs, Security Hub standards coverage, Guard Duty feature coverage, and consistent finding handling across regions.
- Own secrets management on Secrets Manager and SSM Parameter Store, including rotation.
- Support incident response and Cloud Trail forensics, and convert findings into durable engineering fixes.
- Operate core infrastructure: monitoring, alerting, AWS Backup, and patch and lifecycle management.
- Improve our Prometheus, Grafana, and Alert manager stack and its Cloud Watch alarm coverage so real issues page quickly.
- Debug production issues across compute, storage, networking, and databases.
- Automate operational work in Python and Bash, and improve CI/CD for infrastructure.
- Keep Control Tower and the landing zone current, and contribute to cost visibility and optimization.
- Five or more years…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).