×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, DNS

Job in Plano, Collin County, Texas, 75086, USA
Listing for: Optimum
Full Time position
Listed on 2026-07-09
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Unix/Linux, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 83538 - 137241 USD Yearly USD 83538.00 137241.00 YEAR
Job Description & How to Apply Below

Job Summary

The Role DNS Engineer – SRE is a high‑impact role responsible for the architecture, scalability, and reliability of the mission‑critical DNS infrastructure powering our ISP and core network services. This position is designed for an engineer who views infrastructure through the lens of Site Reliability Engineering (SRE) prioritizing automation, observability, and self‑healing systems over manual intervention. You will combine deep IP networking and DNS expertise with modern security protocols to ensure our platforms remain resilient against evolving threats and perform at the highest level for millions of users.

Beyond core engineering, you will serve as a technical authority, leading cross‑functional initiatives with Product, Security, and Service Assurance teams to deliver a carrier‑grade DNS ecosystem that balances cutting‑edge privacy standards with uncompromising availability required by Tier‑1 network operations.

Responsibilities
  • Architectural Ownership – lead the design and evolution of global DNS architectures, ensuring high availability through Anycast routing, multi‑provider redundancy, and automated failover mechanisms.
  • Strategic Vendor Relations – act as the primary technical authority in engagements with DNS and infrastructure vendors, driving roadmaps that align with our long‑term reliability and security goals.
  • Lifecycle & Capacity Management – oversee the full lifecycle of DNS platforms, including automated software deployments, hardware refreshes, and proactive capacity planning.
  • Standardization & Policy – optimize, define, and enforce organization‑wide standards for DNS record management, security protocols (DNSSEC), and traffic steering policies to optimize user latency.
  • Reliability Engineering – define Service Level Objectives (SLOs) and error budgets for all core name services to convert strategic design into operational reality.
  • Protocol Management – manage nuances of UDP/TCP port 53, recursion vs. iteration, and complex record types (A, AAAA, CNAME, MX, TXT, SRV).
  • Security & Mitigation – implement DNSSEC, mitigate cache poisoning, and serve as subject matter expert in defending against DDoS and DNS amplification attacks.
  • Automation – replace manual updates and pool management with automated workflows using Python, Go, Ansible, or Terraform.
  • Performance Tuning – tune Linux kernel for high‑performance network throughput and conduct deep‑dive log analysis on BIND, Unbound, or PowerDNS systems.
  • Observability – utilize Prometheus, Grafana, and dnstap to monitor query rates, latency, and error codes (NXDOMAIN, SERVFAIL), providing actionable insights.
Qualifications

Minimum Qualifications
  • Bachelor’s degree in Computer Science, Telecommunications, or related field (or equivalent practical experience).
  • 5+ years in networking or systems engineering with a focus on SRE principles (automation, reliability, monitoring).
  • Hands‑on experience configuring and maintaining at least two of: BIND, Unbound, PowerDNS, AWS Route 53, or Azure DNS.
  • Functional understanding of TCP/IP (IPv4/IPv6) and DNS‑specific protocols including DNSSEC and encrypted transport (DoH/DoT).
  • Strong Linux/Unix administration and proficiency in at least one scripting language (Python, Bash, or Go) for task automation.
  • Experience using Grafana and Open Telemetry (or similar) to monitor service health and performance.
Preferred Qualifications
  • Hands‑on experience managing BIND, Unbound, or PowerDNS in high‑traffic environments, plus cloud‑native solutions (AWS Route 53, Azure DNS, Google Cloud DNS).
  • Mastery of DNS‑specific protocols (DNSSEC, DoT, DoH) and underlying transport layers (UDP/TCP) with dual‑stack (IPv4/IPv6) networking.
  • Experience building dashboards and alerts using Prometheus, ELK, or Open Telemetry to monitor DNS query latency and error rates.
  • Automation expertise managing “DNS as Code” with Terraform or Ansible and scripting (Python/Go) to automate routine zone updates.
  • Background in Tier‑1/Tier‑2 service provider environments focusing on service resilience, Anycast distribution, and DDoS protection.
Working Conditions
  • Hybrid remote/on‑site with participation in a 24/7 on‑call rotation.
  • Availability for…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary