
What Is DNS Failover, and How Do You Configure It?
Your server will go down at some point. A host fails, a region has an outage, a deploy goes wrong, or a certificate expires at 3 a.m. The question is whether your users notice for two minutes or two hours. DNS failover is one of the most accessible ways to shrink that window: it watches your primary endpoint and, when it stops responding, automatically points your hostname at a backup. This article explains exactly how DNS failover works, how long it really takes, how to configure it in AWS Route 53 and with a do-it-yourself script, and the mistakes that make failover fail. It goes deeper than the overview in how DNS can be used for load balancing, focusing on the health-check-and-switch mechanism specifically.
What Is DNS Failover?
DNS failover is a DNS configuration in which an authoritative DNS service continuously health-checks one or more endpoints and changes the records it returns based on the results. While the primary is healthy, queries get the primary's address. When health checks fail, queries get the secondary's address instead. When the primary recovers, traffic can move back automatically.
There are two common patterns:
- Active-passive. One primary endpoint serves all traffic. A secondary sits idle, or serves a static maintenance page, until the primary fails.
- Active-active. Several endpoints serve traffic at once, and any that fail their health checks are removed from the answer until they recover.
Both depend on the same building blocks: health checks, routing rules, and the TTL on the records.
How DNS Failover Works Under the Hood
1. Health Checks
The DNS provider runs health checkers, ideally from several regions, that probe your endpoint at a fixed interval. Typical check types are:
- TCP: can a connection be opened on a port?
- HTTP or HTTPS: does a request to a path return a success status code?
- HTTP with string matching: does the response body contain an expected string?
An endpoint is marked unhealthy after a configured number of consecutive failures, called the failure threshold. Checking from several locations protects against a false alarm caused by a network problem between one checker and your server.
2. Record Selection
Failover records are configured as a set. Each record has a role, such as primary or secondary, and is linked to a health check. When the authoritative server answers a query, it returns the highest-priority record whose health check is passing. This decision happens per query, on the authoritative side, in real time.
3. Caching and TTL
Here is the catch. Once a recursive resolver has received an answer, it caches it for the record's TTL. Until that TTL expires, the resolver keeps returning the old address, and users behind it keep trying the dead primary. That is why failover records use short TTLs, commonly 30 to 60 seconds.
How Long Does DNS Failover Really Take?
The total time users experience is the sum of several stages:
| Stage | What determines it | Example |
|---|---|---|
| Detection | Check interval multiplied by the failure threshold | 30 s × 3 = 90 s |
| Provider update | Time for the provider's name servers to start serving the new answer | A few seconds to under a minute |
| Resolver cache expiry | The record TTL | Up to 60 s |
| Client cache | Browser, operating system, and application runtime caches | Usually short, occasionally much longer |
With a 30-second interval, a threshold of 3, and a 60-second TTL, most users are moved within roughly two and a half minutes of the failure. Faster intervals, such as 10 seconds, bring that closer to one and a half minutes, at higher cost.
Two caveats make the real number less predictable:
- Some resolvers enforce a minimum TTL and keep records longer than you asked.
- Some clients cache aggressively. Long-running applications, older Java runtime configurations, and connection pools that never re-resolve can stick to the old IP for much longer. If you control those clients, make sure they respect DNS TTLs.
If you need near-instant failover, DNS alone is not enough. Anycast routing, covered in what anycast DNS is, or a load balancer with a stable IP in front of your servers can fail over in seconds without waiting for caches.
Configuring DNS Failover in AWS Route 53
Route 53 supports failover routing natively. The setup has two parts: a health check and a pair of failover records. For general Route 53 CLI usage, see how to manage DNS records in AWS Route 53.
Step 1: Create a Health Check
Save this health check definition as hc.json:
{
"IPAddress": "203.0.113.10",
"Port": 443,
"Type": "HTTPS",
"ResourcePath": "/healthz",
"FullyQualifiedDomainName": "www.example.com",
"RequestInterval": 30,
"FailureThreshold": 3
}
Then create it:
aws route53 create-health-check \
--caller-reference "www-primary-20261001" \
--health-check-config file://hc.json
Route 53's checkers will request https://www.example.com/healthz from 203.0.113.10 every 30 seconds, sending www.example.com as the host name for TLS and HTTP, and mark the endpoint unhealthy after three consecutive failures. --caller-reference must be unique for each health check you create. The command's output includes the health check's Id, which you need for the next step.
Step 2: Create Primary and Secondary Records
Save this change batch as failover.json, replacing the HealthCheckId with the one returned above:
{
"Comment": "Active-passive failover for www.example.com",
"Changes": [
{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "www.example.com",
"Type": "A",
"SetIdentifier": "primary",
"Failover": "PRIMARY",
"TTL": 60,
"ResourceRecords": [{ "Value": "203.0.113.10" }],
"HealthCheckId": "11111111-2222-3333-4444-555555555555"
}
},
{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "www.example.com",
"Type": "A",
"SetIdentifier": "secondary",
"Failover": "SECONDARY",
"TTL": 60,
"ResourceRecords": [{ "Value": "198.51.100.20" }]
}
}
]
}
Apply it:
aws route53 change-resource-record-sets \
--hosted-zone-id Z0123456789EXAMPLE \
--change-batch file://failover.json
While the health check passes, Route 53 answers with 203.0.113.10. When it fails, Route 53 answers with 198.51.100.20. The secondary has no health check here, so Route 53 always treats it as available, which is a common choice when the secondary is a static maintenance site. If you attach a health check to the secondary too and both fail, Route 53 returns the primary anyway, on the reasoning that some answer is better than none.
Step 3: Verify
Check the health check status as Route 53's checkers see it:
aws route53 get-health-check-status \
--health-check-id 11111111-2222-3333-4444-555555555555
This lists each checker region and its latest observation. Then confirm what the authoritative servers are actually returning by querying one of your zone's Route 53 name servers directly:
dig @ns-123.awsdns-15.com www.example.com A +noall +answer
Replace the server with one from your hosted zone's NS record. Querying authoritatively bypasses resolver caches, so you see the current answer immediately.
Active-Active with Multivalue Answers
For active-active setups, Route 53's multivalue answer routing lets you create several records for the same name, each with its own health check. Route 53 returns up to eight healthy records per query, chosen at random, and drops any whose health check fails. Weighted and latency-based records can also carry health checks, and GeoDNS rules can fail over from one region to another the same way.
Other Managed Options
Most managed DNS and edge providers offer health-checked failover, often as part of a load-balancing product. Cloudflare Load Balancing, for example, uses health monitors and pools with ordered fallback, and many other providers expose a simple primary-and-backup failover setting in their dashboards. When comparing them, check the number and location of health checkers, the minimum check interval, the check types supported, and whether failback is automatic.
A Do-It-Yourself Failover Script
If your DNS provider has an API but no built-in failover, you can build a basic version. This Python script uses only the standard library and the Cloudflare DNS API. It checks a health endpoint, switches the www record to a backup after three consecutive failures, and switches back after five consecutive successes:
import json
import os
import time
import urllib.request
API = "https://api.cloudflare.com/client/v4"
TOKEN = os.environ["CF_API_TOKEN"]
ZONE_ID = os.environ["CF_ZONE_ID"]
RECORD_ID = os.environ["CF_RECORD_ID"]
PRIMARY_IP = "203.0.113.10"
BACKUP_IP = "198.51.100.20"
HEALTH_URL = "https://primary.example.com/healthz"
FAIL_THRESHOLD = 3
RECOVER_THRESHOLD = 5
INTERVAL = 10
def primary_healthy() -> bool:
try:
with urllib.request.urlopen(HEALTH_URL, timeout=5) as resp:
return resp.status == 200
except Exception:
return False
def set_record(ip: str) -> None:
req = urllib.request.Request(
f"{API}/zones/{ZONE_ID}/dns_records/{RECORD_ID}",
data=json.dumps({"content": ip, "ttl": 60}).encode(),
method="PATCH",
headers={
"Authorization": f"Bearer {TOKEN}",
"Content-Type": "application/json",
},
)
with urllib.request.urlopen(req, timeout=10) as resp:
result = json.load(resp)
if not result.get("success"):
raise RuntimeError(result.get("errors"))
current = PRIMARY_IP
fails = passes = 0
while True:
if primary_healthy():
passes, fails = passes + 1, 0
else:
fails, passes = fails + 1, 0
if current == PRIMARY_IP and fails >= FAIL_THRESHOLD:
set_record(BACKUP_IP)
current = BACKUP_IP
print("Primary unhealthy: switched to backup")
elif current == BACKUP_IP and passes >= RECOVER_THRESHOLD:
set_record(PRIMARY_IP)
current = PRIMARY_IP
print("Primary recovered: switched back")
time.sleep(INTERVAL)
Set CF_API_TOKEN to an API token with DNS edit permission for the zone, and CF_ZONE_ID and CF_RECORD_ID to the IDs of the zone and the www A record. primary.example.com is a separate, fixed A record pointing at the primary server, so the health check always tests the primary even after www has been switched. The different thresholds for failing over and failing back add hysteresis, which stops the record from flapping when the primary is unstable. A non-200 response raises an exception in urlopen, which the check treats as unhealthy.
The big weakness of a script like this is that it runs from one place. If the machine running it loses connectivity, it will wrongly fail over, and if it crashes, nothing fails over at all. Run it from at least two locations with a shared state, or prefer a managed provider whose checks are distributed. If you are using Cloudflare for your DNS already, how to set up Cloudflare DNS for your website covers the basics of managing the zone.
Designing a Good Health Check Endpoint
The health check is only as good as what it tests:
- Test what users need. A
/healthzendpoint that returns 200 while the database is down will never trigger failover. Check the critical dependencies, quickly. - Keep it cheap. Health checks hit the endpoint constantly from several regions. Avoid heavy queries.
- Do not make it too strict. If a non-critical dependency such as an analytics service fails, failing over the whole site is usually the wrong move.
- Return proper status codes. Return 200 when healthy and 503 when not, rather than a 200 with an error message in the body.
- Allow the checkers through. Firewalls and rate limiters must allow the provider's health checker IP ranges.
Common Mistakes and Best Practices
- Long TTLs. A 24-hour TTL on a failover record makes failover almost useless. Use 30 to 60 seconds.
- An undersized secondary. If the backup cannot handle full production load, failover just moves the outage. Load-test it.
- Stale data on the secondary. Replicate databases, uploads, and configuration so the backup actually works.
- Never testing it. Schedule failover drills by deliberately failing the primary's health check and confirming traffic moves and returns.
- No alerting. You should get a notification the moment a failover happens, not discover it later. See how to monitor DNS uptime and performance.
- Forgetting the name servers themselves. Failover protects your web servers, not your DNS provider. If your DNS host goes down, failover records go down with it.
- Flapping. Use sensible thresholds and different fail and recover thresholds so brief blips do not bounce traffic back and forth.
DNS Failover FAQ
Usually one to three minutes, depending on the health check interval, failure threshold, and record TTL. Some clients that cache DNS longer than the TTL can take longer to follow the change.
Most failover setups use 30 to 60 seconds. Lower values speed up failover slightly but increase query volume, and some resolvers enforce a minimum anyway.
In active-passive, a single primary serves all traffic and a standby takes over only when the primary fails. In active-active, several endpoints serve traffic at once, and unhealthy ones are simply removed from DNS answers.
No. Failover is performed by your DNS provider, so if the provider's name servers are unavailable, failover is unavailable too. Using a provider with a strong anycast network, or adding a secondary DNS provider, addresses that risk.
They overlap. Load balancing distributes traffic across endpoints, while failover redirects traffic away from failed ones. Many DNS load-balancing products include health-checked failover as a feature.
Usually you do not need to. MX records already have priorities, and sending servers automatically try the next MX host when the first is unreachable, which acts as built-in failover for mail delivery.
Common causes are long TTLs, resolvers or applications caching the old address, a secondary that was not actually working, or a health check that passed even though the site was broken. Query the authoritative server directly to confirm what DNS is returning.
Often yes, but with a stricter recovery threshold so an unstable primary does not cause flapping. For complex systems with data replication, some teams prefer manual failback after confirming the primary is fully healthy.
Conclusion
DNS failover is a practical, affordable way to keep a site reachable when a server or region fails. Health checks detect the problem, the authoritative DNS service starts handing out the backup address, and resolvers pick up the change as soon as their cached answers expire. With a sensible check interval, a short TTL, and a backup that can really carry the load, most users are moved within a couple of minutes without anyone touching a keyboard.
The details decide whether it works when you need it. Build health checks that test what users depend on, add hysteresis to avoid flapping, keep the secondary in sync, alert on every failover, and rehearse it regularly. For outages that cannot tolerate even a couple of minutes, pair DNS failover with anycast or load balancers that fail over below the DNS layer.
Here are some useful references for going deeper on DNS failover:
- AWS Documentation: Configuring DNS failover — Route 53's guide to failover routing and health checks.
- AWS Documentation: Creating, updating, and deleting health checks — health check types, intervals, and thresholds in Route 53.
- Cloudflare Developers: Cloudflare Load Balancing — health monitors, pools, and failover with Cloudflare.
- Cloudflare API Documentation: Cloudflare API — reference for the DNS records endpoints used in the script.
- RFC 1035: Domain Names - Implementation and Specification — defines TTL and caching behavior that governs how quickly failover takes effect.


