
How to Monitor DNS Uptime and Performance?
Most website monitoring checks whether https://example.com returns a 200. That's useful, but it hides DNS inside a single pass/fail signal: when the check fails, you don't know if the web server died or the name stopped resolving, and when it passes, you don't know that one of your four nameservers has been timing out for a week. DNS problems are also sneaky because caching masks them — a broken nameserver can go unnoticed for hours while resolvers serve cached answers, then everything fails at once when TTLs expire.
This article covers what to monitor about your DNS specifically, why each signal matters, and how to build checks for it — from a quick shell script to a Prometheus setup with alerting. It assumes you already understand how DNS works; the focus here is on keeping your own domain's DNS observable.
What "DNS Uptime" Actually Means
For a domain owner, DNS is up when resolvers anywhere in the world can get correct answers for your names quickly. That breaks down into several separate things that can each fail independently:
- Nameserver availability. Each authoritative nameserver answers queries for your zone. You need to check every one, not just "the domain," because resolvers spread queries across all of them.
- Answer correctness. The servers return the records you expect — the right IP addresses, MX hosts, and TXT values. A server returning the wrong answer is worse than one returning nothing.
- Consistency. All nameservers serve the same version of the zone, which you can track through the SOA serial.
- Delegation health. The TLD's NS records for your domain match the NS records in your zone, and all listed servers actually serve it.
- Latency. Authoritative responses arrive quickly from the regions where your users are. Slow DNS directly delays first page loads, as covered in whether DNS settings affect website speed.
- DNSSEC validity. If the zone is signed, signatures haven't expired and the DS record at the parent matches your keys.
- Registration status. The domain hasn't expired and isn't on hold — the failure that takes down everything at once.
A complete monitoring setup checks each of these. Below, they're grouped into checks you can run from any machine.
Check 1: Query Every Authoritative Nameserver
The foundation is a non-recursive query sent to each nameserver individually, recording whether it answered, whether the answer was authoritative, how long it took, and what it returned.
Here's a quick manual version using dig:
ZONE="example.com"
for ns in $(dig "$ZONE" NS +short); do
printf "%-30s " "$ns"
dig @"$ns" "$ZONE" SOA +norecurse +time=2 +tries=1 +noall +answer +stats \
| awk '/SOA/ {serial=$7} /Query time/ {time=$4} END {print "serial=" serial, "time=" time "ms"}'
done
For each nameserver in the zone's NS set, this sends an SOA query with recursion disabled and prints the SOA serial and response time. A missing serial means the server didn't answer. dig +nssearch example.com gives similar output in one command, but scripting it yourself lets you add thresholds and alerts.
For anything you'll run on a schedule, a Python script is easier to extend. This one uses dnspython (pip install dnspython) and exits non-zero on any problem, so it plugs straight into cron, a CI job, or a monitoring agent:
import sys
import time
import dns.exception
import dns.flags
import dns.message
import dns.query
import dns.rcode
import dns.rdatatype
import dns.resolver
ZONE = "example.com"
EXPECTED = {
("example.com", "A"): {"203.0.113.10"},
("www.example.com", "A"): {"203.0.113.10"},
("example.com", "MX"): {"10 mail.example.com."},
}
LATENCY_LIMIT_MS = 250
def nameserver_ips(zone):
servers = []
for ns in dns.resolver.resolve(zone, "NS"):
host = ns.target.to_text()
for family in ("A", "AAAA"):
try:
for addr in dns.resolver.resolve(host, family):
servers.append((host, addr.address))
except dns.resolver.NoAnswer:
pass
return servers
def ask(server_ip, name, rtype):
query = dns.message.make_query(name, rtype)
query.flags &= ~dns.flags.RD
start = time.perf_counter()
response = dns.query.udp(query, server_ip, timeout=3)
elapsed_ms = (time.perf_counter() - start) * 1000
return response, elapsed_ms
problems = []
serials = {}
for host, ip in nameserver_ips(ZONE):
label = f"{host} ({ip})"
try:
soa, ms = ask(ip, ZONE, "SOA")
except (dns.exception.Timeout, OSError) as exc:
problems.append(f"{label}: no response ({exc.__class__.__name__})")
continue
if not soa.flags & dns.flags.AA:
problems.append(f"{label}: answer not authoritative (lame?)")
if ms > LATENCY_LIMIT_MS:
problems.append(f"{label}: slow SOA response {ms:.0f} ms")
for rrset in soa.answer:
if rrset.rdtype == dns.rdatatype.SOA:
serials[label] = rrset[0].serial
for (name, rtype), wanted in EXPECTED.items():
try:
resp, _ = ask(ip, name, rtype)
except (dns.exception.Timeout, OSError):
problems.append(f"{label}: timeout for {name} {rtype}")
continue
got = {rr.to_text() for rrset in resp.answer for rr in rrset}
if resp.rcode() != dns.rcode.NOERROR or got != wanted:
problems.append(f"{label}: {name} {rtype} returned {sorted(got) or dns.rcode.to_text(resp.rcode())}")
if len(set(serials.values())) > 1:
problems.append(f"SOA serial mismatch: {serials}")
if problems:
print("DNS CHECK FAILED")
for p in problems:
print(" -", p)
sys.exit(1)
print(f"DNS OK: {len(serials)} nameserver addresses, serial {next(iter(serials.values()))}")
The script discovers every IPv4 and IPv6 address of every nameserver, then checks each one for an authoritative SOA response, latency under a threshold, the exact expected record values, and a matching SOA serial across all servers. Note that IPv6 checks will fail if the machine running them has no IPv6 connectivity — run it from a dual-stack host, or drop "AAAA" from the family list.
Comparing exact answers has a second benefit: it doubles as change detection. If a record changes without anyone touching it, that's a possible sign of account compromise, which the post on detecting DNS hijacking explores in more depth.
Check 2: Verify the Delegation from the Parent
Your zone's own NS records and the TLD's NS records for your domain should match. When they drift — typically after a provider change where one side wasn't updated — some resolvers end up at servers that no longer serve your zone.
# What the .com TLD delegates to
dig @a.gtld-servers.net example.com NS +norecurse +noall +authority | awk '{print $5}' | sort
# What your zone itself publishes
dig example.com NS +short | sort
The first command reads the authority section of a TLD server's referral; the second reads the NS set from your own zone. Diffing the two sorted lists in a scheduled job catches delegation drift. If you use multiple DNS providers for redundancy, this check is essential, because each provider's servers must appear in both lists.
Check 3: Measure Resolution from Multiple Regions
A query from your office tells you nothing about users in another continent. DNS providers use anycast to route each query to a nearby node, and a single node can be degraded while the rest are fine. You need vantage points in the regions you care about.
Options, roughly in order of effort:
- Commercial uptime services. Many uptime monitoring platforms offer a DNS monitor type that queries a nameserver from several regions and alerts on failures or unexpected answers. This is the fastest way to get multi-region coverage.
- Small cloud instances. Run the Python script above from a few cheap VMs or scheduled serverless functions in different regions, and send results to one place.
- RIPE Atlas. The RIPE NCC's measurement network lets you schedule DNS queries from thousands of probes worldwide, which is excellent for understanding global latency and anycast behavior.
Whichever you choose, track latency as a percentile (p50, p95) over time rather than a single number. A creeping p95 is often the first sign of an overloaded or misrouted anycast node.
Check 4: Prometheus and the Blackbox Exporter
If you already run Prometheus, the Blackbox Exporter has a built-in DNS prober. Define a module per record you want to validate:
# blackbox.yml
modules:
dns_example_a:
prober: dns
timeout: 5s
dns:
query_name: "example.com"
query_type: "A"
transport_protocol: "udp"
preferred_ip_protocol: "ip4"
recursion_desired: false
valid_rcodes:
- NOERROR
validate_answer_rrs:
fail_if_not_matches_regexp:
- "IN\tA\t203\\.0\\.113\\.10$"
This module sends a non-recursive A query for example.com and marks the probe as failed unless the response code is NOERROR and an answer record matches the expected IP. Then point Prometheus at each nameserver through the exporter:
# prometheus.yml (excerpt)
scrape_configs:
- job_name: "dns-authoritative"
metrics_path: /probe
params:
module: [dns_example_a]
static_configs:
- targets:
- ns1.example-dns-host.net
- ns2.example-dns-host.net
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
For the DNS prober, the target is the nameserver to query; the relabeling passes it to the exporter running on port 9115. Finally, add alerting rules:
# dns-alerts.yml
groups:
- name: dns
rules:
- alert: DNSNameserverDown
expr: probe_success{job="dns-authoritative"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Nameserver {{ $labels.instance }} failing DNS checks"
- alert: DNSNameserverSlow
expr: probe_duration_seconds{job="dns-authoritative"} > 0.5
for: 10m
labels:
severity: warning
annotations:
summary: "Nameserver {{ $labels.instance }} responding slowly"
The first rule fires if a nameserver fails its probe for two minutes straight; the second warns on sustained slow responses. The for duration avoids paging on a single dropped packet.
Check 5: Watch DNSSEC Signature Expiry
Signed zones carry RRSIG records with explicit expiration timestamps. If your signer stops re-signing — a broken cron job, a key management failure, an expired HSM credential — the signatures expire and validating resolvers start returning SERVFAIL for your entire domain. This kind of outage is entirely preventable with a simple check:
import datetime
import dns.message
import dns.query
import dns.rdatatype
NAMESERVER_IP = "198.51.100.53"
ZONE = "example.com"
WARN_DAYS = 5
query = dns.message.make_query(ZONE, "SOA", want_dnssec=True)
response = dns.query.udp(query, NAMESERVER_IP, timeout=3)
for rrset in response.answer:
if rrset.rdtype == dns.rdatatype.RRSIG:
for sig in rrset:
expires = datetime.datetime.fromtimestamp(sig.expiration, datetime.timezone.utc)
remaining = expires - datetime.datetime.now(datetime.timezone.utc)
status = "WARN" if remaining.days < WARN_DAYS else "OK"
print(f"{status} RRSIG over SOA expires {expires:%Y-%m-%d %H:%M} UTC ({remaining.days} days)")
The query sets the DNSSEC OK bit so the server includes signatures, then prints how long the SOA record's signature has left. Signers usually refresh signatures well before expiry, so if the remaining time ever drops below a few days, re-signing has stopped. You can do the same check by eye with dig @198.51.100.53 example.com SOA +dnssec — the RRSIG line contains the expiration as a YYYYMMDDHHMMSS timestamp. For more on what's being signed and why, see what DNSSEC is.
Check 6: Track Domain Expiry
Monitor the registration itself. A simple daily job can read the expiry date from WHOIS:
whois example.com | grep -iE "Registry Expiry Date|Expiration Date" | head -n 1
WHOIS output formats vary between registries, so many teams use the RDAP protocol instead, which returns structured JSON. For .com:
curl -s https://rdap.verisign.com/com/v1/domain/example.com \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print([e["eventDate"] for e in d["events"] if e["eventAction"]=="expiration"])'
The RDAP response includes an events list; the script prints the entry whose action is expiration. Alert at 30 days and again at 7, even with auto-renew enabled — payment cards expire.
Turning Checks into Useful Alerts
Good DNS monitoring is quiet until something real happens. A few rules help:
- Alert per nameserver, not per domain. One dead server out of four isn't an outage yet, but it's a ticket. All four failing is a page.
- Require consecutive failures. UDP packets get lost. Two or three failed checks in a row filters out noise.
- Treat wrong answers as critical. An unexpected IP is either a mistake or an attack, and either way you want to know immediately.
- Keep history. Latency and error-rate graphs make it easy to answer "has this been getting worse?" and to hold a managed DNS provider to its SLA.
- Run checks from outside your own infrastructure. If your monitoring depends on the same network as your DNS, both fail together.
If your provider supports automated DNS failover, that's a separate mechanism — it health-checks your origin servers and changes DNS answers. You still need the checks above to confirm DNS itself is healthy.
DNS Monitoring FAQ
A website check can't tell you whether a failure was DNS or the web server, and it won't notice a single failing nameserver while others still answer. Dedicated DNS checks pinpoint the layer and catch partial failures before caching hides them.
Every one to five minutes is typical for nameserver availability and answer correctness. DNSSEC signature and domain expiry checks only need to run daily.
Both, for different reasons. Querying authoritative servers directly tells you what your zone is actually serving. Querying public resolvers shows what end users experience, including caching and DNSSEC validation.
For authoritative servers on a major anycast network, responses under 50 ms from nearby regions are typical. Consistent responses over a few hundred milliseconds from a region where you have users are worth investigating.
It means your nameservers are serving different versions of the zone, usually because a secondary server failed to transfer the latest update. Users may get different answers depending on which server their resolver queries.
Yes, partly. Checking exact expected values for important records and watching for NS changes at the parent will alert you if records change without authorization. It's one of the best early-warning signals you can set up.
Yes. Many resolvers query nameservers over IPv6, and an IPv6-only outage on one server can go unnoticed if you only check IPv4. Run checks from a dual-stack host.
Validating resolvers treat the zone as bogus and return SERVFAIL for every name in it, which is effectively a full outage for a large share of users. Monitoring RRSIG expiration gives you days of warning.
Conclusion
DNS rarely fails all at once in a way a simple website check will explain. More often, one nameserver goes lame, a secondary stops transferring, an anycast node in one region slows down, or a signing job quietly stops — and caching hides the problem until it becomes an outage. Monitoring DNS properly means checking each of those failure points on its own: every nameserver, every address family, exact answers for the records that matter, SOA consistency, parent delegation, signature lifetimes, and registration.
You don't need an elaborate system to start. A scheduled script that queries each nameserver and compares answers catches most real problems, and the Blackbox Exporter or a commercial DNS monitor adds multi-region coverage and alerting on top. Whatever you build, run it from outside your own infrastructure and tune it so an alert always means something.
Here are useful references for building DNS monitoring:
- Prometheus: Blackbox Exporter — the official exporter, including the DNS prober configuration reference.
- dnspython: Documentation — reference for the message, query, and resolver modules used in the scripts above.
- RIPE Atlas: atlas.ripe.net — a global measurement network for running DNS queries from thousands of vantage points.
- RFC 4034: Resource Records for the DNS Security Extensions — defines the RRSIG record and its inception and expiration fields.
- RFC 9083: JSON Responses for the Registration Data Access Protocol (RDAP) — specifies the RDAP response format used for expiry checks.


