Diagnosing a Partial Outage After a DNS Change: A Walkthrough
Updated Published by Kishan Prajapat, SEO & Content Lead
Drafted with Citeya, an AI writing tool built by KPThink.
This is an illustrative scenario, not a report of a real incident. The people, companies, IP addresses (from the ranges reserved for documentation) and figures are made up to show how the investigation works.
An e-learning site moves to a new load balancer overnight and points its A record at the new address. By morning, support is full of messages from students in some countries saying the site times out, while the team in New York can load it fine. Servers, databases and monitoring are all green.
Works here, fails there
When a site works for some users and fails for others, and the application itself is healthy, suspect DNS caching. The error students see is a connection timeout, not “site not found”. That detail matters: their browsers are resolving the name, but to an address that no longer answers.
Compare what different resolvers return
Our DNS lookup queries Google Public DNS, and it returns the new address, which explains why the team sees a working site. The next step is to ask other resolvers directly, for example with dig @resolver-address example.com A against a few ISP and public resolvers used by the affected students. Several still return the old address, 192.0.2.48.
The zone file shows why: the old A record had a TTL of 86,400 seconds, 24 hours. Resolvers that fetched the record shortly before the change are allowed to keep serving the old answer until that TTL runs out. Nothing is broken; the cache is doing what it was told.
Fix it now, prevent it next time
You can’t force other people’s resolvers to drop a cached record, so the fix is to make the old address work again. In this scenario the team brings the old load balancer back up and has it forward traffic to the new one. Within minutes everyone can reach the site, whichever address their resolver has. The old server stays up for a full TTL period after the change, and only then is shut down.
For next time: lower the record’s TTL to around 300 seconds at least a day before any planned change, wait for the old TTL to expire, make the change, then raise the TTL again once everything is stable. Our explainer on TTL covers the details.
Check a Domain’s DNS Records
Look up A, AAAA, CNAME, MX, NS and TXT records for any domain.
Run a DNS LookupSpotted a mistake? Tell usand we'll correct it.
Related Articles
Tracing a Suspicious IP During a Database Leak: A Walkthrough
An illustrative walkthrough: a replica database starts streaming data to an unknown IP at 2 AM. How to stop it, identify the IP’s owner, and find the misconfiguration that let it in.
Using IP and ASN Checks to Cut Checkout Fraud: A Worked Example
An illustrative example of how a small online store could use ASN and network-type checks to hold suspicious orders for review, and where those checks fall short.
Investigating an “Impossible Travel” Login Alert: A Walkthrough
An illustrative walkthrough of triaging an impossible-travel alert with WHOIS, reverse DNS and a port scan, and deciding whether it’s an attack or an employee on a VPN.
WestJet first told Boeing about 737 MAX software g
Explains how WestJet first reported a 737 MAX flight-computer software glitch, why regulators delayed MAX 10 certification, and what airlines and crews must do now.
