Back to Blog
Case Study

How DNS Lookup Solved a Website Outage

By Kishan PrajapatJul 07, 2026
How DNS Lookup Solved a Website Outage

At exactly 8:42 AM on a Monday morning, our customer support channel blew up. Dozens of students in Tokyo, London, and Munich were posting screenshots of error pages saying "This site can't be reached" and "DNS_PROBE_FINISHED_NXDOMAIN". The site was completely down for them, preventing them from accessing their course materials and video lectures. But when I tried loading the e-learning platform from my desk in New York, it loaded in milliseconds, resolving perfectly. Our server CPU, memory, and database health metrics were green. The backend was fine, but our users outside the US couldn't find us. We were looking at a global DNS routing outage.

Is the Server Actually Dead?

When a site works in one country but fails in another, the problem isn't the application; it's the network path or the Domain Name System. An NXDOMAIN error means the client's browser asked the DNS system for our IP address, and the root/top-level authority or local resolver replied that the domain did not exist. If they can't resolve the name, they can't connect to our load balancers. I set up an incident response channel and ran Ping Tests and DNS checks from multiple global locations using IPDekho. I queried our domain from test nodes in London, Frankfurt, Tokyo, Singapore, and Sydney to check how their local recursive resolvers responded.

The tests confirmed the issue. Pings from US-based test nodes resolved our domain name to our AWS Application Load Balancer IP and returned clean responses. But pings from Germany, the UK, Japan, and Singapore failed immediately with "host not found". The global DNS registry was out of sync. European and Asian internet users were receiving incorrect lookup directions, pointing them to a defunct IP address that was no longer responding to HTTP requests.

Tracking the DNS Cache Mismatch

We ran a detailed DNS Lookup on our domain using IPDekho's records tool, querying the active A records across major nameservers. The results showed a major discrepancy: Google DNS (8.8.8.8) and Cloudflare DNS (1.1.1.1) servers in the US returned our correct AWS IP. However, several regional ISP DNS servers in Europe and Asia were still returning an old IP address: 104.24.120.48. This IP belonged to an old, decommissioned Cloudflare reverse proxy that we had shut down the night before during a planned database migration. Because the old proxy was offline, any browser trying to connect to it simply timed out with a connection error instead of reaching our system.

We checked our WHOIS records to ensure our nameserver transition had completed. It was active. But when we inspected our zone records, we spotted the mistake: the Time-to-Live (TTL) value for the old A records had been set to 86,400 seconds (24 hours). The TTL tells DNS resolvers how long they should cache records before asking the nameserver for updates. Because it was set to 24 hours, regional ISPs in Europe and Asia were caching the old, offline IP and refusing to look up the new one. The US servers had refreshed because high traffic forced an early purge, but other regions were stuck. Standard local ISP DNS caches are notoriously sticky, often ignoring the authoritative nameserver directives if a high TTL is set.

Resolving the Outage

We couldn't wait 24 hours for the cache to clear on its own. We logged into our legacy DNS panel, temporarily turned the proxy routing back on for the old IP address, and configured it to forward traffic to our new AWS balancer. Within 10 minutes, site access was restored globally because the old cached IP was now forwarding traffic to the right place. Next, we updated our active records, reducing the TTL from 24 hours to 300 seconds (5 minutes) so future changes sync in under five minutes. We kept the legacy forwarder active for a full day until all old cached records had expired, then shut it down safely.

This outage taught us a simple but important rule: always reduce your DNS TTL values to a low setting 48 hours before you migrate any server. By setting the TTL to 300 seconds ahead of time, any changes you make during the transition will propagate globally in near real-time. This prevents local internet service providers from caching old network locations for days. Using global DNS lookup and WHOIS checks allowed us to see the cached mismatch and fix the outage before we lost a day of business, confirming the value of having diagnostic tools on hand.

Check Your DNS Records Globally

Run a live DNS lookup to query A, MX, CNAME, and TXT records to debug cached routing issues or propagation delays.

Check DNS Records
KP

About Kishan Prajapat

Kishan Prajapat is the founder of IPDekho and an expert in IP intelligence, geolocation APIs, and website security diagnostics with over 6 years of experience helping businesses block fraud and secure local servers.

Share this article: