<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy" -->

---
title: Day 6/30 Real-World DevOps Problems | daily.dev
description: A scenario-based DevOps puzzle describes a load balancer migration where DNS TTL was lowered from 3600 to 60 seconds before the cutover, yet two days later 4%...
canonical: https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Day 6/30 Real-World DevOps Problems | daily.dev
og:description: A scenario-based DevOps puzzle describes a load balancer migration where DNS TTL was lowered from 3600 to 60 seconds before the cutover, yet two days later 4%...
og:url: https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy
og:image: https://api.daily.dev/og/posts/QPolk7bdY.png
og:image:alt: Day 6/30 Real-World DevOps Problems
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Day 6/30 Real-World DevOps Problems

**[Bobby Iliev](https://daily.dev/sources/9adj6fgmy)** · [@bobbyiliev](https://daily.dev/bobbyiliev) · 1 min read · 16 upvotes · 13 comments

## Summary

A scenario-based DevOps puzzle describes a load balancer migration where DNS TTL was lowered from 3600 to 60 seconds before the cutover, yet two days later 4% of traffic still hits the old load balancer despite dig lookups from multiple networks showing the new address. Readers are asked to pick the correct next action among waiting it out, trying to purge public resolver caches, lowering TTL further, or inspecting the old load balancer's logs to identify lingering clients and keeping it running.

## Content

You moved your app to a new load balancer.

The day before, you lowered the A record's TTL from 3600 to 60. On Monday at 10:00 you pointed the record at the new load balancer.

It is now Wednesday. 4% of requests still arrive at the old load balancer. dig from three different networks returns the new address.

What do you do next?

**A) Wait. DNS propagation can take up to 72 hours to reach every resolver on the internet**

**B) Purge the old record using the public resolvers' cache flush pages**

**C) Lower the TTL again, from 60 to 5 seconds, so clients refresh faster**

**D) Check the old load balancer's logs for who still connects, and keep it forwarding**

What's your next step?

## Community discussion

Top comments from developers on daily.dev.

**@bobbyiliev** · 2 upvotes

> ![GIF](https://static.klipy.com/ii/9294a2e836d178ddc22430dd7765727e/90/a2/HvQPzzhihrFjBww1.gif)

**@michalfilo** · 2 upvotes

> I would go with D since some places could have hardcoded IP or pinned DNS.

**@ahmetozel** · 1 upvotes

> I would choose D and use the old load balancer as a source of evidence rather than treating three dig results as proof that every client has moved. The remaining requests might come from long-lived connections, application-level DNS caches, or hardcoded addresses, so another TTL reduction would not necessarily affect them.
>
> I would compare source addresses, user agents, request paths, and connection age where available, then keep the old endpoint safely forwarding while the affected clients are identified. The useful retirement signal is that those clients have been corrected or drained, not...

**@endon98** · 1 upvotes

> B. YOLO

**@bobbyiliev** · 0 upvotes

> The answer is D.
>
> Lowering the TTL a day ahead gave ordinary caches time to drop the old 3600-second value, and the new 60-second TTL means normal resolver caching cannot explain traffic two days later. Whatever still reaches the old load balancer is holding on to the old address some other way.
>
> Common causes:
>
> - long-lived connections (HTTP keep-alive pools, WebSockets, gRPC) that never reconnect
> - clients that cache DNS results in the process for a long time
> - the old IP address hardcoded in a client or proxy configuration
> - a few resolvers that do not honour TTLs
>
> The old load balancer's...

## Similar posts on daily.dev

- [DNS TTL Was Still 300. We Rolled Back in Forty Seconds. Clients Hit Dead Pods for Five Minutes](https://daily.dev/posts/dns-ttl-was-still-300-we-rolled-back-in-forty-seconds-clients-hit-dead-pods-for-five-minutes-gdxqrecqu) · Medium · 3 upvotes · 0 comments
- [How To Get DNS Right: A Guide to Common Failure Modes](https://daily.dev/posts/how-to-get-dns-right-a-guide-to-common-failure-modes-4gzax6vbj) · The New Stack · 3 upvotes · 0 comments
- [Protect identity infrastructure in cloud-native environments](https://daily.dev/posts/protect-identity-infrastructure-in-cloud-native-environments-3ooefrd27) · Red Hat Developer · 0 upvotes · 0 comments
- [More Than DNS: The 14 hour AWS us-east-1 outage](https://daily.dev/posts/more-than-dns-the-14-hour-aws-us-east-1-outage-m4b0fv6uu) · Lobsters · 1 upvotes · 0 comments
- [Medium](https://daily.dev/posts/medium-mnet7hexv) · Medium · 0 upvotes · 0 comments

---

Tags: [#devops](https://daily.dev/tags/devops), [#networking](https://daily.dev/tags/networking), [#dns](https://daily.dev/tags/dns)

[View this post on daily.dev](https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy","headline":"Day 6/30 Real-World DevOps Problems","text":"A scenario-based DevOps puzzle describes a load balancer migration where DNS TTL was lowered from 3600 to 60 seconds before the cutover, yet two days later 4% of traffic still hits the old load balancer despite dig lookups from multiple networks showing the new address. Readers are asked to pick the correct next action among waiting it out, trying to purge public resolver caches, lowering TTL further, or inspecting the old load balancer's logs to identify lingering clients and keeping it running.","url":"https://daily.dev/posts/day-6-30-real-world-devops-problems-qpolk7bdy","datePublished":"2026-10-10T11:58:08.029Z","dateModified":"2026-10-10T11:58:23.828Z","author":{"@type":"Person","name":"Bobby Iliev","url":"https://daily.dev/bobbyiliev","image":"https://avatars3.githubusercontent.com/u/21223421?v=4","description":"DevOps, DevEx, cloud & open source\n","worksFor":{"@type":"Organization","name":"Materialize","logo":"https://res.cloudinary.com/daily-now/image/upload/s--4mL2CrlK--/f_auto/v1725263959/companies/materialize"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":83240}},"image":"https://media.daily.dev/image/upload/s--GEV4TCMf--/f_auto/v1791633488/posts/QPolk7bdY?_a=BAMAMicg0","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":16},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":13}],"comment":[{"@type":"Comment","text":"","datePublished":"2026-10-10T11:59:20.916Z","url":"https://daily.dev/posts/QPolk7bdY#c-zodyB42Vd","author":{"@type":"Person","name":"Bobby Iliev","url":"https://daily.dev/bobbyiliev","image":"https://avatars3.githubusercontent.com/u/21223421?v=4"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2}},{"@type":"Comment","text":"I would go with D since some places could have hardcoded IP or pinned DNS.","datePublished":"2026-10-10T17:43:20.492Z","url":"https://daily.dev/posts/QPolk7bdY#c-L7sArsm93","author":{"@type":"Person","name":"Michal Filo","url":"https://daily.dev/michalfilo","image":"https://lh3.googleusercontent.com/a/ACg8ocJw7RT2L5Tl9Qa2OnmI7uwXbrBSH2qT6cLpHU8fAuHxrbOUHA=s96-c"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2}},{"@type":"Comment","text":"I would choose D and use the old load balancer as a source of evidence rather than treating three dig results as proof that every client has moved. The remaining requests might come from long-lived connections, application-level DNS caches, or hardcoded addresses, so another TTL reduction would not necessarily affect them.\nI would compare source addresses, user agents, request paths, and connection age where available, then keep the old endpoint safely forwarding while the affected clients are identified. The useful retirement signal is that those clients have been corrected or drained, not simply that another propagation window has elapsed.","datePublished":"2026-10-11T09:51:53.190Z","url":"https://daily.dev/posts/QPolk7bdY#c-GfKPiT1Y7","author":{"@type":"Person","name":"Ahmet Özel","url":"https://daily.dev/ahmetozel","image":"https://avatars.githubusercontent.com/u/70992231?v=4"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"B. YOLO","datePublished":"2026-10-10T22:08:08.837Z","url":"https://daily.dev/posts/QPolk7bdY#c-FH4g68DwY","author":{"@type":"Person","name":"Endon98","url":"https://daily.dev/endon98","image":"https://avatars.githubusercontent.com/u/67847029?v=4"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"The answer is D.\nLowering the TTL a day ahead gave ordinary caches time to drop the old 3600-second value, and the new 60-second TTL means normal resolver caching cannot explain traffic two days later. Whatever still reaches the old load balancer is holding on to the old address some other way.\nCommon causes:\n\nlong-lived connections (HTTP keep-alive pools, WebSockets, gRPC) that never reconnect\nclients that cache DNS results in the process for a long time\nthe old IP address hardcoded in a client or proxy configuration\na few resolvers that do not honour TTLs\n\nThe old load balancer’s access logs help you find them: source IPs, user agents, and paths. Keep the old endpoint forwarding to the new backend, contact the owners of the remaining clients, and switch it off only when the logs show no more traffic.\nA) There is no universal 72-hour propagation period. Ordinary caches expire after the TTL they received, which here was at most an hour.\nB) Not the first step. You do not yet know which resolvers the stuck clients use, and flushing unrelated ones does nothing for clients that never look up the name.\nC) A lower TTL does not shorten lifetimes already cached, and it does not move existing connections or pinned addresses.\nThe mental model: DNS only moves new lookups. Plan for the clients that never look again, and never switch off the old endpoint until its logs say nobody uses it.\nGo deeper: DNS simulator\nhttps://devops-daily.com/games/dns-simulator","datePublished":"2026-10-11T10:31:55.117Z","url":"https://daily.dev/posts/QPolk7bdY#c-OJKoFeGZU","author":{"@type":"Person","name":"Bobby Iliev","url":"https://daily.dev/bobbyiliev","image":"https://avatars3.githubusercontent.com/u/21223421?v=4"}}],"isPartOf":{"@type":"WebPage","url":"https://daily.dev/sources/9adj6fgmy","name":"Bobby Iliev"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Bobby Iliev","item":"https://daily.dev/sources/9adj6fgmy"},{"@type":"ListItem","position":3,"name":"Day 6/30 Real-World DevOps Problems"}]}
```

