Ignoring DNS as a Security Layer Is Costing Companies More Than They Know
A buggy agent update, a records change, and a user routed to the wrong side of the world. All three were visible in DNS. Nobody was looking.
When a service goes down, the incident bridge fills with teams trying to establish which part of the network failed. Andrew Wertkin knows those calls well. The DNS team blames the firewall, the firewall team blames DNS, and executives wait on a separate bridge for an update.
Andrew Wertkin is Chief Strategy Officer at BlueCat, where he previously led product and technology. His work with enterprise customers has made him wary of diagnosing a problem from its traffic graph alone. Two of the apparent attacks he describes in this episode turned out to be changes someone had deliberately deployed.
On Episode 17 of the Full Metal Packet podcast, hosts Yegor Sak and Alex Paguis ask what DNS records can establish during an investigation, where encrypted DNS changes visibility, and why knowing which device asked a question can matter more than the domain's reputation.
TL;DR
- Check recent changes when DNS traffic spikes. Andrew describes a faulty endpoint agent and a DNS records change that each produced traffic resembling an attack.
- Log answers as well as queries. Investigators need to know what a device was told, not just which name it requested.
- Compare queries with the device's normal behavior. An unfamiliar lookup from a device with a narrow purpose can warrant investigation even when the domain has a good reputation.
- Choose resolution placement by workload. Andrew's examples balance latency, local survivability, and the cost of operating services across many sites.
- Keep visibility at a resolver you control. Managed endpoints can use an organization-operated DoH or DoT service. Encryption protects the connection; the resolver can still log the requests it handles.
Recognition Has Grown. Logging Gaps Remain.
Yegor asks whether DNS is still treated as plumbing rather than as a security or control layer. Andrew says recognition has grown and companies are implementing more solutions. He credits some of that attention to larger vendors marketing into the category, though he remains skeptical of their claims.
"I wouldn't start a cybersecurity architecture with DNS. It's additive and it needs to be added."
In his own customer base, he still encounters organizations that do not log DNS traffic. He cannot quantify the proportion, and his point is narrower than saying nothing has improved: greater recognition has not consistently produced better hygiene.
"I mean, I have customers that have port 53 open from the entire company, so people can choose a different DNS server if they want to."
When the Spike Is Not an Attack
Asked for war stories, Andrew starts with operational failures that customers initially suspected were attacks.
"They thought something was going on, but actually something was just misbehaving."
In the first case, an endpoint management agent had been upgraded across a few hundred thousand clients. A bug left the agent repeatedly trying to download additional information. The resulting DNS traffic looked alarming, but the excess queries pointed to one domain, giving investigators a place to start.
In the second, a customer increased the size of TXT records at zone apexes. Andrew describes servers using a 1232-byte UDP buffer following DNS Flag Day. Responses no longer fit, so resolution fell back to TCP. A default configuration with too few TCP workers could not handle the load.
"And you know what DNS clients do? They just keep retrying and retrying and retrying and retrying."
TCP connections and queries per second climbed. Establishing the cause required connecting the records change to the larger responses, TCP fallback, and server capacity. These were real operational incidents, even though neither was malicious.
Ben Lipczynski discusses checking findings before acting in Episode 10: a scanner result deserves investigation, but the team still needs to establish whether it applies to its environment.
Everybody on the Call Is Proving It Was Not Them
Andrew describes two bridges during a severe service outage: one where executives want updates every twenty minutes, and another with DNS, network, firewall, proxy, application, and cloud teams.
"They're less trying to solve the problem and they're more trying to prove it wasn't the firewall or it wasn't DNS."
He calls it the joke about lowering mean time to innocence and includes BlueCat in it. When his team is pulled into those calls, it needs evidence to show whether its service is responsible. The broader problem is that the teams lack a shared view of the query path.
That gap affects performance investigations too. Andrew describes enterprises that adopted direct internet access and SD-WAN while leaving DNS routing unchanged. Application traffic could exit locally, but a lookup still traveled across the world. A user in Singapore could consequently be directed toward a service location near Illinois. Understanding that mismatch required following the DNS query, not just checking the application's route.
From Securing the Service to Reading the Traffic
For BlueCat, Andrew says, security originally meant keeping DNS and DHCP services healthy and hardening the appliances that ran them. The company later added threat feeds through DNS response policy zones, allowing customers to block known-bad destinations.
Those feeds have a limit: someone must already have identified the threat. Andrew is also interested in behavior that looks unusual for the device making the request.
"Blood pressure monitors don't look up google.com, but google.com nobody's looking at as a potential bad domain."
In his example, a monitor resolves a handful of names in bursts when it is being used or its data is being retrieved. A query outside that pattern deserves attention even if the domain itself is reputable.
When Yegor describes this as an allow list with blocking by default, Andrew makes a distinction: the policy can alert rather than block. The response depends on the customer's needs. For these use cases, he finds device identity more useful than user identity.
The Logging Problem Came First
Customers had asked BlueCat to send DNS queries to a SIEM for compliance and security analysis. Andrew says full-time logging then imposed a substantial performance cost and generated large volumes of data. It also mostly captured questions without their answers.
A lookup may be answered from cache or involve several servers. That makes the collection point important: the record needs to show enough of the transaction to explain what happened.
BlueCat's initial work addressed logging at scale and providing DNS-specific analysis alongside tools such as Splunk. Customers who opted in also supplied data for secondary analysis. Andrew describes volumes reaching hundreds of billions of queries a month from some customers.
Mo Balakrishnan describes the value of telemetry collected before an incident in Episode 14. His team used records from across the environment to reconstruct an attack while the intruder was still inside.
The Domain Looks Fine. Which Device Asked For It?
Andrew recalls customers that had deployed compromised Orion software during the SolarWinds attack. Once indicators became public, their retained DNS records helped them identify the relevant queries.
He describes attackers choosing ordinary-looking domains and generating activity before using them maliciously, making domain age and appearance less useful signals. His response is to examine the device making the request: has it queried that domain before, and does the lookup fit its normal behavior? He particularly emphasizes devices that are not user-driven.
Alex raises a different way familiar assumptions can fail: phishing that abuses IPv6 reverse-DNS names. Infoblox documented the campaign in February 2026. The case shows why infrastructure namespaces should not be assumed harmless. It does not establish that all security products overlook them.
Andrew also explains why DNS responses matter in an investigation: malware can use them to receive instructions, rather than simply to locate a website.
"It's not just about stopping a connection to a bad site. You're shutting down a command and control chain."
Devon Ackerman's discussion of containment in Episode 8 gives that point a broader incident-response context: limit the access and communication an attacker can use.
Where Resolution Should Actually Happen
Andrew describes a mix of local and cloud services across BlueCat's customers. Placement depends on what needs to keep working, the latency it can tolerate, and the operational cost of running it.
| Workload | Placement Andrew describes | Reason |
|---|---|---|
| Trading platforms and high-frequency finance | On-premises, including hardware deployments | Latency and round-trip time matter |
| Hospitals and manufacturing plants | Local services | Operations need to survive a lost external link |
| Large retail estates | Cloud or regional core services, with caching resolvers at stores | A full DNS deployment at every site is expensive to operate |
| Application server looking up a nearby database | Local resolution | Avoid sending an internal lookup through an external service |
Dedicated hardware is not always a performance requirement. Andrew says many customers use virtual machines. Some choose appliances because the network team can manage them directly instead of depending on another team's access to the virtualization platform.
Encrypted DNS Solves One Problem and Creates Another
The hosts challenge the assumption that internal DNS should remain unencrypted when other internal traffic has moved to HTTPS. Andrew acknowledges the operator's preference for being able to observe DNS directly, but describes a way to retain records while encrypting the client connection: provide a DoH service that managed devices use, and log at that resolver.
He prefers DoT's dedicated port because it makes DNS traffic easier to distinguish from HTTPS traffic. That does not make the encrypted queries readable in transit. With either DoH or DoT, an organization-operated resolver can retain visibility into the requests it handles.
The client-to-resolver connection is also separate from the resolver's connection to authoritative servers. RFC 9539 describes experimental encrypted transport on that second hop. That standard concerns the resolver-to-authoritative hop, not the encrypted connection from an employee’s device to its resolver.
Andrew also mentions ordinary DNS over TCP when discussing spoofing. The distinction matters: TCP's handshake helps against certain address-spoofing attacks, but it does not encrypt queries or authenticate the resolver. DoT can provide those protections when server authentication is enforced. See RFC 7766 and RFC 7858.
His other concern is dependency. If every lookup relies on one external provider with no working fallback, an outage there can prevent new names from resolving across the organization. This is a provider and architecture choice, not an inherent property of DoH.
"All of a sudden there's a half hour where nothing's being resolved, and that's a problem."
BlueCat supports DoH and DoT, Andrew says. His skepticism is about their value and operational tradeoffs in the enterprise deployments he sees.
Under Zero Trust, the Queries Move Somewhere Else
Andrew describes finance customers with tightly separated internal and public DNS. A laptop cannot resolve public names through the corporate DNS service; proxies perform those lookups instead.
As these organizations adopt service-edge products, the location of DNS visibility changes. Depending on the architecture, BlueCat may not see outbound public lookups. For private application access, it may see queries from the intermediary service rather than the original client. Investigators need to know where those records now live.
Device attribution has another limit: a query does not prove a person visited a site. Andrew uses restaurant searches as an example. A results page can trigger several lookups for images and other resources without the user opening those restaurants' websites.
He acknowledges customers' concern about queries escaping onto the public network. Encryption can protect their contents in transit, but keeping queries on an approved path requires resolver selection, endpoint configuration, and network policy. DoT alone does not keep a query inside the building.
Something Changed and He Wanted to Know Why
Andrew recalls finding roughly fifty queries using private-use record types among a hundred billion queries. He initially suspected a security scanner checking for an old DNS-server flaw. Similar traffic appeared at other customers, and he eventually traced it to Chrome testing whether those queries could reach Google's resolvers.
BlueCat's heuristics surfaced the unusual pattern. The investigation still depended on someone wanting to explain it.
"We're all pattern matchers. And you find patterns really quickly."
His concern about AI is how people will build that expertise if they stop investigating unfamiliar behavior themselves. He also says organizations should use the tools available to them. The tension is between solving today's problem efficiently and giving people the experience to recognize tomorrow's.
Asked what technical leaders should take away, he returns to the practical requirement:
"You need an observable DNS environment."
He means performance and application availability as well as security.
Making the Query Path Inspectable
The episode gives teams a concrete starting question: can you identify the resolver a device uses and retrieve the DNS evidence needed to investigate a problem? Control D offers tools for parts of that work.
- The DNS Leak Test identifies the resolver handling its test queries. It is a useful spot check, not proof that every application on the device follows the same path.
- Dragonfly helps investigate a domain's category and technical details.
- Real-time DNS logging provides query records with device and user context.
- SIEM data streaming sends DNS activity into a team's existing analysis system.
The Control D guide to DNS logging best practices covers queries, responses, collection points, and retention. When evaluating a service, check which of those fields its actual logs contain. Query attribution alone does not establish that full responses are retained.
Also ask what happens when the resolver provider is unreachable. Those two checks follow directly from the episode: preserve useful evidence without overlooking the dependency introduced by the service.
Continue the Argument
- Episode 8: Ex-FBI Agent: One Phone Call Gave Hackers Full Network Access discusses containment and limiting the attacker's access.
- Episode 10: Navy Officer Reveals the Threat Modeling Mindset Most Cybersecurity Teams Are Missing examines why teams should check whether scanner findings apply before acting.
- Episode 14: I Unplugged The Entire Company To Stop A Live Breach shows how telemetry supports decisions during a live incident.
Browse the full Full Metal Packet series.
Common Questions Answered in Episode 17
What does DNS observability actually mean?
A usable record of queries, responses, and the devices involved, collected where an investigator can reconstruct what happened. Query-only records or incomplete coverage leave gaps.
Should enterprises block DNS over HTTPS on their network?
Andrew describes an alternative: provide an organization-operated DoH service and configure managed clients to use it. The episode does not establish one policy for every enterprise.
Is DoH or DoT better for enterprise networks?
Andrew prefers DoT's dedicated port for distinguishing DNS traffic. Both encrypt queries in transit, and both can support logging at an organization-operated resolver. Protocol choice alone does not determine who holds the logs.
Should DNS resolution run on-premises or in the cloud?
Andrew's customer examples use both. Latency, local survivability, and the cost of operating services across sites shape the decision.
What causes a sudden spike in DNS queries?
The traffic graph alone cannot establish the cause. Andrew describes one deployment that sent clients into a retry loop and another change that overwhelmed the servers' TCP capacity. Recent changes belong in the investigation alongside possible malicious activity.
Can DNS logs identify what kind of device is on a network?
Andrew says BlueCat can infer device types such as printers or VoIP phones from query behavior, though not specific models or versions. That can help establish an expected pattern for a device and identify unusual requests.
How does DNS relate to command and control?
Malware can use DNS exchanges to communicate with an operator. Blocking relevant DNS traffic can interrupt that channel, but does not by itself establish that the device is clean or every communication path is closed.
About the guest: Andrew Wertkin leads strategy at BlueCat. His customer experience spans finance, retail, manufacturing, and education.
Full Metal Packet is hosted by Control D co-founders Yegor Sak and Alex Paguis. Watch Episode 17 on YouTube, or listen on Apple Podcasts or Spotify.