BLOG
From dirt to data. Your digital estate is now his dig site.
Interview with Bill Burns, Director of Digital Archaeology, Diskover
September 23, 2026 · 9 min read

Before Bill Burns opened his first hard drive in a forensics lab, he was walking a grid in Pennsylvania and New Jersey, digging shovel tests and excavation units ahead of pipelines, highways, and cell towers. The job was to find out what sat in the ground before the bulldozers arrived.
Twenty years later, Burns is Director of Digital Archaeology at Diskover, asking the same question of different ground.
As an archaeologist, “defensible” was literal. Before a machine broke ground, the survey had to show what was down there, whether it mattered, and what would be done about it, in a record that could stand up to a state review.
On digital ground, “defensible deletion” is the practice of removing data under documented rules, so that every file removed can be explained afterward to an auditor, a regulator, or opposing counsel. It begins the way every dig begins, with a survey of ground that has no written record.
The dig was always about compliance.
The archaeology was not treasure hunting. Burns worked in cultural resource management, the compliance side of the field. When a United States project uses federal money, needs a federal permit, or touches federal land, Section 106 of the National Historic Preservation Act requires the agency to find out whether the work will destroy anything historically significant.
So the survey came first: walk the ground, dig test holes on a grid, decide whether the site matters, and excavate fully only when the project cannot avoid it. In the Northeast that meant Lenape camps along the river terraces, and colonial mills and iron furnaces.
“Mostly what we were digging for was a defensible answer to the question: is there anything here worth saving, and if so, what do we do about it?”
— Bill Burns, Director of Digital Archaeology, Diskover
The connection to his next career arrived on paper. He moved into digital forensics, starting at a utility, and the first time he filled out a chain-of-custody form for a piece of evidence, he recognized it, almost line for line, as the one he had completed in the field.
“On a dig, you bag each artifact and write down everything about it so it’s documented and preserved: what was found, what type of soil it came out of, what color, what else was in that layer. I was doing exactly the same thing for a hard drive.”
The fragments are metadata, and some of them lie.
An archaeologist reads a site from fragments and their position, because the site came with no instructions. A data estate works the same way. Burns has a question he asks every room: “If your files could talk, what would they say?” They can’t, of course. But they leave metadata: the file name, the type, where it sits, how big it is, when it was created, changed, and last touched, and who owns it. Each of those facts exists whether or not a person ever opens the file.
They say more than people expect. Burns has found a folder named “current” where most of the data was years old and already labeled as archive, and a 1 terabyte (TB) log file untouched for years. He has found two spreadsheet files, 18 TB between them and too large for any application to open, sitting side by side. They turned out to be copies of each other.
At one organization, one of the largest content types across employees’ company-issued OneDrive accounts was wedding videos. Personal files sat on corporate storage and were backed up at corporate expense.
Part of the discipline is knowing which fragments to trust. Automated processes touch files and reset the last-access date, so last-modified is the better signal for stale data. Timestamps also get corrupted, which is why his reports carry files dated before computers existed, and files dated centuries into the future.
Then the estate gets context the ground cannot supply. A file knows who owns it, but only as a username. The company’s employee directory, Active Directory in most organizations, knows the rest. It has their department, their business unit, and whether they still work there. Match the two, and every file now carries that information alongside its own. Add the legal department’s hold list the same way, and the file also says whether its owner is under a legal hold. None of that was on the file to begin with, and that is what turns a file listing into an enriched metadata catalog.
And because Diskover indexes every storage system into one catalog, the survey covers the whole estate at once. One question reaches every vendor, tier, and location, so a file stays findable wherever it sits.
Ask Burns what two decades of excavation taught him about how much a company knows about its own data, and the answer is two words.
“Almost nothing. Storage teams care about how much space they have and how it’s allocated. Nobody owns the question of what the data is.”
He has had a customer estimate their own file count and be off by billions. And, as he puts it, “everyone has a retention policy; almost no one enforces it.”
Which of your shares could you describe right now by age, owner, and size, without opening a single file?
Defensible means someone will try to break it.
Burns has written expert reports and given deposition testimony, so the industry’s favorite adjective has a specific meaning for him.
Defensible is what survives a person whose job is to find the step you skipped.
“Opposing counsel goes after everything, down to the tiniest detail. It’s like having a screw turned into you to check that you really understood your own findings and make you defend every one of them.”
One exchange stayed with him. A client wanted him to testify that a particular person had done a particular thing. Opposing counsel asked: “Can you put that person in the chair? Can you put them behind the keyboard?” He had no video of anyone at a computer. What he had was activity in the data that traced back to that user, and that is what he said. State what the evidence supports, and stop there.
The same standard governs deletion. Legal teams hesitate because they have lived through litigation, and the hesitation is rational. The way through it is a protocol, and the order reverses how most programs start. Rather than classifying every file and working out its retention schedule, find the paths where business records are genuinely stored. Often there are a few dozen. Tag those and mark them protected. Holds then apply to the files and paths owned by the custodians under hold, so legal can see exactly what is covered. The rest of the estate falls under the policy the company already wrote. Sign-off comes before removal. Legal and the data owners approve what is protected, what is under hold, and what is eligible, and every file removed is logged with the rule that removed it.
Defensible deletion becomes a rule applied at scale, not a judgment call made file by file.
He bought the tool before he became part of it.
Burns was a Diskover customer before he was a Diskover employee. He ran it inside EY engagements from around 2020, while also helping clients evaluate and buy the other tools on the market. That gave him an unusual view of what the software was worth.
“I’d used tools that did less, had more errors, and cost several times as much.”
One engagement made the case better than any evaluation could. A customer had a data-loss incident and no idea how far it reached. Because they had an index from before the event and another from after, Burns pinpointed exactly which folders and files were affected. They restored those, and only those, instead of the whole volume.
He had been sending the team product feedback for years, and much of it ended up in the product. Joining Diskover meant putting the ideas in directly.
The people who know the data are the ones who delete it.
At EY’s forensic and integrity services practice, Burns spent eight years doing governance work for Fortune 500 clients. At an agriculture company, he disposed of a couple hundred TB and millions of files, and he expected the fight that normally comes with it. It never came. The client was unusually well organized, and its data owners were specialists who understood their own data. What they had never had was a way to see it. He profiled the estate and built dashboards showing what was there, how old it was, and who owned it. They read them, marked what was obsolete, and actioned it themselves.
The first find arrived almost immediately: 150 TB of old security footage that was supposed to have been migrated and was forgotten instead. Assessed, confirmed outside retention, removed, with a record.
Most programs move slower. Run a small phase, show the result, then work from easy capacity first to risky decisions second. Getting a client comfortable with that loop can take a year, and after that it runs as ordinary business. At one client, Burns says, 40% of the data went, and in his words, a year later “nobody had said I’m missing a file.”
“We don’t often see people having success getting their clients all the way to deletion. Every client I’ve run the program with has.”
What breaks at scale is ownership, not technology. Across five business units, you can find the person who knows what a share is for. Across fifty, the data outlives the people who created it, so each unit needs a named owner with real authority. Burns says the business’s first job is to name, for every unit, one person who understands its retention and legal hold requirements and has the authority to act.
Survey the whole site, then open one trench.
Burns came from EnCase and FTK, where you preserve a drive with a forensic image, then examine it down to the byte: file contents, deleted fragments, system artifacts. Diskover reads metadata only. Coming from full forensic analysis, is that enough?
His answer separates two jobs.
- Forensic analysis (the trench) is a deep look at a few things, and it takes a long time.
- Metadata (the survey) is how you assess an entire estate quickly and narrow the scope to what deserves that attention.
The arithmetic makes the order obvious. In Burns’s experience, content-scanning a petabyte (PB) for personally identifiable information (PII) takes a year. A metadata pass over the same ground takes a fraction of that time, and he has watched a billion files index in a single weekend. The survey then points the expensive tool at the right ground, cutting the target from a hundred terabytes to a few.

He is equally plain about the limit. Metadata will not look inside a file for PII. It tells the content tool or the data loss prevention (DLP) team where to go, which is why security teams find it an easier yes: metadata only, read-only, inside their own controls. When a question needs certainty rather than speed, the two work together. A fast check flags likely duplicates, then a content hash confirms them on the tagged set alone.
Do not point artificial intelligence at unsurveyed ground.
Organizations are now aiming artificial intelligence (AI) at estates they have never mapped.
Burns puts duplication at 40% in one client’s cloud object store and as high as 67% at another. Push that into an AI pipeline, and you pay to process and store the same data twice, and the duplicates crowd out the answers you need.
“You can’t feed or govern what you’ve never seen.”
A surveyed estate flips that, and fast. Sitting on the Diskover index, the AI connector answers questions in plain language. The files, in other words, can now talk back. Burns asked it whether the estate held any export-controlled data. It wrote its own search terms, ran them, and came back with candidates. A privacy first pass worked the same way and produced a risk assessment without opening a file.
Each answer arrives with a recommended action attached, and the action waits. AI-assisted, human-approved is the whole arrangement: Diskover proposes, a person decides, and only then does anything happen to a file.
What will look outdated in five years.
Keeping everything forever. Burns has watched this curve once already: organizations kept every email until the volume forced retention on them, and unstructured data is on the same path. The second thing he expects to age badly is clicking through screens to ask a question an AI assistant can answer in one sentence.
Three things you can do this week, each of which produces a fact rather than an opinion:
- Search your file shares for file names containing “password.” Burns on the results: “We’ve never had a client that didn’t have them.”
- Pull your largest-files list and look at the outliers. That is where the fastest capacity win sits.
- Ask a room of your own people how many have gone back to a file more than two years old, and count the hands. When Burns asks, barely any go up.
Each one tells you something true about your own ground. As Burns puts it: “The best time to start managing your unstructured data was yesterday. The second-best time is today.”
KEY TAKEAWAY
Evidence beats assumption.
Every estate holds something its owners would not predict: copies of copies, forgotten backups, passwords sitting in plain text.
Reading the metadata first tells you what is actually there. A documented rule and process for what goes is what makes removing it defensible.
Start with one search and one list.
Know the ground before you dig.
Diskover indexes every file across every storage system into one catalog, so your team can see what is there, show what it is, and act on it with a record.
