BLOG
Extending Snowflake Horizon Catalog to the enterprise unstructured data estate:
metadata first, content on demand.
August 6, 2026 · 5.5 min read
Snowflake Horizon Catalog has a clear ambition: one catalog where every data and AI asset is discoverable, governed, and ready to produce trusted answers. Tables, apps, models, listings, even data in external catalogs.
The last frontier is unstructured data. Roughly 80% of enterprise data lives in files and objects, most of it on-prem. Media footage, chip design, medical imaging, contracts, sensor data, spread across NAS, object stores, and archive tiers from a half-dozen storage vendors. That data has never been cataloged anywhere. Not by a warehouse, not by a storage vendor, not by anyone. It was never mapped at the source.
For years, the answer was “copy it all in.” Nobody does it. It’s petabytes, it changes daily, and most of it isn’t worth moving. So the biggest slice of the data estate stays off-catalog, and out of reach of the AI built on top.
The fix isn’t a bigger pipe. It’s a smarter map, delivered to the catalog. Snowflake brings the governed platform and the AI capabilities. Diskover brings the map of the file estate. Together, they put the entire data estate under one catalog.
A map, not a copy: the global metadata engine.
Diskover indexes every storage and cloud tier from every vendor and builds one thing from it: a global metadata engine that lives as a logical abstraction layer above the storage. One searchable catalog of every file and object the enterprise owns, across NAS, cloud, object, and archive, enriched with business context like owner, project, age, usage, and custom attributes.
That catalog is a map of the data estate, not a copy of it. Nothing gets moved to build it. And it’s a powerful complement to the robustness of the Horizon Catalog: the unstructured estate, described in rows – in Snowflake.
A protocol shares what you point it at. A catalog is how you know what’s worth pointing at.
See it, curate it, deliver it.
It starts with visibility. Harvesting metadata from dozens of storage systems doesn’t make data actionable on its own. The global view is what lets the business assess which data is valuable, find it fast, and organize it in terms knowledge workers actually use: project, owner, age, usage. Without that layer, you’re just relocating a mess.
Delivery then works in two lanes.
First, curated file metadata lands in Snowflake as ordinary queryable rows, on a schedule. Teams define what qualifies by file type, age, size, path, or owner, so Snowflake receives only what workloads need. Curated, not dumped.
Second, file content moves only when something asks for it. Snowflake fetches files on demand through the Diskover REST API, from a single document to a bulk load. Cortex Search and Cortex Agents get exactly the content they need.
For your teams, that means an analyst can ask Snowflake CoWork a question and get an answer grounded in documents that were invisible last quarter, and no data engineer built a pipeline to make it happen.
Under the hood, scheduled scans stream the curated metadata into your Snowflake account via Kafka. No custom pipelines, no per-storage scripts, no connector sprawl. That’s the whole technical story.
The economics follow the architecture. Because metadata leads and content follows the query, ingest cost stays controlled from day one, and AI readiness drops from months of plumbing to minutes. And because none of this is a new protocol, there is nothing for your storage vendors to adopt, no native integration to build, and no partner list where your vendor is marked “coming soon.” It works across every tier, every vendor, today.
One catalog, from tables to files.
Here’s what changes once file metadata lives in Snowflake: it stops being a storage inventory and becomes a governed asset.
The same Horizon Catalog toolset that governs tables now applies to the unstructured estate. Classification and tagging flag sensitive file data. Access policies control who sees which files and their metadata. Discovery puts forty years of file history next to last quarter’s revenue tables. And when an agent reaches for file content, it does so under the same guardrails as every other query.
In the agentic era, that’s the whole game. An agent is only as governed as the data it touches, and a share is only as trustworthy as the catalog behind it. Extending the catalog to files means extending trust to the 80% that never had it.
The correlation is the payoff.
Metadata alone is an inventory. Content alone is a swamp. Together, inside Snowflake, joins become possible that no storage tool and no warehouse can deliver alone.
Unstructured data, joined to business data. Structured tables, semi-structured logs and JSON, and file metadata, all queryable side by side. A semiconductor design team joins job telemetry with the design files behind each run, turning failure investigations that took weeks into queries. A major studio links its production systems to petabytes of footage and project files, so every asset carries its business context.
Files, joined to files. One queryable estate across every storage brand and tier. No single storage vendor can offer this view, because each one only sees its own systems.
Storage economics, joined to data value. Capacity, growth, and redundant, obsolete, and trivial data become Snowflake tables, joined to ownership and usage. Storage teams get cost analytics on a platform the business already trusts, for a problem they already have budget to solve.
Keep the storage you have. Gain the platform you want.
The question every data leader asks first: do we have to migrate all of it? No. Your data never leaves the storage you chose. You keep the performance you bought and the capacity you’re growing. What changes is that the estate finally shows up in your catalog of record, governed and queryable next to everything else the business runs on.
Every team gets something out of that. Storage teams walk into AI and governance conversations they were locked out of. Data scientists and business teams reach data they could never see before. And nobody rips out anything to get there.
Proud to support Diskover Data as they help companies uncover their most valuable data across legacy systems with a unified, searchable view. Together with Snowflake’s easy, trusted, and connected platform, we’re helping customers seamlessly ingest critical data and build a strong, AI-ready foundation.
KEY TAKEAWAY
Start today.
Diskover is live on Snowflake Marketplace today, with direct streaming of curated file metadata and on-demand file content for Cortex AI and Snowflake CoWork. And as open sharing over Iceberg matures, the same curated estate becomes shareable across clouds and engines, with Horizon Catalog as the connective tissue.
Openness is a property of your data, not of a protocol announcement. Nothing here locks you in: the metadata is plain queryable rows, and the content never left the storage you own.
Your files already hold the answers, and you already own the storage they sit on. Put them where the catalog can see them.
Your estate already holds the intelligence. Bring it into view.
Diskover is live today on Snowflake Marketplace, with direct streaming of curated file metadata and on-demand file content for Cortex AI and Snowflake CoWork.

