Skip to main content
Microsoft Solutions Partner — Data & AI · 11,000+ engagements

Microsoft Purview Data Catalog & Data Map Enterprise Hub (2026)

The Microsoft-native unified data catalog — Data Map auto-scan across 70+ source types, business glossary, Data Domains, Data Products, steward and owner workflows, end-to-end column-level lineage, Microsoft Fabric integration, and Copilot in Purview. Activated by a senior-architect-led Microsoft Solutions Partner founded in 1997.

What is Microsoft Purview Data Catalog and how do enterprises deploy unified data discovery, classification, and lineage? Microsoft Purview Data Catalog is the Microsoft-native unified data catalog spanning the Data Map scan plane (70+ source connectors), business glossary, Data Domains for federated ownership, Data Products for curated consumption, steward and owner workflows, AI-driven classification, end-to-end column-level lineage into Microsoft Fabric and Power BI, and Copilot in Purview for natural-language search. Enterprises deploy it through a five-phase Assess, Data Map, Glossary + Domains, Lineage + Products, Operate program that turns a metadata dump into a managed asset — and stands up the catalog layer that is the prerequisite for label-aware Microsoft 365 Copilot grounding in regulated industries.

Microsoft Purview Data Catalog is the unified Microsoft data catalog — Data Map scan plane across 70+ sources, business glossary, Data Domains, Data Products, steward + owner workflows, end-to-end column-level lineage, native Microsoft Fabric integration, and Copilot in Purview for natural-language discovery. EPC Group activates the full catalog under a fixed-fee five-phase accelerator and supplies the steward operating model that prevents the post-launch catalog-rot failure mode that kills most monolithic catalog programs.

Key Facts

  • Data Map scans 70+ source types — Fabric, Synapse, Snowflake, Databricks, on-prem SQL, Oracle, SAP, Salesforce, S3, BigQuery
  • 200+ system-defined sensitive-information classifiers plus trainable classifiers and customer custom regex
  • Business glossary, Data Domains, and Data Products operationalize the data-mesh pattern inside one Purview tenant
  • Column-level lineage native for Fabric, Synapse, Power BI, Azure SQL, Snowflake, Databricks, dbt; OpenLineage for the rest
  • Copilot in Purview delivers natural-language catalog search with citation back to source assets
  • Sensitivity labels propagate from catalog through Fabric lineage into Power BI exports and Copilot grounding
  • EPC Group five-phase Accelerator delivers full activation in 12 to 20 weeks, fixed-fee $150K to $600K
  • Microsoft Solutions Partner founded in 1997, 70+ Fortune 500 clients, 216+ M&A tenant consolidations

The six Microsoft Purview Data Catalog components — what each does and who owns it

Microsoft Purview Data Catalog is one product surface that spans six tightly integrated components — the Data Map scan plane, the business glossary, the Data Domain federation primitive, Data Products, the steward and owner accountability fabric, and AI-driven classification including Copilot in Purview. Each component has its own owner archetype; deployment success depends on naming a human accountable for each before go-live.

Purview Data Map — automated scan across 70+ source types

What it does: Data Map is the scan-and-catalog plane. Connectors cover Microsoft Fabric OneLake, Azure SQL, Azure Synapse, Azure Data Lake Storage Gen2, Azure Cosmos DB, Azure Database for PostgreSQL and MySQL, Power BI workspaces, Dataverse, on-premises SQL Server, Oracle, Teradata, SAP S/4HANA, SAP HANA, SAP ECC, Snowflake, Databricks Unity Catalog, Amazon S3, AWS RDS, Amazon Redshift, Google BigQuery, Google Cloud Storage, MongoDB Atlas, Salesforce, ServiceNow, Hive Metastore, Looker, Tableau, Erwin, and 40+ other sources. Scans run on schedule, classify against 200+ system-defined patterns plus customer-authored custom classifiers, and emit technical metadata plus end-to-end lineage to the unified catalog.

  • Capacity-based provisioning sized to source-system footprint — start at 1 capacity unit, scale on consumption
  • Self-hosted integration runtime (SHIR) for on-premises and VNet-isolated source scanning
  • Incremental scans on schedule — full scan weekly, delta scan daily, deep classification monthly
  • 200+ system-defined sensitive information types plus custom regex and dictionary classifiers
  • Multi-cloud connectors for Snowflake, Databricks Unity Catalog, BigQuery, Redshift, S3, and Cloud Storage

Owner archetype: Data platform owner, catalog admin, regulatory compliance officer

Business glossary — semantic layer above the catalog

What it does: The business glossary is the semantic layer above the technical catalog. Glossary terms — Customer, Patient, MNPI, PHI, Net Revenue, Loss-Given-Default — carry definitions, parent-child hierarchies, acronyms, related terms, stewards, owners, expert reviewers, and approval workflow status. Terms bind to physical assets (tables, columns, files, Power BI semantic-model fields) and bind upward to data domains. This is the layer a business stakeholder browses; the technical catalog is the layer an engineer browses.

  • Bulk import via CSV from existing glossaries — ASUG, DAMA, customer-authored taxonomies
  • Term-to-asset binding for tables, columns, files, Power BI semantic-model fields, Fabric items
  • Multi-stage approval workflow — author → steward → expert → owner → published
  • Synonyms, acronyms, related terms, parent-child hierarchy, and contextual definitions
  • Glossary adoption metrics — terms defined, terms bound, terms in steward backlog

Owner archetype: Data steward, business analyst, regulatory subject-matter expert

Data Domain — federated catalog ownership by business unit

What it does: Data Domain is the federated organizing unit inside Purview Unified Catalog. Each domain (Finance, Sales, Clinical, R&D, Supply Chain) has its own owner, stewards, data products, glossary terms, and policies. Domains are the data-mesh primitive — they let one Purview tenant host multiple business-unit catalogs without forcing a single global glossary or a single global ownership model. This is the answer for an enterprise that has tried and failed at a top-down monolithic catalog.

  • Domain-scoped glossary, data products, policies, and stewardship
  • Domain owner approval workflows for cross-domain term publication
  • Per-domain role-based access — domain admin, data product owner, steward, reader
  • Federated discovery — searches surface results across all domains the user has read access to
  • Foundation for the Pattern 3 data-mesh deployment described below

Owner archetype: Chief Data Officer, business-unit data leader, data mesh architect

Data Products — curated, discoverable, business-grade data assets

What it does: A Data Product is a curated bundle of assets (tables, columns, Power BI semantic models, dataflows, files, KQL queries) packaged as a single discoverable unit with documentation, SLAs, owner, quality metrics, and access-request workflow. Where the Data Map answers "what data exists?" Data Products answer "what data can I rely on?" Discovery, request access, and consumption all flow through the Data Product, not the raw asset.

  • Documentation, sample queries, schema contracts, SLAs, freshness commitments — all attached to the product
  • Owner-approved access requests routed through Purview policy and provisioned in source systems
  • Quality scorecards from Microsoft Purview Data Quality engine bound to the product
  • Product version history and deprecation lifecycle with reader-impact notification
  • Foundation for self-service analytics — analysts shop the catalog, request the product, and consume

Owner archetype: Data product owner, analytics consumer, citizen data scientist

Steward + Owner workflows — accountability fabric

What it does: Steward and Owner are the accountability roles that turn the catalog from a metadata dump into a managed asset. Owners hold business accountability for a data product, domain, or glossary term; stewards hold operational accountability for quality, classification accuracy, lineage completeness, and access decisions. Purview routes approval requests, quality alerts, classification disputes, and access requests through these roles via in-product workflow and Microsoft Teams notifications. This is the layer that fixes the "catalog rots after launch" failure pattern.

  • Per-asset steward and owner assignment with optional delegate chain for vacation coverage
  • Approval workflows for glossary publication, classification change, lineage edit, access grant
  • Teams notifications routed to the steward channel with one-click approve / reject / escalate
  • Steward dashboard surfacing backlog, SLA breaches, quality alerts, and pending access requests
  • Steward performance metrics rolled into the Data Estate Insights executive view

Owner archetype: Data steward, data owner, data office program manager

AI-driven classification + Copilot in Purview

What it does: AI-driven classification combines 200+ system-defined sensitive-information types, machine-learning trainable classifiers (PHI clinical notes, MNPI deal memos, attorney-client privileged content), customer-authored regex and dictionary rules, and natural-language Copilot suggestions for glossary term binding, steward assignment, and classification rule authoring. Copilot in Purview is the natural-language search layer — "find regulated customer-PII columns in the Snowflake estate that lack a steward" — that turns the catalog into a queryable system rather than a navigation tree.

  • System-defined classifiers for SSN, NI number, credit card, IBAN, passport, driver license, ICD-10, NPI, MRN
  • Trainable classifiers for domain-specific content — PHI clinical notes, MNPI deal memos, CUI markings
  • Copilot natural-language search across the entire catalog with citation back to source assets
  • Copilot-suggested classifications and term bindings, reviewed by steward, approved or rejected with one click
  • Copilot-authored regex starter rules for customer-specific patterns — supplier IDs, member IDs, internal codes

Owner archetype: Catalog admin, classification analyst, data office automation lead

Lineage visualization

End-to-end lineage — system level for all sources, column level for the ones that matter

Lineage is the answer to the auditor question "where did this number come from?" and to the engineer question "what breaks if I change this column?" Purview delivers system-level lineage universally across all 70+ connectors and column-level lineage natively for the source types that hold most regulated content. The remaining sources are covered through OpenLineage emitters that EPC Group ships as part of Phase 4.

System-level lineage — universal

Every Data Map scan emits system-level lineage — which job populated which asset from which upstream source, on what schedule, with what success rate. This covers every connector including SAP, Salesforce, ServiceNow, Tableau, Looker, MongoDB Atlas, and the rest of the long tail. It is enough for impact analysis at the table-to-table grain.

Use cases: impact analysis for schema change, source-system retirement planning, M&A consolidation, GDPR DSAR fulfillment, audit response.

Column-level lineage — Fabric, Power BI, Snowflake, Databricks, dbt

Column-level lineage resolves the path of a single column — Patient_DOB, Customer_NPI, Trade_PnL — from source through every transformation into every downstream consumer. Native support covers Microsoft Fabric, Azure Synapse, Azure Data Factory, Power BI semantic models, Azure SQL, Snowflake via the Snowflake-Purview integration, Databricks Unity Catalog, and dbt via the dbt-Purview integration.

Use cases: regulated-column audit defense, Power BI report certification, AI-grounding-corpus PHI containment, Copilot sensitivity-label propagation chain.

For source types without native column-level support, EPC Group ships OpenLineage emitters as part of Phase 4 — Apache Airflow via the openlineage-airflow integration, Spark via Spline or the openlineage-spark integration, custom Python and Java code via the OpenLineage client libraries. The output is column-level lineage across the full estate, not just the Microsoft-native portion.

Six enterprise Data Catalog deployment patterns

Every Data Catalog engagement composes from six deployment patterns. Most enterprises run the regulated-discovery pattern as the entry point, then sequence Fabric Lakehouse cataloging and the data-mesh pattern in parallel. M&A inventory and GDPR DSAR fulfillment are typically operational by month four.

Pattern 1 — Regulated-industry sensitive-data discovery

The regulated-discovery pattern stands up Purview Data Map across every regulated source system — Epic Clarity, Cerner Millennium, Workday, ADP, SAP HR, Fabric and Synapse warehouses, Snowflake, on-premises SQL Server, the M365 estate, and the Fabric lakehouses — runs the classifier library against the regulated content (PHI, NPI, PII, PCI, MNPI, CUI, CJIS), and produces a sensitive-data inventory the audit committee will accept. EPC Group ships custom trainable classifiers for the specific PHI, MNPI, and CUI patterns the customer regulator cares about (HHS-OIG audit triggers, FINRA Rule 3110 surveillance, CMMC L2 CUI markings), the steward and owner assignment matrix per source system, and the remediation backlog that closes the discovery findings. This is the deliverable a Chief Privacy Officer takes to the audit committee in week 12.

Pattern 2 — M&A data inventory and consolidation playbook

The M&A pattern is the data-side complement to a tenant migration. EPC Group runs Purview Data Map against the acquired entity's source estate during diligence or the first 30 days post-close, produces a complete data inventory (tables, columns, sensitivity classifications, ownership, lineage to consumption), identifies redundant systems (two CRMs, three warehouses, four BI tools), surfaces compliance gaps (PHI flowing to an unencrypted on-premises share, MNPI in a Box account, customer NPI in an Excel attachment on personal email), and produces the consolidation backlog. With EPC Group's 216+ M&A tenant consolidations and 1.83M users migrated, this pattern is repeated dozens of times per year. Cross-link to the M&A playbook hub at /microsoft-ma-90-day-tenant-consolidation-playbook for the broader tenant story.

Pattern 3 — Data mesh with Data Products and Data Domains

The data-mesh pattern uses Data Domain and Data Products to federate ownership across business units while keeping a single discovery surface. Finance, Sales, Clinical, R&D, Supply Chain each own their domain, their glossary, their data products, and their access policies; the federated catalog lets a Finance analyst discover and request access to a Clinical data product through one search, with the access decision flowing to the Clinical domain owner. This is the deployment pattern for enterprises that have tried and failed at a monolithic top-down catalog. EPC Group ships the domain taxonomy, the data-product template library, the federated governance operating model, and the steward training curriculum.

Pattern 4 — Microsoft Fabric Lakehouse cataloging end-to-end

The Fabric Lakehouse pattern catalogs the full Fabric estate — workspaces, lakehouses, warehouses, KQL databases, eventhouses, semantic models, reports, dataflows, notebooks, pipelines — with end-to-end column-level lineage that resolves from Fabric ingestion (Dataflow Gen2, Data Factory pipeline, Eventstream) through transformation (Spark notebook, T-SQL stored procedure, KQL update policy) to consumption (Power BI report, Excel pivot, Copilot answer). Sensitivity labels propagate through the lineage — a PHI label applied at ingestion in the bronze lakehouse carries through silver and gold and into the Power BI report a clinician opens, with the encryption and rights protection traveling along. Cross-link to the broader Fabric story at /microsoft-fabric-expertise and /services/fabric-consulting.

Pattern 5 — GDPR data subject access request (DSAR) fulfillment

The GDPR DSAR pattern uses Purview Data Map and the business glossary to answer the regulator-mandated 30-day data subject access request without a fire drill. EPC Group builds the data-subject identifier map (email, customer ID, member ID, employee ID, claim ID — all mapped to the underlying source columns), the source-system query plan (which 17 systems hold data on a given subject), the fulfillment runbook (export, redact, package, deliver inside the SLA), and the right-to-be-forgotten propagation chain (delete in source → confirm in lakehouse → confirm in warehouse → confirm in Power BI semantic-model refresh). The same plumbing serves CCPA, LGPD, Quebec Law 25, and the state-level US privacy laws now in force across 19 states.

Pattern 6 — Healthcare HIPAA data inventory and BAA evidence

The healthcare HIPAA pattern produces the comprehensive PHI data inventory required for the Security Rule risk analysis (45 CFR 164.308(a)(1)(ii)(A)) and the documentation trail that satisfies the BAA obligations under 45 CFR 164.504(e). EPC Group catalogs PHI across the Epic, Cerner, Meditech, or Allscripts source environment, the Fabric and Synapse warehouses that feed Population Health and quality measure reporting, the Power BI semantic models clinicians and operations leaders consume, the M365 estate where PHI escapes into Word, Excel, and Teams, and the third-party SaaS connections (Salesforce Health Cloud, ServiceNow Healthcare, Smartsheet) that often hold uncatalogued PHI. Cross-link to /healthcare-it-consulting-hipaa-microsoft-2026 for the broader HIPAA Microsoft architecture and to /microsoft-purview-data-governance-enterprise-2026 for the broader Purview platform context.

Copilot in Purview

Copilot in Purview — natural-language search, classification suggestions, glossary co-authoring

Copilot in Purview is the natural-language layer over the catalog. Instead of building and saving a faceted search, an analyst asks: "Find regulated customer-PII columns in the Snowflake estate that lack a steward." Copilot resolves the query against the Data Map, the glossary, the steward assignment matrix, and the classification index — and returns the answer with citations back to specific assets. The same surface co-authors classification rules ("suggest a regex for our internal member-ID pattern based on these 50 examples"), suggests glossary term bindings ("which technical columns should this Customer term bind to?"), and accelerates the steward backlog. Copilot in Purview is included with Microsoft 365 Copilot licensing for users who hold a Purview catalog role.

Natural-language data search

Ask the catalog in plain English. Copilot resolves against Data Map, glossary, and classification index and returns ranked results with asset citations.

Classification suggestions

Copilot proposes classifications and glossary term bindings based on the column profile and the existing taxonomy. Steward reviews and approves or rejects with one click.

Glossary co-authoring

Draft new glossary terms with Copilot, including definitions, synonyms, and candidate asset bindings. Route through the steward approval workflow.

Pricing — pay-as-you-go capacity plus per-Data-Product SKU

Microsoft Purview Data Catalog is metered on Azure consumption. The catalog plane bills on capacity units per hour, scans bill on data scanned (per MB or per vCore-hour), classification bills on classified assets, and the Unified Catalog Data Products SKU bills per published Data Product. The Microsoft 365 E5 Compliance licensing customers already own covers sensitivity labels, DLP, eDiscovery, IRM, and the Records Management stack; Data Map is the Azure-billed exception.

Mid-market deployment — $4K to $8K per month Azure consumption

Scanning 30 to 50 source systems, 250K to 500K cataloged assets, weekly full scans plus daily delta scans, 1-2 capacity units of provisioned Data Map, classification on the regulated content types, and 10 to 25 published Data Products. Total Azure consumption typically runs $4K-$8K per month plus the existing M365 E5 Compliance licensing.

Enterprise deployment — $15K to $45K per month Azure consumption

Scanning 200 to 500+ source systems, 2M to 10M+ cataloged assets, the full Data Products SKU with 100+ published products, multi-region Data Map capacity, and customer-authored trainable classifiers across the regulated content library. Total Azure consumption typically runs $15K-$45K per month, with the cost-attribution model tying spend back to the Data Domain consuming it.

EPC Group accelerator — fixed-fee $150K to $600K

The five-phase Data Catalog Accelerator described below is delivered fixed-fee between $150,000 and $600,000 depending on source-system count, regulatory scope, domain taxonomy complexity, lineage-emitter scope, and managed-service tail. Senior-architect led, no offshore handoff, no T&M overruns.

The EPC Group Data Catalog Accelerator — five phases, fixed fee

The accelerator anchors on The EPC Group Lifecycle — Assess, Data Map + Scan, Glossary + Data Domains, Lineage + Data Products, Operate. Fixed-scope between $150,000 and $600,000. Senior-architect led, named on-record from kickoff through go-live, no offshore handoff, no T&M overruns.

Phase 1 — Assess

Data estate inventory and catalog readiness in three weeks

Phase one inventories every source system in scope, every existing metadata or catalog tool the customer has tried (Alation, Collibra, Atlan, Informatica EDC, data.world, Erwin Data Intelligence Suite, SAP DataSphere catalog), every regulatory driver (HIPAA, FINRA, GLBA, GDPR, CCPA, FedRAMP, CMMC, GxP), and every active failure mode (monolithic top-down attempt that rotted, classification false-positive flood, steward backlog never cleared). EPC Group ships a costed roadmap, a domain taxonomy proposal, a steward and owner staffing plan, and a board-ready decision package anchoring on the Assess stage of the EPC Group Lifecycle.

  • Source-system inventory across Fabric, Synapse, Snowflake, Databricks, on-premises, SAP, Salesforce
  • Existing-tool teardown — Alation, Collibra, Atlan, Informatica, data.world, Erwin, SAP DataSphere
  • Regulatory driver mapping — what each regulator requires, how Purview satisfies it
  • Data Domain taxonomy proposal scoped to the customer org structure and data-mesh appetite

Phase 2 — Data Map + scan

Stand up the catalog plane with classifier library

Phase two provisions Data Map capacity, deploys the self-hosted integration runtime for VNet and on-premises sources, configures connectors and scan schedules across the source estate, and tunes the classifier library — system-defined plus trainable plus customer custom — to a defensible false-positive rate. EPC Group ships the scan-health monitoring runbook, the classification review workflow, and the steward escalation path for false-positive disputes. The deliverable at the end of Phase 2 is a populated catalog with classifications a steward will sign off on.

  • Data Map capacity provisioning sized to source-system footprint, with scaling alarms
  • Self-hosted integration runtime deployed for on-premises and VNet-isolated sources
  • Connector configuration for Fabric, Synapse, Snowflake, Databricks, on-prem SQL, Oracle, SAP, Salesforce
  • Classifier library tuned to <5% false-positive rate on customer-specific regulated content

Phase 3 — Glossary + Data Domains

Semantic layer and federated ownership operating model

Phase three stands up the business glossary and Data Domain structure. EPC Group authors the initial glossary import covering the regulated terms the customer's industry requires (PHI, NPI, MNPI, PCI, CUI, attorney-client privileged) plus the customer-specific business terms (Net Revenue, Loss-Given-Default, Patient Day, Active Subscriber, Daily Active User), defines the Data Domain taxonomy, assigns initial domain owners and stewards, and trains the steward community on the approval workflows. This is the phase where the catalog acquires meaning.

  • Glossary import covering 200-500 initial terms with definitions, stewards, owners, hierarchy
  • Data Domain taxonomy operationalized with domain owners and per-domain stewards
  • Approval workflows configured and tested for term publication, classification change, lineage edit
  • Steward training curriculum delivered, with one-week shadow program before steady-state handoff

Phase 4 — Lineage + Data Products

End-to-end lineage and curated data products

Phase four delivers end-to-end column-level lineage across the catalog plus the first wave of curated Data Products. Lineage resolves from ingestion (Fabric Dataflow Gen2, Azure Data Factory, Databricks job, dbt run, Stored Procedure) through transformation (Spark notebook, T-SQL, KQL, Power Query) to consumption (Power BI semantic model, Tableau workbook, Excel pivot, Copilot grounding). Data Products are authored for the highest-leverage business analytics — customer 360, patient 360, sales pipeline, supply chain visibility — and routed through owner-approved access requests.

  • Column-level lineage validated end-to-end across Fabric, Synapse, Snowflake, Databricks, Power BI
  • Five to ten curated Data Products authored, with documentation, SLAs, and access workflow
  • Quality scorecards bound to Data Products via Microsoft Purview Data Quality engine
  • Self-service consumption pattern operationalized — analyst search → product request → grant → consume

Phase 5 — Operate

Managed catalog with senior-architect escalation

Phase five is steady-state operation. EPC Group provides managed catalog services — classifier tuning, scan-health monitoring, glossary evolution, steward coaching, Data Product backlog curation, quality alerting, and quarterly governance steering committee output. Senior-architect escalation is the differentiator; tier-one analysts handle routine cases, but every customer has named senior architects on call for the cases that matter — a regulator inquiry, an M&A diligence, a Copilot rollout gate.

  • Monthly catalog-health report — scan coverage, classification accuracy, steward backlog, lineage completeness
  • Quarterly governance steering committee output with Chief Data Officer review pack
  • Quarterly regulatory change review — new state privacy laws, sector regulator changes, Fabric updates
  • Senior-architect on-call escalation tied to data-governance incident severity matrix

Continue exploring the EPC Group enterprise Microsoft library

Purview Data Catalog sits inside the broader Microsoft data and AI governance plane. These hubs and analyses cover adjacent and complementary territory.

Why EPC Group leads enterprise Data Catalog deployments

1997
Founded · Microsoft consulting
70+
Fortune 500 clients
216+
M&A tenant consolidations
1.83 million
Users migrated

Microsoft Solutions Partner — Data & AI

Microsoft Solutions Partner with the Data & AI and Security designations plus four additional designations across Modern Work, Infrastructure, Digital & App Innovation, and Business Applications. Senior architects average two decades of Microsoft platform delivery experience.

Four-time author for Microsoft Press and Sams

Founder Errin O’Connor has nearly three decades of Microsoft consulting leadership and is a four-time author for Microsoft Press and Sams across Power BI and SharePoint.

Fixed-fee accelerators

Every Data Catalog engagement is fixed-fee with a costed roadmap and a named senior architect on-record from kickoff through go-live. No T&M overruns, no offshore handoff, no junior-analyst-led production cutover.

Compliance-native

EPC Group is compliance-native across HIPAA, SOC 2, FedRAMP-aligned, FINRA, CMMC, and GxP. Data Catalog deployments ship with auditor-ready Compliance Manager assessments, sensitive-data inventories, and exception-management workflows.

Regulatory and standards coverage

HIPAA
SOC 2
FedRAMP
FINRA
CMMC
GxP

Frequently asked questions — Microsoft Purview Data Catalog

How does Microsoft Purview Data Catalog compare to Alation?

Alation is the legacy enterprise data-catalog incumbent, built before the lakehouse era around a query-log-driven Behavioral Analysis engine that infers popularity, stewardship, and lineage from query traffic. Purview Data Catalog is the Microsoft-native unified catalog built around Data Map scans, the business glossary, Data Domains, Data Products, and Copilot-in-Purview natural-language search. The decision is rarely feature-by-feature parity; it is platform alignment. Customers whose data estate centers on Microsoft Fabric, Azure Synapse, Power BI, Microsoft 365, and Snowflake or Databricks on Azure consistently land on Purview because lineage and sensitivity-label propagation flow natively into Fabric and Power BI. Customers with heavy Tableau, Snowflake-on-AWS, Looker, and a 10-plus-year Alation investment often keep Alation. EPC Group has migrated 14 enterprises off Alation to Purview in the last 24 months — the trigger is almost always a Fabric or Power BI lineage-integration requirement Alation could not deliver natively.

How does Microsoft Purview Data Catalog compare to Collibra?

Collibra is the workflow-and-governance heavyweight — strong on stewardship, approval workflows, and the integration with external systems like ServiceNow and Jira for change management. Purview is now competitive at the workflow layer thanks to Data Domains, steward + owner role primitives, and the Teams-native approval notifications. The differentiation is licensing and integration: Collibra is a separate enterprise contract often running $300K-$1M+ per year, whereas most Purview Data Catalog capabilities are bundled into the broader Microsoft Purview platform agreement with Data Map metered by Azure consumption. For a Microsoft-aligned enterprise, the total-cost-of-ownership math has shifted enough that Collibra renewals are increasingly contested. The exception is highly federated multi-cloud enterprises with non-Microsoft analytics platforms at the center — Collibra still leads there.

How does Microsoft Purview Data Catalog compare to Atlan?

Atlan is the modern, design-led, dbt-native catalog favored by analytics-engineering teams who live in dbt, Snowflake, BigQuery, and Looker or Mode. Atlan's strengths are dbt-model-as-first-class-citizen, GitHub-pull-request-based glossary review, and a UX that analytics engineers genuinely use. Purview is now reasonably competitive at the UX layer post the unified-catalog redesign, and Purview's edge is sensitivity-label propagation through Microsoft Fabric and Power BI plus the regulated-industry classifier library. For a Fabric-centric enterprise the Purview integration is decisive; for a dbt-and-Snowflake-centric mid-market analytics team, Atlan remains the better fit. EPC Group sometimes recommends a hybrid where Atlan owns the analytics-engineering catalog and Purview owns the Microsoft Fabric and M365 estate with federation between the two.

How does Microsoft Purview Data Catalog compare to data.world?

data.world is the knowledge-graph-native catalog with strong semantic-layer ambitions and a SPARQL query interface. data.world's strengths are linked-data semantics, the ability to model complex many-to-many relationships across the catalog, and the public-data catalog community. Purview is more pragmatic, more directly tied to operational source-system scanning, and natively integrated with Microsoft Fabric and Power BI lineage. For an enterprise that wants a semantic knowledge graph over the catalog and is willing to invest in linked-data modeling, data.world remains differentiated. For the typical Fortune 500 trying to inventory the regulated estate, satisfy GDPR DSARs, and stand up label-aware Copilot grounding, Purview is the faster and better-integrated answer.

How accurate is Purview classification, and how do you tune false positives?

Out-of-the-box classification with the system-defined sensitive-information types runs at roughly 70% to 85% precision depending on the source content profile — strong on PCI and SSN, weaker on free-form clinical notes or deal memos. The tuning workflow is the false-positive review cycle: a steward reviews a sample of classification matches per scan run, marks false positives, and Purview retrains the trainable classifiers on the labeled feedback. EPC Group standard delivery targets sub-5% false-positive rate on the regulated content types most relevant to the customer (PHI, MNPI, CUI, PCI) within 60 days of Data Map go-live. Custom trainable classifiers authored against customer-specific patterns (member IDs, supplier codes, internal proprietary markings) typically reach 90%+ precision after two training cycles.

How deep is Purview lineage, and does it cover column-level for all sources?

System-level lineage (which job populates which table from which source) is universal across all 70+ connectors. Column-level lineage is supported natively for Microsoft Fabric, Azure Data Factory, Azure Synapse pipelines, Power BI semantic models, Azure SQL, Snowflake (via the Snowflake-Purview integration), Databricks Unity Catalog, and SQL Server. Column-level lineage for dbt is supported through the dbt-Purview integration. Column-level lineage for sources without native support is delivered through OpenLineage emitters that Purview ingests — Apache Airflow, Spark via Spline, and custom code via the OpenLineage Python and Java clients. EPC Group ships OpenLineage emitters for the customer's remaining lineage gaps as part of Phase 4 of the accelerator.

What does Purview Data Map cost, and how is it priced?

Purview Data Map is metered on Azure consumption — capacity units per hour for the catalog plane, scan vCore-hours for active scans, and per-data-product pricing for the Unified Catalog Data Products SKU. A mid-market deployment scanning 50 source systems and holding 500K assets in the catalog typically runs $4K-$8K per month in Azure consumption plus the Microsoft 365 E5 Compliance licensing customers already own for sensitivity labels and DLP. Enterprise deployments with 500-plus source systems, multi-million-asset catalogs, and the full Data Products SKU run $15K-$45K per month. EPC Group ships the consumption forecast, the capacity-scaling alarms, and the cost-attribution model that ties Azure spend to the data domain consuming it.

Why is Purview Data Catalog the prerequisite for label-aware Copilot in regulated industries?

Microsoft 365 Copilot and Copilot Studio agents honor sensitivity labels on every piece of grounding content they retrieve from Microsoft Graph and the Fabric estate. But labels only protect what the catalog knows about — content that has not been scanned, classified, and labeled by Purview Data Map sits outside the protection envelope. The Purview Data Catalog is the prerequisite layer because it discovers the regulated content, classifies it accurately, propagates the sensitivity label to the source asset, and then propagates the label through downstream lineage into Fabric, Power BI, and the Microsoft Graph corpus Copilot grounds against. Without that chain, Copilot deployment in a regulated industry is an unredacted exfiltration risk. With it, Copilot is a compliant productivity multiplier. Cross-link to /microsoft-purview-data-governance-enterprise-2026 for the broader Purview platform context and to /microsoft-purview-data-loss-prevention-insider-risk-2026 for the DLP and Insider Risk plane that pairs with the catalog.

Make the catalog canonical — and make Copilot safe to ground on it

Book a Data Catalog briefing with an EPC Group senior architect. Two-hour working session — source-system inventory, existing-tool teardown (Alation, Collibra, Atlan, data.world, Informatica), Fabric and Power BI lineage gap analysis, classifier-library scoping, Data Domain taxonomy review, and accelerator scoping. Zero obligation, board-ready output.

AI assistant — not human