Dataobservability

BUYER GUIDE

Data Lineage Tools Pricing: Data Lineage Cost and Data Catalog Pricing Compared

Almost every commercial lineage and catalog vendor quotes by sales call. Here is what each one actually publishes, what drives the number they will quote you, and how to build a defensible 3 year budget before you take the meeting.

See pricing

14-day trial, no credit card, read-only connection

SNOWFLAKE · PROD
247 tables |
Break a monitor:

Alerted #data-eng 0.8s ago.

Downstream impact · consumers at risk

INCIDENT #1042 OPEN · owner @you

How much do data lineage tools cost?

Data lineage tools fall into three price bands. Cloud native catalogs meter usage and publish real rates: the AWS Glue Data Catalog is free to one million metadata objects and one million requests a month, then charges 1 dollar per 100,000 objects a month, with crawlers and table statistics at 0.44 dollars per DPU hour. Commercial lineage and catalog platforms, including Atlan, Alation, Collibra, Secoda and Select Star, publish no figure at all and quote per named user or per governed asset after a discovery call. Open source lineage, meaning DataHub, OpenMetadata and Marquez, carries no license fee and costs engineering and hosting time instead. Every row below was checked on the vendor's own pricing page in August 2026.

Last updated August 2026

// COMPARE

Side by side

Data lineage tools pricing compared

Swipe to see all columns →

Tool Publishes a price? What is actually published (August 2026) Pricing model
Dataobservability Yes Starter 99 dollars, Team 299, Scale 799 per month, with column level lineage included on every tier and a 14-day trial that takes no card Flat monthly by tier
AWS Glue Data Catalog Yes Free to 1 million metadata objects and 1 million requests a month, then 1 dollar per 100,000 objects per month. Crawlers, statistics generation and Iceberg optimization each bill at 0.44 dollars per DPU hour Metered usage
Microsoft Purview Partly The meters and units are public (per governed asset per day, per capacity unit per hour, and Basic, Standard and Advanced processing units) but the dollar figures render as a dash on the pricing page itself and only resolve in the regional calculator. The first 1 MB of Data Map metadata storage is free Metered, priced by region
Atlan No The pricing page carries no tiers and no figures, only an invitation to start a conversation with sales Quote only
Alation No The pricing URL resolves to a contact form for an Alation expert. No tiers, no figures Quote only
Collibra No The pricing URL returns a 404. Nothing is published anywhere on the site Quote only
Secoda No Three named tiers, Core, Premium and Enterprise, with a full feature comparison. Every tier links to contact sales and no tier shows a figure Quote only
Select Star No The pricing page publishes no tiers, figures or limits Quote only
Informatica No No public figure for the catalog and lineage products Quote only
Datafold No The pricing link resolves to a contact page Quote only
DataHub Free to license Apache 2.0, no license fee. The managed DataHub Cloud offering is quote only Self host or quote
OpenMetadata Free to license Apache 2.0, no license fee. Collate is the managed offering and is quote only Self host or quote
Marquez Free to license Apache 2.0. The reference backend for the OpenLineage standard, no license fee Self host

Positioning and pricing models are summarized in good faith from each vendor's public pages, August 2026. Verify current terms with the vendor.

// CAPABILITY

What you get

What actually drives a data lineage quote

Named users is the usual catalog unit

Most catalog led lineage vendors price per named user per year, then split those users into editors and viewers at different rates. It sounds cheap until governance works the way it is supposed to and every analyst needs read access. Ask early whether read only seats are free, discounted or full price, because that single answer moves a quote more than any feature on the comparison sheet.

Governed assets is the other unit

The cloud native tools meter differently. Purview bills per unique governed asset per day plus processing units per run, and Glue bills per 100,000 metadata objects a month once you pass the free million. Under that model your bill tracks how many tables, columns and partitions exist, so a warehouse that generates partitions daily grows the bill without anyone adding a user.

Scan frequency is a compute bill, not a license bill

Lineage has to be extracted from somewhere, usually query logs or a scan. Scan heavy tools bill the compute to you: Glue crawlers and statistics run at 0.44 dollars per DPU hour, and a warehouse side scanner burns your own credits. Metadata first approaches that read query history avoid most of this. Ask which one you are buying, then price a week of it before you sign.

Column level lineage is usually an upgrade

Table level lineage is table stakes. Column level is where the money sits, and on several platforms it is gated behind the top tier. dbt puts column level lineage in Enterprise and Enterprise Plus only, and Snowflake exposes the ACCESS_HISTORY views that make it possible on Enterprise Edition and above. Confirm which tier carries it before you compare two quotes as though they cover the same thing.

Rollout and stewardship is the line nobody budgets

A catalog is only as good as the people curating it. Budget the implementation, the connector work, and the ongoing stewardship hours, because a lineage graph nobody maintains stops being trusted within about two quarters. On a quote only platform this is frequently a bigger number than year one license, and it recurs.

A published price removes the whole exercise

The reason all of this analysis exists is that twelve of the thirteen tools above will not tell you a number. Dataobservability publishes 99, 299 and 799 dollars a month with column level lineage on every tier, so the budget question takes about four seconds instead of four weeks of procurement.

// 4 STEPS

How it works

From connected to caught

01

Count what you would actually govern

Pull a table and column count from your warehouse information schema, and separately count the humans who need read access versus edit access. Those two numbers are what every quote is built from, and walking in with them stops the discovery call from becoming a discovery quarter.

02

Decide table level or column level up front

Column level lineage costs more everywhere, sometimes by a whole tier. If you need it because you are doing impact analysis on individual fields, say so before you get a table level quote you will have to renegotiate.

03

Price the compute separately from the license

Ask whether the tool scans your warehouse or reads its query history, then measure a week of the warehouse it would run against. Scan heavy tooling can cost more in credits than the subscription it came with.

04

Model three years, not one

Add year one license, implementation, connector work, and the stewardship hours per week times a loaded engineer rate. Then add the escalation clause from the contract for years two and three. That number is what you compare across vendors, and it is often ranked differently than the year one quotes are.

Why almost no data lineage vendor publishes a price

Of the thirteen tools checked in August 2026, exactly three publish a usable figure on their own site, and two of those are cloud provider meters rather than platforms. Atlan, Alation, Collibra, Secoda, Select Star, Informatica and Datafold all route pricing to a form. Collibra does not even keep a pricing page live, the URL 404s. This is a deliberate commercial choice, not an oversight. Value based pricing works best when the seller sets the anchor after learning your team size, your warehouse footprint and your compliance deadline, and a public number would cap the quote for the largest buyer while scaring off the smallest. The practical effect for you is that comparison shopping is genuinely hard, every evaluation costs weeks of calls, and two quotes for the same product can differ by a multiple depending on what the buyer disclosed and when their fiscal year ends.

What Microsoft Purview actually publishes, which is less than people assume

Purview gets described as the transparent option because it is a cloud meter rather than a sales motion, and that is only half right. The pricing page does publish the structure honestly: Unified Catalog bills per unique governed asset per day, the Data Map bills per capacity unit per hour where a capacity unit covers 25 operations per second and 10 GB, and enterprise data management bills per data governance processing unit across Basic, Standard and Advanced. The first 1 MB of Data Map metadata storage is free. What the page does not do is show the dollar amount. Checked in August 2026, the rate cells render as a dash and the actual figure only appears through the regional calculator, because the rate varies by region. So Purview is more predictable than a sales quote and still not a number you can put in a budget from the pricing page alone.

The open source lineage bill is real, it just lands on a different budget line

DataHub, OpenMetadata and Marquez are Apache 2.0 and cost nothing to license. Marquez is the reference backend for OpenLineage, the open lineage standard governed under LF AI and Data, and it is the cheapest credible way to get lineage events flowing if you already emit them. The cost shows up in three places. Someone hosts the metadata store and keeps it upgraded. Someone writes and maintains the ingestion connectors for every source, and connector coverage is exactly where open source catalogs are thinnest. And someone curates, because an uncurated catalog decays into a stale table list that nobody trusts. Price those hours at a loaded engineer rate and compare the annual figure honestly against a subscription. For a small stack with a handful of critical sources, self hosting genuinely wins. Past a few hundred tables and a few dozen connectors, the maintenance line usually crosses a self serve subscription inside the first year.

Coverage limits that quietly change what you are paying for

Two quotes that both say column level lineage may not be buying the same coverage, because the underlying platforms have documented gaps. dbt column level lineage fails on complex SQL, Python models, JSON unpacking, lateral joins, and any model that hardcodes a table name instead of using ref. Databricks Unity Catalog captures column level lineage automatically but holds nothing before September 2024, does not preserve lineage through renames, and does not capture RDDs, global temp views, Spark SQL checkpointing or UDFs. BigQuery tracks top level columns only, skips load jobs and routines, gives no upstream column lineage for external tables, and falls back to table level above 1,500 column links in a job. Redshift has no native column level lineage graph at all. None of that is a reason to avoid these tools, but it is a reason to ask a vendor which of your specific transformation patterns they resolve before you price the tier that claims to cover them.

A worked three year model you can take into procurement

Take a mid sized case: 800 tables, 25 people who need read access, 6 who curate, and a warehouse already modeled in dbt. On a per user catalog platform at a typical enterprise motion you are pricing 31 seats, an implementation engagement, connector work for the sources the catalog does not cover natively, and stewardship. Assume the license escalates annually under the contract clause, because it almost always does. On a metered cloud catalog you are instead pricing 800 tables worth of governed assets every day, plus the crawler or scan compute, which scales with partition count rather than headcount. On a self serve platform with a published price you are pricing the tier and nothing else. The ranking between these three flips depending on headcount and table count, which is exactly why a year one quote is a bad comparison instrument. Build the three year number for all three shapes before you decide, and insist that any quote you receive is expressed in the same shape.

Where a monitoring first tool changes the calculation

Catalog platforms sell lineage as a governance artifact, something people browse. That is a real need and it is priced like an enterprise seat product. A monitoring first tool treats lineage as the thing that makes an alert actionable: when a table breaks, the graph tells you which dashboards and models are downstream, so the on call engineer knows what to tell people. Dataobservability takes that second shape. It connects read only to Snowflake, BigQuery, Databricks and Redshift, generates freshness, volume, schema and distribution monitors from warehouse metadata rather than hand written assertions, builds column level lineage from query history, and routes alerts to Slack and PagerDuty. It publishes 99, 299 and 799 dollars a month with a 14-day trial and no card, so the three year model above takes one line instead of a spreadsheet. If what you want is a curated business glossary with stewardship workflows, buy a catalog. If what you want is to know when data breaks and what it broke, the pricing works out very differently.

// FAQ

Questions buyers ask

Data lineage tools pricing FAQ

How expensive are data lineage solutions?

They span from zero to six figures a year. Open source lineage such as DataHub, OpenMetadata and Marquez has no license fee. Metered cloud catalogs are cheap at small scale, with the AWS Glue Data Catalog free to a million metadata objects a month and 1 dollar per 100,000 objects after that. Commercial catalog platforms are quote only and typically land in the tens of thousands a year once seats, implementation and stewardship are counted. Self serve platforms that publish a price start around 99 dollars a month.

How do pricing plans differ for automated data lineage tracking solutions?

They differ mainly in the billing unit. Catalog led platforms bill per named user per year, often splitting editors from viewers, so the bill tracks headcount. Cloud native services bill per governed asset per day or per metadata object per month plus the compute for scanning, so the bill tracks how many tables and partitions exist. Self serve monitoring platforms bill a flat monthly tier. The same organization can rank these three models completely differently depending on whether it has more people or more tables.

Why do data lineage vendors hide their pricing?

Because value based pricing works better when the seller sets the anchor after learning your team size, warehouse footprint and compliance deadline. A public figure would cap what the largest buyer pays and deter the smallest. In August 2026, checking each vendor pricing page directly, Atlan, Alation, Secoda, Select Star, Informatica and Datafold all route to a form, and the Collibra pricing URL returns a 404. The practical consequence is that two buyers can pay very different amounts for the same product.

What is the typical pricing model for data lineage and catalog tools?

The most common commercial model is an annual subscription priced per named user, tiered by capability, with column level lineage and advanced governance reserved for the higher tiers, plus a separate implementation fee in year one. The most common cloud native model is metered consumption against governed assets or metadata objects, with scanning compute billed separately by the hour. Open source has no license and shifts the cost to hosting and engineering time.

How do you estimate the 3 year cost of a data lineage tool?

Add four lines and project them forward. Year one license or metered consumption, the implementation and connector engagement, the stewardship hours per week times a loaded engineer rate, and the annual escalation clause in the contract for years two and three. Then add warehouse compute if the tool scans rather than reading query history. The result frequently ranks vendors differently than the year one quotes do, which is why year one quotes are a poor comparison instrument.

Is open source data lineage actually free?

Free to license, not free to run. DataHub, OpenMetadata and Marquez are Apache 2.0 with no license fee, and for a small stack with a handful of critical sources they are genuinely the cheapest credible option. The recurring costs are hosting the metadata store, writing and maintaining ingestion connectors for every source, and curating the result so people keep trusting it. Connector coverage is where open source catalogs are thinnest, and it is the line that grows fastest as the stack does.

Does Microsoft Purview publish a price for data lineage?

It publishes the structure but not the number. The pricing page names the meters and units clearly, billing per unique governed asset per day for Unified Catalog, per capacity unit per hour for the Data Map where a capacity unit is 25 operations per second and 10 GB, and per processing unit across Basic, Standard and Advanced tiers, with the first 1 MB of Data Map metadata storage free. Checked in August 2026 the dollar rates themselves render as a dash and resolve only in the regional calculator, because rates vary by region.

Which data lineage tools publish a real price?

Very few. Of thirteen lineage and catalog tools checked in August 2026, only the AWS Glue Data Catalog and Dataobservability publish a figure you can budget from without a sales call, and Microsoft Purview publishes the meters without the rates. Everything else, meaning Atlan, Alation, Collibra, Secoda, Select Star, Informatica and Datafold, is quote only. The open source projects have no license fee at all but shift the cost to hosting and engineering time.

Catch broken data before your stakeholders do

Connect your warehouse and get data lineage tools pricing live from one read-only connection. Transparent pricing, no credit card.