Analytics and Tracking

High Cardinality

Also called Cardinality explosion

A dimension with so many distinct values that the reporting tool merges the rare ones into one catch-all row.

Quick facts: High Cardinality

Category
Analytics and Tracking
Also called
Cardinality explosion
Level
Intermediate
Affects
Report readability, custom dimensions, long-range analysis
Where to see it
GA4 (any report showing an (other) row), BigQuery export, Looker Studio
In this article4
  1. How high cardinality happens
  2. Why high cardinality matters
  3. Common mistakes with high cardinality
  4. How to act on it

How high cardinality happens

Cardinality is simply the count of distinct values a dimension can hold. Country has a small, stable set. Device category has three. Page path has one value for every page you publish, which is already larger but still manageable. The trouble starts when a value is close to unique for every visit: a full URL carrying a query string, a session identifier, a timestamp, a product variant code, a search term typed by a visitor.

Reporting tools are built to summarise, not to list. To keep report tables to a workable size, a platform such as GA4 keeps the values it saw most often for a given day and files everything else under a single bucket named (other). Nothing is deleted, and the totals still add up. The rare rows are merged, and once merged they cannot be separated again in the standard reports.

Why high cardinality matters

The damage is quiet, which is what makes it dangerous. Sessions, users and conversions all still reconcile, so nothing looks broken. What breaks is the breakdown you actually wanted. A landing page report that has collapsed tells you traffic arrived and almost nothing about where it landed. A campaign report built on unique identifiers instead of clean names turns into one giant unlabelled row halfway through the month.

It also arrives late. Cardinality builds up over a reporting period, so a dimension that looked fine on a single day can be useless across a quarter. Teams often discover the problem when they finally sit down to analyse a long window, which is exactly when the history they need has already been merged.

Common mistakes with high cardinality

The first is treating a custom dimension as a place to store an identifier. User IDs, order numbers, transaction references and email hashes are all near-unique by design, so a dimension built on them is guaranteed to collapse. They belong in a data warehouse or an export, not in a summary report.

The second is assuming the merged rows can be recovered by shortening the date range or changing the report. They cannot. The merge is applied when the report table is built, so the only real recovery is a raw event export where every row survives intact.

How to act on it

Design dimensions to answer a question, not to record everything. Strip query strings from page paths unless a parameter genuinely changes the content. Group values before they are sent: a plan tier rather than a plan identifier, a category rather than a stock code, a campaign name rather than a click identifier. If you truly need row-level detail, send it to BigQuery and analyse it there, and keep the analytics interface for grouped views.

When a report already shows a heavy merged row, treat it as a tracking design fault rather than a reporting bug, and fix it at the point of collection. A short review of what each dimension is for, run as part of your analytics and tracking setup, prevents most of it before any data is lost.

Do and do not

Do

  • Group values before sending them, not after
  • Strip query strings from page paths where they change nothing
  • Send row-level detail to a warehouse instead of a report

Do not

  • Store user or order identifiers in a custom dimension
  • Assume a shorter date range will restore merged rows
  • Add a dimension without deciding which question it answers

Questions people ask about this

Does high cardinality mean my analytics data is wrong?

No. Totals such as sessions, users and conversions stay accurate. What you lose is the ability to break those totals down by the affected dimension, because the rare values have been merged into a single catch-all row. The data was collected correctly; the summary report simply cannot show it at that level of detail.

Can I get the merged rows back later?

Not from the standard reports. The merge happens when the report table is built, and it is not reversed by changing the date range or the report layout. The only reliable way to keep row-level detail is a raw event export to a warehouse, set up before the data you care about is collected.

Which dimensions cause this most often?

Anything close to unique per visit. Full page URLs with query strings, internal search terms, session or user identifiers, order numbers, and custom dimensions built on database keys. Dimensions with a small fixed set of values, such as country, device category or traffic source, almost never cause the problem.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.