How high cardinality happens
Cardinality is simply the count of distinct values a dimension can hold. Country has a small, stable set. Device category has three. Page path has one value for every page you publish, which is already larger but still manageable. The trouble starts when a value is close to unique for every visit: a full URL carrying a query string, a session identifier, a timestamp, a product variant code, a search term typed by a visitor.
Reporting tools are built to summarise, not to list. To keep report tables to a workable size, a platform such as GA4 keeps the values it saw most often for a given day and files everything else under a single bucket named (other). Nothing is deleted, and the totals still add up. The rare rows are merged, and once merged they cannot be separated again in the standard reports.
Why high cardinality matters
The damage is quiet, which is what makes it dangerous. Sessions, users and conversions all still reconcile, so nothing looks broken. What breaks is the breakdown you actually wanted. A landing page report that has collapsed tells you traffic arrived and almost nothing about where it landed. A campaign report built on unique identifiers instead of clean names turns into one giant unlabelled row halfway through the month.
It also arrives late. Cardinality builds up over a reporting period, so a dimension that looked fine on a single day can be useless across a quarter. Teams often discover the problem when they finally sit down to analyse a long window, which is exactly when the history they need has already been merged.
Common mistakes with high cardinality
The first is treating a custom dimension as a place to store an identifier. User IDs, order numbers, transaction references and email hashes are all near-unique by design, so a dimension built on them is guaranteed to collapse. They belong in a data warehouse or an export, not in a summary report.
The second is assuming the merged rows can be recovered by shortening the date range or changing the report. They cannot. The merge is applied when the report table is built, so the only real recovery is a raw event export where every row survives intact.
How to act on it
Design dimensions to answer a question, not to record everything. Strip query strings from page paths unless a parameter genuinely changes the content. Group values before they are sent: a plan tier rather than a plan identifier, a category rather than a stock code, a campaign name rather than a click identifier. If you truly need row-level detail, send it to BigQuery and analyse it there, and keep the analytics interface for grouped views.
When a report already shows a heavy merged row, treat it as a tracking design fault rather than a reporting bug, and fix it at the point of collection. A short review of what each dimension is for, run as part of your analytics and tracking setup, prevents most of it before any data is lost.