Keyword clustering decides whether your keyword list becomes a focused content plan or a sprawl of overlapping thin posts. The single decision that matters: which keywords belong on the same page, and which need separate pages. SERP overlap is the only reliable signal Google gives you for that decision, and most clustering tools either skip it or use it wrong. Get the method right and a 200-keyword list collapses to 40 to 60 high-intent pages instead of fragmenting into duplicate content competing with itself.
You’ll get the SERP overlap threshold that decides cluster boundaries, the three clustering methods ranked by accuracy on real SERPs, and the spreadsheet workflow that produces a publishable content map in under 90 minutes. Every threshold is calibrated against the Ahrefs 2025 keyword clustering study and tested against 14 client content audits I’ve run since January 2026.
What Keyword Clustering Actually Decides for Your Content Plan
Keyword clustering decides three things: which keywords share a page, which keywords need separate pages, and which keywords belong in the same topical cluster but on different supporting pages. The first decision is the one most teams get wrong. They group by topic (“all wordpress plugin keywords on one page”) instead of by SERP intent, which produces pages that target keywords with completely different ranking pages.
The right grouping signal is SERP overlap. If two keywords share 4 or more URLs in the top 10 Google results, they’re the same query intent and belong on one page. If they share 3 or fewer URLs, they’re different intents and need separate pages. The 4-URL threshold comes from the 2025 Ahrefs clustering study, which tested overlap thresholds from 2 to 8 against actual ranking outcomes and found 4 produced the best precision-recall balance.
This SERP overlap rule overrides every other clustering signal. Two keywords can be lexically similar (“seo audit” and “seo audit checklist”), share the same parent topic, and still belong on different pages because Google interprets them as different intents. Or two keywords can look unrelated (“local seo” and “google business profile optimization”) and belong on the same page because Google ranks the same URLs for both. Trust the SERP, not the keyword strings.
The Three Clustering Methods Ranked by Accuracy
Three clustering methods exist in production tools today: lexical clustering, topic clustering, and SERP overlap clustering. Lexical clustering groups keywords by string similarity, topic clustering uses semantic embeddings, and SERP overlap groups by ranking URL overlap. Their accuracy gap on real SERPs is large enough to make method choice the most important decision in the workflow.
Lexical clustering performs the worst, agreeing with manual ground-truth at 54% in the Ahrefs 2025 test. The method clusters “best seo tools” with “seo tool reviews” because they share words, even though Google ranks completely different pages for each query. Anyone using Ubersuggest or Mangools default clustering is getting lexical results, and the time savings come at the cost of fragmented content plans that produce internal cannibalization.
Topic clustering with semantic embeddings (the method most Python libraries use) hits 71% agreement, which is better but still misses high-stakes intent splits. SERP overlap clustering hits 89% agreement when measured at the 4-URL threshold, and it’s the only method that catches the cases where lexically dissimilar keywords share intent. The accuracy gap means SERP overlap is worth the API costs ($0.03 to $0.08 per keyword on most tools) for any project where the clustering decision drives content investment. For implementation, our walkthrough on competitor keyword research methodology shows where to source the underlying SERP data.
The 90-Minute Spreadsheet Workflow That Produces a Content Map
The workflow that produces a publishable content map runs in five steps and takes 90 minutes for a 200-keyword list. Step one: pull SERP data for all keywords from Ahrefs, Semrush, or SE Ranking. Step two: build a matrix of top-10 URLs per keyword. Step three: calculate URL overlap counts between every keyword pair. Step four: cluster keywords with overlap counts of 4 or more. Step five: review and adjust manually for the 5 to 10% of edge cases.
For step three, the Python script that calculates overlap takes about 12 lines of code: load the matrix, iterate keyword pairs, count shared URLs, write the overlap score to a new sheet. ChatGPT, Claude, or any open source model writes this script in under 60 seconds when you paste the matrix structure. Running it on 200 keywords produces a 200×200 overlap matrix in 3 to 4 seconds.
For step four, the clustering itself is a graph traversal: connect any two keywords with overlap of 4 or more, then identify connected components. Each component is a cluster. Running this on 200 keywords typically produces 40 to 60 clusters, with cluster sizes ranging from 1 keyword (unique intent) to 12 keywords (broad commercial query). The cluster size distribution matches the long-tail shape of search demand, which is what tells you the clustering worked correctly.
How to Choose the Page-Level Target Keyword for Each Cluster
Every cluster needs one primary keyword that becomes the page’s focus keyword in RankMath. Pick the keyword in each cluster with the highest search volume and the lowest difficulty score that still represents the cluster’s central intent. The volume gives you the search opportunity, and the difficulty score tells you whether the page can rank within 90 to 180 days.
The mistake to avoid: picking the highest-volume keyword regardless of intent fit. If the cluster contains “best wordpress plugins for seo” (4,400 volume) and “wordpress seo plugins” (3,600 volume) and “yoast vs rankmath” (1,900 volume), the right primary is “wordpress seo plugins” not “best wordpress plugins for seo,” because the search results show informational comparison content ranks better than ranked-list content for these terms. Reading the SERP tells you which keyword represents the cluster’s intent best.
Once the primary keyword is locked, the secondary keywords in the cluster become subheadings, internal anchor text, and natural mentions throughout the body. They don’t need separate optimization. They get covered as part of the article’s semantic territory, which matches how Google’s BERT and MUM models score content relevance. For more on the underlying semantic model, our piece on keyword difficulty scores and when they lie covers the difficulty-side calculation. Ahrefs published the original SERP overlap study at ahrefs.com/blog/keyword-clustering with the threshold testing methodology in full.

