Topic clusters and cannibalization: deciding what counts as one topic

Two keywords belong on one page when they want the same answer. claude-seo decides that by diffing the SERPs of a pair; seodraft decides it by scoring how close two phrasings are, before it spends anything measuring them. Our run returned both answers for the same ten terms, and they disagreed.

Module 3 · 6 of 13 · 18 min

This course is independent. It has no affiliation with Anthropic or with AgriciDaniel, who writes claude-seo, and nobody on their side reviewed it. Run with claude-seo v2.4.1 on 2026-10-03.

What you will be able to do

  • Lay out a pillar and its spokes, and say what each page owns.
  • Read a SERP overlap score against the thresholds that produced the grouping.
  • Catch a term that collides with a page you already published, by hand.

Every planning tool eventually has to answer one question: are these two keywords the same topic? claude-seo answers it by running both searches and counting the URLs they return in common. seodraft answers a different version of it, scoring how close two phrasings are before it spends a cent measuring either. On 2026-10-03 we put the same seed through both, and the two answers pulled in opposite directions: eleven measured keyword pairs came back with almost no SERP overlap, while seven of ten terms were withheld for being too much alike. This lesson reads that disagreement, because it is the normal state of a young topic and the plan you build depends on which answer you believe. It ends with the check neither tool did for us, and the thirty seconds it takes by hand.

A pillar owns the topic and a spoke owns one question

A cluster is one page that defines the whole subject and a set of pages that each answer a single question under it, wired together so the reader and the crawler can get from any page to the centre. The pillar carries the broad term and links down; each spoke carries one specific term and links back up. The value is in the wiring: a topic covered by twelve disconnected posts looks like twelve opinions, and the same twelve under one pillar look like a body of work.

Our run proposed 1 pillar and 17 spokes across 5 clusters, with a link matrix of 72 internal links: 34 mandatory pillar-to-spoke links in both directions, 34 recommended links inside each cluster, and 4 bridges between clusters. Every spoke sits one click from the pillar, every post receives at least 3 links, and the plan reports no orphan page. The estimated total is about 28,800 words, which is the first number worth arguing with: that is a quarter of writing proposed off one afternoon of research.

How claude-seo decides two keywords belong together

The method is a diff of two result sets, scored on how many URLs they share. Its bands run like this: 7 to 10 shared URLs means both keywords want the same post, 4 to 6 means the same cluster, 2 to 3 means two posts that should link to each other, and 0 to 1 means two separate pages with nothing in common. The appeal of the rule is that it reads a decision Google has already made, rather than an opinion about language.

That rule needs a stable, position-ranked top ten to work on, and this run had something else. WebSearch returns roughly eight or nine synthesized source links per query and the set moves between sessions, which the report says in its own limitations. A threshold applied to a sample like that is weak evidence in both directions: a low score can mean two topics are genuinely separate, and it can mean the two searches happened to surface different pages from the same publishers.

Our eleven measured pairs all landed in the bottom band

Eleven pairs were actually compared, out of a matrix of eight keywords where every other cell is null. Every measured pair scored between 0 and 2, so not one of them reached the 4 the method asks for before it groups two terms.

PairShared URLsScore
seo agent · what is an seo agentlyzr.ai1
seo agent · ai seo agent vs seo toolwritesonic.com, lyzr.ai2
seo agent · best ai seo agents 2026none0
seo agent · how to build an ai seo agentahrefs.com, marketermilk.com2
how to build an ai seo agent · agentic seo workflow examplesnone0

Read the architecture again with that in front of you. The report built 5 clusters on a measurement that supports none of them, said so, and flagged the groupings as inferred from intent, shared domains and near-duplicate text rather than from URL overlap. The honesty is real and the plan is still an opinion. Treat the clusters as a draft architecture, verify by hand the three or four pairs you are about to spend a quarter on, and promote only those to measured.

seodraft measures a different kind of sameness

Phrasing similarity runs before any search and asks whether two strings pose the same question. add_topics scores every term against the bank and against the others in the same call, and withholds anything at or above 0.70 so the batch never pays to measure its own duplicates. At that threshold, seven of our ten terms were held back and none of the seven was stored:

  • "ai seo agent", "what is an seo agent", "ai agents for seo", "what is agentic seo", "best ai seo agents" and "agentic seo workflow" all collided with "seo agent".
  • "seo audit ai agent" collided with "how to build an ai seo agent".

Six collisions against one term is a signal about the plan rather than about the tool. The expansion produced many ways of saying the seed, and the SERP diff grouped almost none of them together, because the pages answering them today live on different URLs. Both measurements are correct about what they measured, and the pages you would write from them look nothing alike.

A plan of eighteen posts can collide with itself

The plan reports a cannibalization status of pass, with 0 duplicate primary keywords after 5 merges. That is true inside its own frame: each of the 18 posts carries a distinct primary keyword. The same 18 posts carry primary keywords that another tool read as the same question six times over, and the ones held back included the pillar phrasing itself.

What survives both readings is a smaller plan. Take the terms that collided and ask which of them a person would type and what the answer would open with; the ones whose first paragraph is identical become secondary keywords on one page, and the rest earn their own. Our seed is the clearest case: "seo agent", "ai seo agent" and "ai agents for seo" open the same way, so one page owns them and the other two phrasings live in its H2s.

The published-page check missed our own home page

seodraft can also compare a new term against the pages your site already has. import_sitemap reads your sitemap into an inventory, and add_topics withholds a phrase that resembles one of those pages, naming the URL, the path, the title inferred from the slug and the score. Nothing is downloaded, which is why every title is inferred and every consumer labels it that way.

Ours returned an empty published list for all ten terms. "seo agent" is the term our own home page targets, and the check never raised it, because the home page has no slug for a title to be inferred from. That is the gap worth carrying out of this lesson: an inventory built from slugs is blind to exactly the page you most want protected, since the money pages of most sites sit at the shortest paths.

Merge, separate, or let the older page keep it

Three decisions cover every collision, and the one that gets forgotten is the third. Merge when two phrasings would open with the same first paragraph: fold one into the other with merge_topics, which keeps the measured metrics and the paid brief on the survivor and archives the loser instead of deleting it. Separate when each phrasing needs a different opening to be useful, and send both through with allowSimilar: true so the bank records that you decided it on purpose.

The third decision is to write nothing. A term already owned by a page you published belongs to that page, and the work is improving it rather than planning a rival. The way to make that decision reachable is a list you maintain by hand: your home page, your pricing page and each main landing page, with the one term each one targets. Read the plan against that list before anybody writes, and seodraft's gate handles the rest at draft time (features/keyword-cannibalization).

What to settle before the pillar gets written

You now have an architecture, two disagreeing measurements and a list of collisions. Settle three things before a draft exists: which page owns each term, which pairs you verified yourself, and which of your existing pages already answers something on the list. Write those three on the plan, because a plan that records why a page exists survives the week somebody asks why there are two.

With the architecture settled, one page is ready to be specified in full. Lesson 7 builds the brief for it: the competitors, the gaps nobody covered, the outline with a word budget, and the rules that check the finished article against the plan it promised.

Steps

  1. Step 1

    claude-seo

    Get the architecture, not only the list

    This is the same run lesson 5 opened, read further down. Past the expansion the report lays out one pillar, the spokes under it, the clusters they sit in and a link matrix connecting them. Read the architecture before the word counts: the shape is the claim, and the word budget is a consequence of it.

    /seo cluster <your seed keyword>

    What you should see

    A pillar with its spokes and a link plan. Ours proposed 1 pillar and 17 spokes across 5 clusters, with 72 internal links, no orphan page, every spoke one click from the pillar and roughly 28,800 estimated words in total.

  2. Step 2

    claude-seo

    Separate the measured pairs from the inferred groups

    Paste this as a prompt in the same session. A cluster plan looks uniform on the page, and underneath it some pairs were compared and the rest were grouped on a tie-break rule. The matrix is where that difference is recorded, with a null in every cell nobody measured.

    Show me the SERP overlap matrix: which pairs were measured, what each pair scored, and which groupings are inferred.

    What you should see

    A scored matrix and a limitation. Ours measured 11 pairs, every one scoring 0 to 2 shared URLs, under the 4 its own method asks for before grouping, and the report marks the resulting groupings "inferred" rather than measured.

  3. Step 3

    seodraft

    Let the batch refuse its own duplicates

    `add_topics` compares every term against the bank and against the others in the same call before it measures anything, and withholds the ones that compete. The response names the score, the term and the topic it collided with. Re-sending with `allowSimilar: true` stores both on purpose, which is the escape hatch for two phrasings you decided are separate pages.

    Add these terms with add_topics and show me every term it withheld and what it collided with.

    What you should see

    A withheld list with scores. At the 0.70 threshold ours held back 7 of 10 terms: six collided with "seo agent" and one with "how to build an ai seo agent". None of them was measured, and none was stored.

  4. Step 4

    seodraft

    Rehearse a merge before you commit it

    `merge_topics` folds one question into another when two phrasings are the same query. Without `confirm` it is a dry run: the same envelope comes back with `applied: false` and a plan naming what would move. The survivor keeps the fresher paid brief and the measured metrics, an estimate never overwrites a measured number, and the losing side ends archived.

    Run merge_topics without confirm to show me what folding the second phrasing into the first would move.

    What you should see

    A plan and nothing written. A merge where both sides already have an article is refused by name, `article-conflict`, with both post ids, so the decision about which article survives stays yours.

  5. Step 5

    seodraft

    Check the terms against what you already published

    `import_sitemap` is free and reads your sitemap into the inventory of pages the site already has. No page is downloaded, so each title is inferred from its slug and every consumer labels it that way. `add_topics` then withholds a phrase that resembles one of those pages, naming the URL, the path, the inferred title and the score.

    Run import_sitemap, then re-send the terms and show me anything withheld for resembling a page I already published.

    What you should see

    A `published` list, which can be empty for reasons worth reading. Ours came back empty and never flagged "seo agent" against our own home page, which targets exactly that term: the home has no slug for a title to be inferred from.

Checklist

Tick every line and the lesson marks itself as completed.

Checkpoint

claude-seo scores a pair 1 and seodraft withholds the same pair at 0.78. Which one is right?Show the answer

Both, about different questions. The overlap score counts how many URLs the two searches returned in common, so it reports what Google currently treats as one answer. The similarity score reads how close the two phrasings are, before any search runs, so it reports what a reader would call the same question. A pair can be worded almost identically and still pull different results on a young topic, which is exactly our case. Take the low overlap as weak evidence when the sample is one synthesized result set, and let the phrasing score stop you from writing the same article twice.

The published-page check comes back empty. Is your plan clear of your own site?Show the answer

Only for the pages that have a slug to infer a title from. The inventory reads your sitemap and downloads nothing, so a URL whose path carries no words gives the check nothing to compare. Ours missed the home page, which targets the same term we were about to plan 18 posts around. The fix is thirty seconds of your own time: list your home page, your pricing page and your main landing pages with the term each one targets, and read that list against the plan before anybody writes.

What this lesson said

  • A cluster is one pillar that owns the topic and a set of spokes that each own one question, linked so no page is an orphan.
  • claude-seo groups by SERP overlap: 7 to 10 shared URLs means one post, 4 to 6 one cluster, 2 to 3 an interlink, 0 to 1 separate pages.
  • Our 11 measured pairs all scored 0 to 2, under the method's own grouping threshold, so the report marks its groupings inferred and says why.
  • seodraft's `add_topics` answers a different question before spending: at the 0.70 phrasing threshold it withheld 7 of our 10 terms, six against "seo agent" alone.
  • The published-page check reads slugs, so it misses a page without one. Ours never flagged our own home page, and that check stays manual.

Questions

What is keyword cannibalization, in one paragraph?
Two of your own pages competing for the same query, so the search engine picks one and splits the signals of both. It shows up as a page that ranked and then slid, as two URLs trading places week to week, or as a new article that never gains traction because an older one already answers its question. Planning is where it is cheap to fix: one term, one owner, decided before anybody writes. The /features/keyword-cannibalization page covers how seodraft blocks it at the gate.
Should I trust a cluster plan built on inferred groupings?
As a draft of the architecture, with the pairs you care about verified by hand. An inferred grouping is a reading of intent and shared domains, which is how an experienced strategist groups terms anyway; the difference is that a measurement can be checked and a reading has to be argued. Our plan proposed 18 posts off 11 measured pairs, so most of the shape rests on inference. Search the three or four pairs you are about to commit budget to, look at the ten results yourself, and promote those groupings to measured.
My plan has two posts for what feels like one question. Merge or separate?
Merge when the two phrasings would open with the same first paragraph, and separate when each one needs a different first paragraph to be useful. That test beats any threshold, because it asks what the reader gets. If you merge, pick the survivor by which phrasing a person would type, fold the other in with `merge_topics` so its metrics and brief survive, and redirect nothing, since neither page exists yet. If you separate, write the two intros first and read them side by side.

Next lesson

Build a content brief with Claude

What a brief has to contain, what claude-seo returns from today's SERP, and how seodraft checks the finished article against the plan it promised.

The same lesson, as plain Markdown: /learn/claude-seo/topic-clusters-and-cannibalization.md

Back to the course