Every planning tool eventually has to answer one question: are these two keywords the same topic? claude-seo answers it by running both searches and counting the URLs they return in common. seodraft answers a different version of it, scoring how close two phrasings are before it spends a cent measuring either. On 2026-10-03 we put the same seed through both, and the two answers pulled in opposite directions: eleven measured keyword pairs came back with almost no SERP overlap, while seven of ten terms were withheld for being too much alike. This lesson reads that disagreement, because it is the normal state of a young topic and the plan you build depends on which answer you believe. It ends with the check neither tool did for us, and the thirty seconds it takes by hand.
A pillar owns the topic and a spoke owns one question
A cluster is one page that defines the whole subject and a set of pages that each answer a single question under it, wired together so the reader and the crawler can get from any page to the centre. The pillar carries the broad term and links down; each spoke carries one specific term and links back up. The value is in the wiring: a topic covered by twelve disconnected posts looks like twelve opinions, and the same twelve under one pillar look like a body of work.
Our run proposed 1 pillar and 17 spokes across 5 clusters, with a link matrix of 72 internal links: 34 mandatory pillar-to-spoke links in both directions, 34 recommended links inside each cluster, and 4 bridges between clusters. Every spoke sits one click from the pillar, every post receives at least 3 links, and the plan reports no orphan page. The estimated total is about 28,800 words, which is the first number worth arguing with: that is a quarter of writing proposed off one afternoon of research.
How claude-seo decides two keywords belong together
The method is a diff of two result sets, scored on how many URLs they share. Its bands run like this: 7 to 10 shared URLs means both keywords want the same post, 4 to 6 means the same cluster, 2 to 3 means two posts that should link to each other, and 0 to 1 means two separate pages with nothing in common. The appeal of the rule is that it reads a decision Google has already made, rather than an opinion about language.
That rule needs a stable, position-ranked top ten to work on, and this run had something else. WebSearch returns roughly eight or nine synthesized source links per query and the set moves between sessions, which the report says in its own limitations. A threshold applied to a sample like that is weak evidence in both directions: a low score can mean two topics are genuinely separate, and it can mean the two searches happened to surface different pages from the same publishers.
Our eleven measured pairs all landed in the bottom band
Eleven pairs were actually compared, out of a matrix of eight keywords where every other cell is null. Every measured pair scored between 0 and 2, so not one of them reached the 4 the method asks for before it groups two terms.
| Pair | Shared URLs | Score |
|---|---|---|
| seo agent · what is an seo agent | lyzr.ai | 1 |
| seo agent · ai seo agent vs seo tool | writesonic.com, lyzr.ai | 2 |
| seo agent · best ai seo agents 2026 | none | 0 |
| seo agent · how to build an ai seo agent | ahrefs.com, marketermilk.com | 2 |
| how to build an ai seo agent · agentic seo workflow examples | none | 0 |
Read the architecture again with that in front of you. The report built 5 clusters on a measurement that supports none of them, said so, and flagged the groupings as inferred from intent, shared domains and near-duplicate text rather than from URL overlap. The honesty is real and the plan is still an opinion. Treat the clusters as a draft architecture, verify by hand the three or four pairs you are about to spend a quarter on, and promote only those to measured.
seodraft measures a different kind of sameness
Phrasing similarity runs before any search and asks whether two strings pose the same question. add_topics scores every term against the bank and against the others in the same call, and withholds anything at or above 0.70 so the batch never pays to measure its own duplicates. At that threshold, seven of our ten terms were held back and none of the seven was stored:
- "ai seo agent", "what is an seo agent", "ai agents for seo", "what is agentic seo", "best ai seo agents" and "agentic seo workflow" all collided with "seo agent".
- "seo audit ai agent" collided with "how to build an ai seo agent".
Six collisions against one term is a signal about the plan rather than about the tool. The expansion produced many ways of saying the seed, and the SERP diff grouped almost none of them together, because the pages answering them today live on different URLs. Both measurements are correct about what they measured, and the pages you would write from them look nothing alike.
A plan of eighteen posts can collide with itself
The plan reports a cannibalization status of pass, with 0 duplicate primary keywords after 5 merges. That is true inside its own frame: each of the 18 posts carries a distinct primary keyword. The same 18 posts carry primary keywords that another tool read as the same question six times over, and the ones held back included the pillar phrasing itself.
What survives both readings is a smaller plan. Take the terms that collided and ask which of them a person would type and what the answer would open with; the ones whose first paragraph is identical become secondary keywords on one page, and the rest earn their own. Our seed is the clearest case: "seo agent", "ai seo agent" and "ai agents for seo" open the same way, so one page owns them and the other two phrasings live in its H2s.
The published-page check missed our own home page
seodraft can also compare a new term against the pages your site already has. import_sitemap reads your sitemap into an inventory, and add_topics withholds a phrase that resembles one of those pages, naming the URL, the path, the title inferred from the slug and the score. Nothing is downloaded, which is why every title is inferred and every consumer labels it that way.
Ours returned an empty published list for all ten terms. "seo agent" is the term our own home page targets, and the check never raised it, because the home page has no slug for a title to be inferred from. That is the gap worth carrying out of this lesson: an inventory built from slugs is blind to exactly the page you most want protected, since the money pages of most sites sit at the shortest paths.
Merge, separate, or let the older page keep it
Three decisions cover every collision, and the one that gets forgotten is the third. Merge when two phrasings would open with the same first paragraph: fold one into the other with merge_topics, which keeps the measured metrics and the paid brief on the survivor and archives the loser instead of deleting it. Separate when each phrasing needs a different opening to be useful, and send both through with allowSimilar: true so the bank records that you decided it on purpose.
The third decision is to write nothing. A term already owned by a page you published belongs to that page, and the work is improving it rather than planning a rival. The way to make that decision reachable is a list you maintain by hand: your home page, your pricing page and each main landing page, with the one term each one targets. Read the plan against that list before anybody writes, and seodraft's gate handles the rest at draft time (features/keyword-cannibalization).
What to settle before the pillar gets written
You now have an architecture, two disagreeing measurements and a list of collisions. Settle three things before a draft exists: which page owns each term, which pairs you verified yourself, and which of your existing pages already answers something on the list. Write those three on the plan, because a plan that records why a page exists survives the week somebody asks why there are two.
With the architecture settled, one page is ready to be specified in full. Lesson 7 builds the brief for it: the competitors, the gaps nobody covered, the outline with a word budget, and the rules that check the finished article against the plan it promised.