In short: a semantic core is not a keyword list — it is a map of "query to page". The working order: collect seeds from your product, competitors and internal analytics, expand them with keyword tools, pull frequency for your target region, strip the noise with negative keywords, verify intent against the actual SERP, group queries into clusters and assign exactly one landing page to every cluster.
What a semantic core is and why it matters
A semantic core is the set of search queries your site should rank for, distributed across pages. The important word is distributed. An export of 12,000 phrases with no "landing page" column is raw material, not a core. Raw material cannot be implemented: nobody knows what to write, where to write it, or whether it should exist at all.
A finished core answers three questions:
- Which pages you need. Every stable cluster is a candidate for its own page, subcategory or section.
- What to put on those pages. The wording of queries shows what people check before buying: price, lead time, compatibility, warranty, installation, returns.
- Which pages should not exist. If an idea has no cluster behind it, do not build the page. It adds noise to the index and brings no traffic.
Everything downstream grows from the core: catalogue structure, the article plan, draft headlines and snippets (see title, description and meta tags) and the internal linking scheme. That is why the core is built before sections are designed, not after — rebuilding structure on a live site always costs more than designing it once.
A core without a page map is useless. If your spreadsheet has no "which page" column after collection, the work is not finished — it has not really started.

Seed sources: where the first queries come from
Seeds are the base phrases everything else expands from. You do not need many: 50–200 good seeds produce tens of thousands of expansions. Bad seeds produce tens of thousands of garbage.
The product and its synonyms
Write down what you sell in every possible wording: category, subcategory, material, purpose, packaging format. The same object is named differently by different people — "server cabinet", "telecom rack", "19-inch enclosure". Each variant is a separate seed with its own demand.
Customer language, not company language
Nobody searches for your internal naming: tier names, SKU codes, "enterprise-grade solution". People search by the problem and by the everyday name. The cheapest source of customer vocabulary is call recordings, live chat and support threads — there people phrase things exactly as they type them into a search box.
Competitor sections
Open the menus and breadcrumbs of three to five competitors that hold the top consistently. Their structure is someone else's finished clustering. Copying it blindly is wrong — they may have a different range and their own mistakes — but as a seed source and as a "did we forget an entire section" check it works well.
Search suggestions and related searches
Suggestions are built on live demand, including recent demand that has not reached historical exports yet. Walk the alphabet: type a seed, then append letters one by one. Collect mobile suggestions separately — they differ from desktop.
On-site search
The most underused source. These people already reached you and typed, in their own words, what they could not find. Zero-result queries point directly at missing pages or missing stock.
Support tickets and sales questions
A repeating question is a ready-made informational cluster and a ready-made FAQ page. If a question comes in ten times a month, it is being typed into search engines too.
The query report in your search console
It shows which phrases already trigger impressions, including ones you never considered. Watch rows with many impressions and a low click-through rate: that is either an intent mismatch or a weak snippet. And watch positions 11–30 — that is the closest reserve to revenue.
Expansion and frequency: operators and what they really count
A keyword tool returns monthly impressions for a phrase in a chosen region. Without operators that number is almost always inflated: broad frequency sums up every query containing your words in any form, with any additions. Yandex Wordstat uses the operator syntax below; Google Keyword Planner and Search Console express the same idea through match types and filters, but the reasoning is identical.
# Broad frequency: all word forms plus any additional words
semantic core
# Phrase frequency: only these words, no tails, forms still free
"semantic core"
# Exact frequency: each word form is pinned
"!semantic !core"
# Negative keywords: cut irrelevant tails during collection
semantic core -free -course -pdf -download
# Grouping alternatives: several seeds in one pass
(build|collect|create) semantic core
# Pinning a preposition: without ! stop words are dropped
windows !in london
# Combination: exact form plus negatives
"!semantic !core" -competitor -generator
Why broad frequency lies
Take a real example from Russian-language data, where the pattern is easy to see. The broad frequency of "semantic core" is roughly 3,700 impressions per month nationwide. Inside that number sit "semantic core of a site", "how to build a semantic core", "semantic core online", "free semantic core", "competitor semantic core" and hundreds of other tails. Each is a distinct intent: one person wants a methodology, another wants a free tool, a third wants to reverse-engineer a competitor. Planning a single page against "3,700 impressions" means planning against traffic that does not exist.
| Notation | How it is written | What it counts | When to use it |
|---|---|---|---|
| W | semantic core | Every query containing both words in any form, plus any extra words | Sizing a topic, finding expansion directions |
| "W" | "semantic core" | The phrase with no additional words, forms still free | Spotting topics where all demand sits in the tail |
| "!W" | "!semantic !core" | The exact form of each word, no additions | Estimating the real ceiling of a specific landing page |
| "!W" with preposition | "!window !repair !in !london" | Same, but the preposition is not discarded | Geo queries and phrases where a preposition changes meaning |
Practical rule: expand on broad frequency, prioritise on exact. A gap of tens of times between W and "!W" is normal for general phrases and is a signal that the topic must be split into subtopics.
Frequency counts impressions, not visits. Even first place does not capture the whole volume: part goes to ads, part to answer blocks, part to refined follow-up queries. Treat frequency as an upper bound, never as a traffic forecast.

Region changes everything
Frequency without a region is a national average, and it almost never matches your situation. Differences come in three kinds:
- Volume. The same service in a capital city and in a regional one differs by multiples, not by percentages.
- Wording. Local terms, local abbreviations and district names live only inside their own region.
- The SERP itself. The same query returns a different mix of results in different cities: in one place marketplaces and aggregators own everything, in another local companies hold the top. This directly affects how realistic the cluster is for you.
So pull frequency in at least two slices: nationwide, to understand the topic and content priorities, and in your actual service region, to plan commercial pages. If the business operates in several cities, pull each one — otherwise it is easy to build a structure around demand you do not have.
Whether to create a page per city depends on one thing: do you have something genuinely different there — an address, a warehouse, staff, prices, delivery times. Without that, a geo page is a templated duplicate. On the related choice between city subdomains and folders, see subdomain vs subdirectory.
Cleaning the core: negatives and junk intent
After expansion the file holds 20–50 times more rows than you need. Cleaning happens in two passes: mechanical and semantic.
The mechanical pass
Lowercase everything, collapse whitespace, drop duplicates, subtract a stop list. Any script does this in seconds and removes the bulk.
# normalise and deduplicate the export
tr '[:upper:]' '[:lower:]' < raw.txt | tr -s ' ' | sed 's/^ //; s/ $//' | sort -u > clean.txt
# subtract stop words: stop.txt holds one word or fragment per line
grep -v -F -i -f stop.txt clean.txt > filtered.txt
# always eyeball what was actually removed
grep -F -i -f stop.txt clean.txt | head -50
# track volume at every step
wc -l raw.txt clean.txt filtered.txt
What belongs in the stop list
- Other companies' brands and models. If you do not sell them, that traffic is not yours. If you do sell them, it is a separate cluster with its own page, not an impurity inside a general category.
- "DIY", "make your own", "blueprint", "assembly drawing". The person plans to build it and will buy nothing — unless you deliberately pursue informational traffic.
- Gaming, TV-series and music noise. The classic homonym trap: "clan", "season", "episode", "mod", "chords". These tails often generate most of the broad frequency.
- Jobs and training. "Salary", "vacancy", "course", "certification", "career" — different intent, different audience.
- Academic noise. "Essay", "thesis", "presentation", "wikipedia", "definition pdf".
- Piracy tail. "Free download", "torrent", "crack", "keygen" — unless you actually ship a free tier.
- Foreign geography. Cities and countries you do not serve.
- Word-form and word-order duplicates. "buy pvc windows" and "pvc windows buy" are one goal. Keep the most frequent variant as the head phrase and the rest as supporting members of the cluster.
Never delete removed rows permanently. Keep a separate "junk" file and review it after every pass: a negative like "diagram" will take out "wiring diagram" too — and you may well have a page for that.
Intent verification: the step most teams skip
This is where most cores break. A query with solid frequency looks attractive, but frequency says nothing about what the person wants. A high-frequency phrase regularly turns out to be off-target: the searcher wants spare parts rather than the product, a repair manual rather than a purchase, a different product with a similar name, or an entirely different subject area.
How to verify intent
The only reliable source is the actual SERP. The search engine has already run the experiment on millions of people and is showing what they choose. The procedure:
- Open the results for the query in the right region, without personalisation — a clean private window.
- Write down the type of the first ten results: product page, category, article, forum, video, aggregator, marketplace, directory, manufacturer site.
- Look at the blocks: shopping carousel, images, video, answer boxes, maps. Their presence is a direct hint about intent.
- Read the titles. If eight of ten promise instructions, a commercial page will not take that slot.
The rule is simple: your page type must match the page types already ranking. Arguing with the SERP is expensive and pointless.
Three intent types
- Informational. "How", "what is", "why", "difference between", "guide", "DIY". Landing page: an article, a guide, a knowledge-base entry. These do not sell directly but build the top of the funnel and are the easiest content to get quoted in answer blocks.
- Commercial. Two subtypes. Transactional — "buy", "order", "price", "cost", "delivery" — landing page is a category or product page. Investigational — "best", "comparison", "reviews", "vs", "which to choose" — landing page is a comparison or review page with commercial blocks on it.
- Navigational. A brand name, "official site", "login", "customer portal", "contacts". Landing page: the specific brand page or section. Chasing other companies' navigational queries is futile — that person is looking for a named company and will go there.
Marker words are useful for a quick first sort, but the SERP always has the final say. "Price" appears in informational queries ("what determines the price"), and "how" appears in fully commercial ones ("how to order").
| Query type | Intent | Landing page type | Realistic expectations |
|---|---|---|---|
| High-frequency generic ("pvc windows") | Mixed, usually commercial-investigational | Root category or section hub | The slowest and most contested ground; approached last, after subcategories are working |
| Mid-frequency ("pvc windows for a house") | Commercial | Subcategory, solution page | The main working zone; movement shows once the page has history and internal links |
| Long-tail ("which profile for an unheated balcony") | Informational | Article, FAQ entry, knowledge base | Fastest response; little traffic per page, but many clusters that add up |
| Branded (your own brand) | Navigational | Home page, about page, login | Usually held without effort; the job is not to lose it to aggregators |
| Branded (someone else's) | Navigational | Brand page in the catalogue, if you actually stock it | Taking first place is effectively impossible; only worth it when the product is in stock |
| Geo ("pvc windows in Manchester") | Local commercial | Branch page with unique data: address, stock, lead times | Works where real presence exists; templated copies do not perform |
One more check: has the SERP been taken over entirely by marketplaces and aggregators? If the top ten contains no site of your type, downgrade the cluster regardless of its frequency.

Clustering: by meaning and by SERP overlap
Clustering groups queries that a single page can serve. There are two approaches, and in practice they work as a pair.
Manual, by meaning
You decide that "windows for a house" and "windows for a cottage" are the same thing. Fast, free, needs no tooling. The downside is subjectivity: the site owner's logic often differs from the search engine's. "Window repair" and "glazing unit replacement" feel identical, yet they return different sets of pages.
By SERP overlap
The formal method: two queries belong to one cluster if they share at least N URLs in the top ten. The threshold is usually 3–5 and is tuned per niche — higher in competitive commercial topics, lower in narrow ones.
- Hard clustering. Every query in the group overlaps with every other one. Clusters are small and clean, errors are rare, but you get many of them.
- Soft clustering. Overlap with the central query is enough. Clusters are larger, but the risk of gluing different intents together through a bridging query grows.
The working scheme: run hard clustering first to get a reliable skeleton, then merge small clusters manually where the meaning and the page type genuinely match. Automation without manual review produces strange groups; manual work without automation cements your own assumptions about how search thinks.
Different intents must never share a page. A page trying to be both a catalogue and a tutorial loses to both: it fully answers neither query, and its engagement signals are worse than those of specialised competitors.
Cluster to page: turning the core into site structure
The final step. Every cluster gets exactly one landing page and one head query the headline is built around.
Where new pages come from
- Catalogue subcategories — for mid-frequency commercial clusters with steady demand.
- Filter pages — only for combinations that have demand, stock and a unique description. Opening every filter combination to indexing is a duplicate generator.
- Solution and use-case pages ("windows for a summer house", "for a panel building") — for clusters describing a situation rather than a product.
- Articles and reference sections — for informational clusters.
- Comparison pages — for investigational clusters: "A or B", "difference between".
Where not to breed duplicates
- Separate pages for word forms and word-order variants.
- Separate pages for synonyms when the SERP for them is the same.
- Copies of one page per city with no real difference in content.
- Filter combinations with zero demand and an empty product list.
If cannibalisation already happened — two pages competing for one cluster — pick the primary one and either rewrite the other for an adjacent cluster or merge them. The mechanics and the common mistakes are covered in the guide to redirects. Pages produced from clusters (product, article, question-answer) should get their markup planned at the same time: see structured data.
# Core file layout: one row = one query
# CSV, ";" as separator, UTF-8
query;freq_broad;freq_exact;region;cluster;landing;intent;priority
semantic core;3725;;national;core-general;/articles/semantic-core-guide;info;1
semantic core of a site;608;;national;core-general;/articles/semantic-core-guide;info;1
how to build a semantic core;482;;national;core-howto;/articles/semantic-core-guide;info;1
semantic core collection;332;;national;core-howto;/articles/semantic-core-guide;info;2
semantic core clustering;61;;national;core-clustering;/articles/semantic-core-guide;info;2
semantic core online;136;;national;core-tool;/keywords;commercial;1
semantic core service;120;;national;core-tool;/keywords;commercial;1
site structure seo;56;;national;site-structure;NEW: structure section;info;3
crawl budget;50;;national;crawl-budget;NEW: crawl budget article;info;3
# The landing column is always filled. Allowed values:
# an existing URL - cluster already covered
# NEW: short description - page has to be created
# DROP: reason - cluster deliberately skipped
There are usually more "NEW" rows than expected — that list is the work plan. Rows with an empty landing column must not exist at all.

Prioritisation: why "biggest first" fails
The intuitive order — take queries by descending frequency — leads to a team spending six months against the most contested phrases and getting nothing. Priority is a product of three factors at once:
- Demand — exact frequency of the whole cluster in the target region, not broad frequency of the head query.
- Feasibility — what the SERP looks like: age and authority of the ranking sites, share of aggregators and marketplaces, presence of sites at your scale.
- Distance to revenue — how close the cluster sits to a purchase. A transactional cluster with 100 impressions is often worth more than an informational one with 3,000.
A practical launch order:
- Quick wins. Queries where the site already appears around positions 11–30 (search console data). The page exists, the intent matches, it needs content work and internal links.
- Mid-frequency commercial clusters where the top ten includes sites of your scale.
- Informational long tail — cheap to produce, brings topical entry points and internal links into commercial pages.
- High-frequency generic queries — last, once the section already has a mass of solid subpages.
No one can guarantee a position or the date it will be reached. Any promise of "top 3 in a month" is either ignorance or an intention to use methods that get sites penalised. Click and engagement manipulation belongs to exactly that risk category: the effect is unstable, and the penalty takes a long time to lift — sometimes it does not lift at all.
Common mistakes and keeping the core alive
- A core for the core's sake. Collected, reported, archived. Zero value: what works is the implemented structure, not the file.
- Thousands of queries with no page map. Export size is not a quality metric. 800 mapped queries beat 40,000 unmapped ones.
- Cannibalisation. Several pages targeting one cluster. Search picks which one to show and keeps changing its mind — rankings oscillate and internal weight is split.
- Geo doorways. Hundreds of city pages differing only by a substituted place name. They are recognised as templated and bring nothing.
- Ignoring intent. A commercial page for an informational query and the reverse. The most common cause of "the page exists but gets no traffic".
- Treating the core as a one-off. Demand is seasonal and shifting: new models, formats and terms appear. Review quarterly, and immediately after launching new products or services.
Crawl budget
A crawler spends a limited number of requests on your site. The more empty pages, duplicates and endless filter combinations you publish, the less attention goes to the pages the whole effort was for. A bloated structure hurts twice: important pages are crawled less often, and the site looks lower-quality overall.
A properly built core reduces crawl waste naturally — you simply never create pages that have no cluster behind them. Regular technical audits help find pages nobody needs; the baseline list of checks is in the SEO audit checklist.
How to check
- Pull frequency and break the list apart — the keyword research tool: frequency, related phrases and a first split into groups before manual intent review.
- Check the pages you assigned as landings — page SEO audit: headings, meta tags, duplicates, technical accessibility. A page with the right cluster and a broken technical layer will not rank.
- Find out why a new page is missing from search — the walkthrough on a site not appearing in search: indexing, robots rules, canonicals, content quality.
Frequently asked questions
How many queries should a semantic core contain
As many as it takes to cover every page you are genuinely ready to build and maintain. For a small service that may be 300–500 queries in 40–60 clusters; for a large catalogue, tens of thousands. File size is not a quality signal — the share of rows with a filled "landing page" column is.
Can a core be built for free
Yes. Keyword tools, search suggestions, the search console and on-site search are free and cover most of the work. Paid services save time on bulk frequency collection and automated SERP-overlap clustering, but they do not replace manual intent review.
How fast will results appear
There is no fixed timeline: it depends on competition in the niche, the age and technical state of the site, and page quality. Low-frequency informational clusters and queries already ranking in the second or third page of results usually respond first. Generic high-frequency phrases take the longest. Any specific deadline promised in advance is invented.
Does every long-tail query need its own page
No. A page belongs to a cluster, not to a query. If five phrasings return the same SERP, that is one cluster and one page; the other phrasings live inside the text, in subheadings and in the question block.
What to do with queries dominated by marketplaces
Assess them soberly. If the top ten contains no site of your type, getting there is unlikely regardless of page quality. Either postpone such clusters or approach them differently: presence on the platforms themselves, narrower subqueries where platforms are absent, or informational pages around the topic.
How often should the core be rebuilt
A full rebuild is rare — when the product range changes or you enter a new market. The routine work is different: every quarter, add new queries from the search console and on-site search, review clusters that lost demand, and check new pages for cannibalisation.
Core-building checklist
- Seeds collected from at least five sources, including on-site search and customer conversations.
- Expansion done with operators; negative keywords applied during collection, not after.
- Frequency pulled in both broad and exact form; priorities calculated from the exact one.
- Frequency pulled per service region, not only nationally.
- Mechanical cleanup done, removed rows saved to a separate file and reviewed.
- Intent verified against the live SERP for every cluster that entered the work plan.
- Clusters produced by SERP overlap and manually checked for mixed intent.
- Each cluster has exactly one landing: an existing URL, "NEW", or "DROP" with a reason.
- Verified that no existing page is assigned to two clusters at once.
- Priority set by demand, feasibility and distance to revenue — not by frequency.
- The "NEW" list converted into a publishing plan with owners and dates.
- A date set for the next core review.