Broad keyword collection for Russia: competitor crawls and planners
A keyword list written from memory gives you a hundred and fifty obvious phrases. A list collected from large competitor sites through a keyword planner gives you tens of thousands. The catch is in how you pick the competitor pages: one of the two ways quietly throws away half the demand.
By Andrey Belokrylov · September 18, 2026 · 9 min read

A keyword list written from memory is short. A marketing team sits down, lists what the product is called, adds a few synonyms, and ends up with a hundred and fifty phrases that every competitor also bids on. To get tens of thousands of relevant phrases, you do not invent them. You take them from large competitor sites and run those sites through a keyword planner.
The method is simple to describe and easy to break. There are two ways to pull the right pages out of someone else’s site, and one of them loses a large share of demand without leaving any trace. This article covers how to collect wide and keep the relevant phrases on the way, and what a foreign company should expect when the list is meant for the Russian market.
Why an inclusion filter fails on someone else’s site
The task looks like this. A competitor runs a site with tens or hundreds of thousands of pages. You need only the category pages, because those are the ones worth feeding to a planner: each category describes a slice of demand in the competitor’s own words. At that scale nobody copies addresses out of a menu by hand. A crawler does the work, Screaming Frog for example, and you give it filtering rules written as regular expressions.
Here the road forks. The first temptation is an inclusion filter: collect only addresses that match a pattern. It feels like the simplest way to narrow the sample. In practice it is the riskiest one, because you do not know in advance how the other site is built. Categories may sit in tidy subfolders. They may also sit directly in the root with no system at all. A rule such as “collect everything within the top two levels” rests on an assumption that folder depth matches page type, and nothing guarantees that. Guess the structure wrong and an entire section of demand never enters the collection. You will not find out, because a missing category produces no error.
The second approach reverses the logic. An exclusion filter drops addresses that are recognisably useless: product cards, service pages, pagination, catalogue filters. It still cuts the volume of collected data many times over. What it does not do is throw away a category by accident, because every exclusion is based on a clear, recognisable pattern rather than a guess about where the valuable pages live.
A wrong exclusion rule is visible. You see the page type you removed and can check it. A wrong inclusion rule is invisible, because the pages it missed never appear anywhere. That asymmetry is the whole argument.
Cleaning the list before it reaches the planner
The exclusion filter does not finish the job. Large catalogues are full of nested subcategories and implicit duplicates: one category can be reachable at several addresses, with sorting parameters, with pagination, through different paths. So the already reduced list goes into a spreadsheet tool, Power Query for instance, where it is cleaned down to the category addresses you need with no duplicates.
Only then does the list go to the planner. Google Keyword Planner accepts a page address and returns phrases related to that page. Run the cleaned category addresses through it and the output is tens of thousands of unique phrases, with wide coverage of synonyms and of the nested depth inside each category.
What changes when the campaign is for Yandex
For a company used to Google, this is where the Russian specifics begin, and they are worth stating plainly.
The planner and the measure are different tools. Google stopped selling advertising in Russia in 2022. Keyword Planner remains useful as a generator of phrases from pages, because it reads the competitor’s content and proposes wording. The demand figures that should drive a Russian campaign, though, come from Yandex Wordstat, the statistics service behind Yandex Direct. Wordstat works from phrases rather than page addresses, so the two tools fit together in sequence: the planner widens the list, Wordstat measures it. I wrote separately about what Wordstat shows and what it hides.
Russian morphology multiplies the list. A Russian noun changes its ending by case and number, and Yandex matches word forms by default. A phrase list collected from Russian competitor pages will contain the same query in many grammatical shapes. That inflates the raw count and makes deduplication by exact string match misleading. The cleaning stage has to work at the level of words and their forms, which is also why the frequency check below runs in stages.
Crawl local competitors, not your own translated site. Large Russian catalogues were structured for Russian buyers, so their category names reflect how people here search. A translated version of your home-market categories often does not. For a company selling into Russia from abroad, the sequence is: local competitors, exclusion crawl, cleaned addresses, planner, then Wordstat.
One algorithm for any language and either ad system
If the project is not Russian-language, does the pipeline have to be rebuilt? No. The order of work is the same whatever the language.
The only real difference for a non-Russian niche is the extra step of translating phrases into the specialist’s working language. Everything else happens in that working language and does not depend on the source language of the keywords:
- Collect the domains of competitors in the industry.
- Pull the initial phrase set for them from the planner.
- Merge the exports into one file and remove duplicates.
- Run a word-level frequency analysis.
- Tag words with meaning labels.
- Filter with a black list and a white list.
- Extend again from the phrases that are already relevant.
- Group the result and load the campaigns into the editor.
The skeleton transfers between ad systems in the same way. The basic logic of collection, cleaning and grouping is shared by Google Ads and Yandex Direct. Only the details of loading finished campaigns into a particular interface differ. Once you have built the pipeline, you use it in both places. For the cleaning step on a Russian list specifically, see how to clean a Russian keyword set.
Frequencies in stages, so you do not drown in captchas
With tens of thousands of phrases, taking frequencies for all of them head-on is slow and runs into captchas. A simple, less obvious saving helps: take frequencies in several passes rather than all at once.
The logic rests on a plain property. If a phrase has zero base frequency, its exact frequency and its fixed-order frequency will also be zero, so there is no point requesting them. In Wordstat terms, base frequency counts every query containing the words, and the narrower checks use operators that fix the word forms and then the word order. The order of work in Key Collector, a desktop tool widely used in Russia for bulk Wordstat collection, is:
- Take base frequency for the whole list.
- Keep only phrases with a non-zero value.
- For that remainder, take exact frequency with fixed word forms.
- Keep only non-zero values again.
- Only for the narrowest remainder, take frequency with both word order and word forms fixed.
The gain is twofold. You do not spend time on exact frequencies where the base is already zero, and the total number of requests to the statistics service falls, and with it the cost of passing captchas. The more obviously empty phrases the list contains, the larger the saving; on a dense, high-frequency set the effect is smaller.
Separately, align the parsing settings: the same number of threads, the same delays between requests, the same number of connected accounts. Then the speed is predictable and the collection does not stall for no visible reason.
New keywords from your own queries, every couple of weeks
A wide collection at launch is only the start. The most underrated source of growth sits inside an account that is already running, in its accumulated search query history.
The ad system records the real wording people used before reaching you, and part of that wording is not covered by the current keyword list. A regular review of these statistics, once every one or two weeks, where you pick out new relevant wording and launch it as separate keywords, gradually widens the demand you cover. Without it, that demand simply passes you by. In practice, when this was applied systematically over several months, lead volume did not grow smoothly. It grew in jumps: a new batch of keywords would suddenly open access to a noticeable amount of extra traffic.
This is also an answer to a common question from companies that hire a contractor for Russian advertising: how do you judge the quality of account management? Not by one magic metric. Look at how often changes were made to the campaigns and to the site, what those changes were, and whether the contractor had to redo their own work from scratch. Passive waiting for statistics to accumulate and copying someone else’s structure do not count. That is imitation of work. Launching new keywords from accumulated queries on a regular schedule is one of the most visible forms of real work. More on this in how to judge a Russian ads contractor.
What to do with your keyword set in the coming weeks
- When collecting from a large competitor site, filter the crawl by exclusion rather than inclusion, and finish the cleaning in a spreadsheet.
- Run the category addresses through a keyword planner to get wide coverage of synonyms and nested phrases.
- For Russia, measure the resulting list in Wordstat rather than trusting planner volumes.
- Do not reinvent the process for each language and each ad system. The skeleton stays the same, only translation and upload details change.
- Take frequencies in stages, removing zero phrases at every step, to save time and captchas.
- Set up a routine: every one or two weeks, pull new relevant wording from the query history and launch it as separate keywords.
A keyword set is not a one-off file prepared at launch. It is a wide net: you cast it far first, then pull it in regularly and take out what has swum into it. Teams that keep pulling grow in jumps. Teams that collected once and forgot wonder why demand seems to have dried up.
If you suspect half of your relevant demand is passing your campaigns by, request a free account review and I will show where the gaps in coverage are. How keyword collection fits into a paid traffic system is described on the Yandex Ads page.
Frequently asked questions
Why use an exclusion filter instead of an inclusion filter when crawling a competitor site?
Because an inclusion filter depends on a guess about how someone else's site is organised, and you cannot check that guess at the scale of a large catalogue. If the categories do not sit where you assumed, a whole section of demand never enters the collection and nothing tells you it is missing. An exclusion filter removes pages you can recognise with certainty, such as product cards, service pages, pagination and catalogue filters, and keeps everything else, so a wrong assumption costs you extra rows rather than lost categories.
Can I use Google Keyword Planner for a campaign that will run on Yandex?
As a source of phrases, yes. The planner accepts a page address and returns phrases related to it, and the underlying method does not depend on the ad system. Volumes are a different matter: Google stopped selling ads in Russia in 2022, and the numbers that should drive a Yandex campaign come from Yandex Wordstat. Use the planner to widen the list and Wordstat to measure it.
Why collect Wordstat frequencies in several passes?
Because a phrase with zero base frequency will also show zero for every narrower operator, so checking it again is wasted time. On a list of tens of thousands of phrases each extra request adds to the time spent and to the captchas you have to pass. Take base frequency first, drop the zeros, and apply the narrower operators only to what remains.
Does the process need to be rebuilt for a language other than Russian?
No. The only extra step for a non-Russian niche is translating the phrases into the specialist's working language. Merging, deduplication, word-level frequency analysis, tagging, filtering by black and white lists and grouping all happen the same way regardless of the source language, and the same skeleton works for Google Ads and Yandex Direct.
Sources
- Google Ads help, using Keyword Planner: support.google.com/google-ads/answer/7337243
- Yandex Direct help, keyword symbols and operators: yandex.com/support/direct/en/keywords/symbols-and-operators
- Russian version of this article on belokrylovo.ru: Широкий сбор семантики: парсинг конкурентов и планировщик