Analytics without guessing: how to find where the budget burns

Ad optimisation usually looks like a dispute between opinions. One person says raise the bids, another says rewrite the ads, a third blames the season. Everyone sounds convincing, nobody is right for certain, and the budget keeps burning. Analytics exists to end that dispute: to see the exact address of the leak instead of guessing where it is. And the address is almost never where people looked by eye.

Ad optimisation usually looks like a dispute between opinions. One person says raise the bids. Another says the ads need rewriting. A third blames the season. Everyone sounds convincing, nobody is right for certain, and while the argument runs the budget keeps burning.

Analytics exists to end that dispute. Its job is to stop guessing where the money leaks and to see the exact address of the leak instead. That address is almost never where people looked by eye, for a structural reason: a human scanning an account sees phrases, campaigns and dates, while the waste hides at a different level entirely.

The good news is that you do not need to search among thousands of individual queries. The search happens at the level of words and distributions. That is where the clusters live that quietly eat a visible share of the budget without a single conversion. Below are three techniques that turn guessing into observation, with the parts that work differently in Yandex Direct than in Google Ads.

Look at words, not phrases

In a large account, reviewing tens of thousands of search queries by hand is not realistic. Yet if you break every query into its individual words and count clicks, spend and conversions per word, the picture clears instantly. This is word-level analysis, sometimes called n-gram analysis, and it works the same way in Yandex Direct and Google Ads.

The logic is simple. Take your keyword and search query statistics. Split each query into words. Aggregate the metrics by word rather than by phrase. What appears is something invisible at the phrase level: distinct clusters of vocabulary that consumed a noticeable share of the budget and produced nothing. In practice this is a recurring pattern: a handful of clusters responsible for the lion’s share of the money spent for nothing. At the phrase level they go unnoticed because they are smeared across hundreds of different wordings. At the word level they collect into one pile and become obvious.

Two things make this harder in Russian than in English, and a foreign team should know them before the first export.

First, morphology. A Russian noun changes its ending across six cases and two numbers, and an adjective agrees with it. The same concept can appear in a dozen surface forms across the query report. If you count raw tokens, one large off-target cluster looks like a dozen small ones, each too minor to act on. Normalise to the stem first, either with a stemming library or with the rough approach of trimming endings, and only then aggregate. The dozen small piles turn back into the one big one they always were.

Second, the source of data. Google Ads gives you a search terms report with its own filtering and, for some advertisers, third-party n-gram scripts that plug straight into the account. Yandex Direct has no built-in word-level view. You take the search query report from the Report Wizard, or pull it through the Reports API, and do the splitting yourself in a spreadsheet, a script or a BI tool. It is not difficult, but it will not do itself, and accounts run by people who expect the interface to surface the answer stay unanalysed for years.

Once the words are counted, they get labelled with meaning and grouped. Target vocabulary stays. Non-converting clusters go into negative keywords, and in Yandex Direct that means appending to the campaign or group list rather than replacing it, because a careless upload wipes what a previous cleanup already found. Entities labelled once are then filtered instantly in any cut, and repeating the analysis for a new period takes minutes instead of hours. That is what manageability at scale means: you do not re-read the queries every time. You work with a map you built once and keep updating.

For a company selling into Russia from abroad, this method has a quiet second benefit. Word-level clusters are the fastest way to see what Russian searchers actually mean when they type your category, which is often not what the translated keyword list assumed. A cluster that converts and that nobody planned for is a segment of demand you did not know you had.

Test beliefs with a distribution, not with faith

A large body of folklore lives around search advertising. That a keyword in the sitelinks moves the position. That the first position decides everything. That cosmetic changes to an ad transform its statistics. All of it is tested with the same technique: build a distribution and look at the facts.

Take the distribution of impressions by position. Export the statistics cut by position, group them in single-unit steps starting from zero, and count impressions in each bucket. Often it turns out that the share of premium impressions is small, and the hypothetical gain from cosmetic edits, measured on a large sample, is statistically indistinguishable from noise. Which means the time you were about to spend polishing the display URL for the sake of “position” would have gone nowhere. A distribution is a sobering instrument. It does not improve the advertising by itself. It saves you weeks of useless action.

This matters in Yandex Direct for a specific reason. The interface talks about premium placement more prominently than Google does, the auction shows traffic volume forecasts per bid, and a cottage industry of advice exists about “holding the first premium slot”. A team arriving from Google Ads absorbs that vocabulary and starts chasing position as a goal. The distribution tells you whether position is where your impressions actually are, and in many accounts it is not.

One technical subtlety, without which the numbers will lie to you. Relative metrics must be averaged with weights, and you must know what the weight is. Average click price is weighted by clicks: total cost over total clicks. Click-through rate is weighted by impressions: total clicks over total impressions. Conversion rate is weighted by clicks. The plain arithmetic mean of a column of rows gives a pretty number and a wrong one, because a keyword with ten impressions counts exactly as much as one with a hundred thousand. Then you make decisions on that number. So before drawing any conclusion from an average, confirm how it was computed and not simply that it exists in a cell.

Foreign teams fail here more often than local ones, and the maths is not the reason. A report arrives from a Russian contractor with averages already computed, nobody on the client side reads the column headers well enough to ask what the weight was, and a row-averaged click-through rate becomes a board-level fact. Ask. The answer takes one sentence, and if the contractor cannot give it, that tells you something too.

Having done this calculation once by hand in a pivot table, you carry the same logic into a business intelligence tool, where the metrics recompute themselves for any period and any project. Built once, then only the data is refreshed. Yandex Direct feeds such tools through the Reports API, so the pipeline is no harder to set up than the Google equivalent. It is just less documented in English.

Keep the raw material while it exists

There is a separate risk that people remember only when it is too late. Ad platforms periodically delete or anonymise historical search query statistics, citing privacy. After the cut-off date that data cannot be restored, exported or requested back. It simply disappears.

For analytics this is the loss of raw material for two tasks at once. Negative keyword collection, where you need to see real queries. And keyword expansion, where the living wording of real people is the source of new keywords. So at any official announcement from a platform about deleting historical data, the export is done in advance, before the announced date, and it is done as fully as possible.

A practical detail: the standard report in the interface is usually narrow, with few columns. It is more useful to assemble a custom report through the report builder and add as many attributes as the “search query” cut supports: dates, campaign, group and keyword statuses, every available dimension. The wider the export, the more analysis scenarios remain open after the raw data is gone. And this should be a habit rather than a one-off reaction to a particular announcement.

For a foreign advertiser this deserves a harder line than it gets at home, for two reasons. The Russian query history is often the only large body of real customer language you hold in Russian, and it cannot be recreated from a keyword planner. And access to the account frequently sits with a local agency or contractor, so when the contract ends the history ends with it unless you exported before. The rule is the same in both cases: export first, argue later.

The point is simple. Analytics lives on data, and data is not eternal. Whoever archives the raw material regularly keeps the ability to understand the past. Whoever wakes up after the deletion loses it for good.

What to do with your analytics this week

A short route:

  • Split the queries into words. Count spend and conversions per word, normalised to the stem, and find the clusters that spend the budget without a single inquiry. This is the cheapest money you can get back.

  • Test the folklore with a distribution. Before believing in “position” or ad cosmetics, build the distribution and look at the facts. In Yandex Direct especially, position is talked about more than it is measured.

  • Average correctly. Weight relative metrics, do not average them row by row, or the decisions will rest on a false number. Ask your contractor what the weight was.

  • Label entities once. Meaning labels on words and groups turn a repeat analysis from hours into minutes and survive a change of contractor.

  • Archive data ahead of time. When a platform announces the deletion of history, export the widest search query report available while the data still exists.

Analytics is not about handsome dashboards. It is about ending the argument and starting to see. Word-level analysis shows where the fire is. A distribution dismantles the myths you would have spent weeks on. A data archive preserves the very possibility of understanding. Together they are the move from “it seems to me” to “I can see”.

If the budget is being spent and you cannot say precisely what brings money and what drains it, the analytics is under-built. I can show you where it burns and where to start the cleanup in a free review of the account. How to build advertising where every ruble is visible and explainable is covered on the Yandex Ads page.

Frequently asked questions

Can I run word-level analysis in Yandex Direct the way I do in Google Ads?

Yes, and the logic is identical: export search query statistics, split every query into words, aggregate clicks, cost and conversions per word. The difference is where the data comes from. Yandex Direct has no built-in n-gram view, so you export the search query report from the Report Wizard or the Reports API, then do the split in a spreadsheet, a script or a BI tool. Russian morphology adds one step: a word appears in several inflected forms, so you normalise to the stem before counting, otherwise one cluster looks like six small ones.

What is the right way to average click price or click-through rate across campaigns?

Weight it. Average click price is total cost divided by total clicks, which means each row is weighted by its clicks. Click-through rate is total clicks divided by total impressions, weighted by impressions. Taking the plain arithmetic mean of the values in a column gives a number that looks tidy and is wrong, because a row with ten impressions counts as much as a row with a hundred thousand. Decisions built on that number inherit the error.

Does Yandex delete historical search query data the way Google Ads hid low-volume terms?

Ad platforms in general periodically trim or anonymise historical query data, citing privacy, and once the cut-off date passes the data cannot be recovered or requested back. Treat any official announcement of that kind as a deadline: export the widest possible search query report before the date, with every dimension the query cut supports. For a foreign company this matters twice over, because the Russian query history is usually the only body of live customer language you have in the language.

We sell into Russia from abroad and nobody on the team reads Russian. Can we still do this analysis?

The arithmetic does not need Russian. Word-level aggregation, weighted averages and position distributions are language-neutral: you find the cluster that spent a large share of the budget with zero conversions by the numbers alone. What needs Russian is the last step, deciding whether a cluster is off-target or a segment you had not planned for. That is one short session with someone who reads the language, not a full-time hire.

Sources

Andrey Belokrylov
Andrey Belokrylov

Independent marketing strategist and digital marketer. 10+ years, 100+ projects, from Marriott to small restaurants. I write about how Russian customers decide and how to run Yandex, VK and Avito without wasting the budget. More about me

From reading to doing

Let's look at your project

Send a link. In 30–40 minutes I'll show where the budget leaks, where Russian customers drop off, and what to fix first. Honest, even if we never work together.