Methodology

How our data and scores work

If a product shows you a number, you should be able to find out where it came from. Here is where all of ours come from.

What we will not show you

Several YouTube tools display search volume, RPM, CPM or projected revenue for a keyword. YouTube does not publish any of those figures. They cannot be derived from any public interface. Every such number you have seen was modelled, extrapolated from a different platform's data, or simply invented.

We do not show them. Not in a smaller font, not with an asterisk, not as a “range”. A confident fabricated number is worse than no number, because you will make decisions with it.

The one place a number could mislead you is the opportunity score, so it carries an Estimate label everywhere it appears, and its inputs are always shown next to it.

The three labels

  • Measured — reported directly by the YouTube Data API v3. View counts, like counts, comment counts, subscriber counts, publish dates, durations and tags. We display them unmodified, with the time they were fetched.
  • Derived — arithmetic we perform on measured values. Views per day is views divided by days since publication. The search-sample outlier factor shown in research results is a video's views divided by the median views of that channel within the same result set — blank when we have fewer than three of that channel's videos to compare against, because two data points are not a baseline. It is not the same number as the watchlist's channel-history outlier factor, which uses a different denominator; both are described below.
  • Estimate — a relative heuristic with no real-world unit. Only the opportunity score falls in this category.

The opportunity score, in full

Version 2. A 0-100 ranking heuristic computed from measured YouTube inputs using the published weights, scales and fallbacks below. The inputs are measured; the way they are weighted and combined is a design choice we made, which is why all of it is written down here. It answers one question: compared with the other queries you are looking at right now, how attractive does this one look for a channel of your size? It is ordinal, not physical. Two queries scoring 71 and 52 tell you which to prefer; the gap of 19 has no unit.

ComponentWeightMeasured input
Reach of ranking results30%Median view count across the videos ranking for the query.
Openness to smaller channels25%The share of ranking videos published by channels smaller than yours that still beat the median view count.
Current velocity20%Median views per day among results published in the last 180 days.
Room for fresher coverage15%Median age of the ranking videos.
Headroom vs incumbents10%Inverse of the median subscriber count of the ranking channels.

Swipe the table sideways for the full measured-input column.

How each component is normalised

Reach of ranking results. Log-scaled between 1,000 views (0 points) and 10,000,000 views (100 points). View counts span orders of magnitude, so a linear scale would let a single viral outlier dominate. We use the median rather than the mean for the same reason. This is observed performance, not search volume: YouTube publishes no query-volume figure, and these views also reflect video age, channel size, recommendation traffic and links from elsewhere. Read it as how much reach the current results are getting, not how many people search the term.

Openness to smaller channels. A direct proportion, 0-100. When you have connected a channel we use your own subscriber count as the threshold; otherwise 100,000. This is the most decision-relevant component: it says whether the results are winnable rather than locked up by incumbents.

Current velocity. Log-scaled between 1 and 5,000 views per day. Captures live interest rather than an accumulated back-catalogue. Scored at a neutral 50 when there are no recent results — absence of fresh uploads is already rewarded by the staleness component and should not be counted twice.

Room for fresher coverage. Linear from 0 points at 0 days to 100 points at 900 days. Older top results mean the query is under-served by current content.

Headroom vs incumbents. 100 minus a log-scaled competition index between 100 subscribers (0) and 10,000,000 (100). Scored at a neutral 50 when channel sizes are unavailable.

The final score is the weighted sum of the five component scores, clamped to 0-100. Every intermediate value is returned with the result and rendered under the score, so you can always audit it against the raw numbers.

Confidence

Each score reports its own confidence. Fewer than eight ranking videos is reported as low. If subscriber counts were readable for fewer than half the channels, it is moderate, because two of the five components fall back to neutral. Otherwise it is good. The reason is stated in words next to the score.

Watchlist outliers

The channel watchlist asks a narrower question than the research tools: did this video do unusually well for the channel that published it? A global comparison would be meaningless, since 40,000 views is a disaster for one channel and a breakout for another. So the only comparison we make here is against the channel's own history.

This is the channel-history outlier factor, and it is a different calculation from the search-sample outlier factor in the research tools. That one divides by the median of a channel's videos inside one set of search results — often only three or four videos, chosen because they ranked for your query. This one divides by the median of the channel's recent uploads generally. The same video can carry two different factors, so the label always says which you are looking at.

  • The refresh runs daily on every plan. Each tracked channel is re-read about once every twenty hours, and the baseline and every stored factor are recomputed from scratch — a factor calculated against last month's median would drift as the channel keeps publishing. The weekly digest of newly flagged outliers is on every plan too. Plans differ in how many channels you can track, not in how often they are checked.
  • The baseline is the median views of that channel's recent uploads — sampled from its 50 most recent — counting only uploads at least 14 days old. Fresh uploads are excluded because they have not accumulated views yet; including them would drag the median down and make every older video look like an outlier.
  • The outlier factor is a video's views divided by that baseline. Both numbers are measured, so the ratio is Derived, not an Estimate.
  • No baseline, no number. A channel with fewer than 5 settled uploads gets no outlier factor at all, and the row says so instead of showing a figure computed from too little data.
  • The flag is set at 2× the baseline or above. The threshold is a product decision, not a discovered constant — it is the point at which a video is far enough from typical to be worth your attention.
  • Shorts and long-form get separate baselines. A video of 180 seconds or less is measured against the channel's other Shorts, and anything longer against its long-form uploads. The two formats have different view distributions, and one median across both sits between them — which made the stronger format read as a run of outliers and could stop the weaker one ever reaching the threshold. A format with fewer than 5 settled uploads of its own falls back to the pooled median, and each row says which baseline produced its number. The 180-second line is YouTube's own definition of a Short, not one we picked.

The outlier explorer, and what it is not

We do not crawl YouTube. Doing that at the scale needed to claim a database of “every outlier” is not something the API permits, and building a shallow version while calling it that would be exactly the sort of overclaim this page exists to rule out.

What we do keep is narrower and we would rather describe it precisely. Every keyword, top-video and niche search run on this platform fetches real videos from the API and pays quota for them. Rather than discard those records when the cache expires, we keep them, and the outlier explorer searches that store. It costs no credits and no quota because the data was already paid for once. Four things follow, and each is printed on the page itself:

  • Coverage is what has been researched, not what exists. A video nobody has ever searched their way to is not in there. An empty result means we have not seen it, which is not the same as it not existing — and we say so rather than let an empty page imply the stronger claim.
  • The factor is the search-sample one, described above: views over the median of that channel's videos inside the result set that surfaced it, over at least three of them. That is a thinner baseline than the watchlist's channel-history median, and the explorer labels it as such. It is a lead worth checking, not a settled measurement.
  • The statistics are as of the last time we fetched them. Views move; our copy does not, until research surfaces that video again. Every row states when it was last read instead of presenting an old number as a current one.
  • It is a rolling 30-day window, not a growing archive. YouTube's Developer Policies allow stored API data 30 days from the moment it was retrieved, after which it has to be re-read or deleted. So a daily job re-reads what it can — statistics first, and videos YouTube no longer serves are dropped outright — and everything it could not renew in time is deleted, channels included. A video nobody has researched in a month leaves. We would rather hold less and describe it accurately, and a two-year-old outlier was not a signal anyway.

For a channel you actually care about, the watchlist remains the better instrument: it reads that channel's uploads directly, on a daily schedule, against a baseline built from its own recent history.

How we check a replication brief is not a copy

A brief turns one of those outliers into a plan for your channel, which is a feature that could easily become a machine for producing knock-offs. So the instruction not to write one is not where we leave it — instructions are a hope. We measure the output against the video it came from, and print the result at the top of every brief:

  • Every proposed title, against the source title. We take the longest run of consecutive identical words shared by the two. More than 4 and that option is flagged in place. The limit is low on purpose: titles run eight to twelve words, so a five-word overlap is most of one.
  • The brief as a whole, against everything we hold of the source. The same check the script tools use for derived work — phrase overlap and longest identical passage, plus a similarity comparison of meaning — so a brief that avoids the title but tracks the framing is caught too.

A flagged brief is still shown to you. You are the one who decides whether an echo matters, and hiding it would be its own kind of dishonesty — but it is never shown quietly. What we will not do is present a near-copy as an original idea.

What we mean by “top 1%”

Tools in this category advertise finding the top 1% of videos in any niche, and almost never say the top 1% of what. Without a stated population the phrase cannot be checked, which is precisely what makes it comfortable to print.

Ours is narrower and checkable. A percentile here is a ranking of one video's search-sample outlier factor against the other videos we hold that share its niche, its format and its publication window. Every one of those constraints exists for a reason:

  • Factors, not view counts. Ranking by views ranks channels by size. The factor is already measured against each video's own channel, so ranking factors asks the question worth asking — how unusually well did this do for whoever made it.
  • One niche. Channels are classified into a fixed list of 27 categories, because a percentile against everything we happen to hold is a percentile against an arbitrary mixture. The list is closed rather than free-form: “gardening”, “garden” and “horticulture” would otherwise split one population into three. A channel we are less than 60% confident about is treated as unclassified and left out entirely, rather than guessed into a cohort it would distort.
  • One format. Shorts and long-form are ranked separately, for the same reason their outlier baselines are separate.
  • A stated window. The last 90 days first, widening to 365 only if the population is too small before that.
  • A minimum population of 500. A distribution of eighty videos does not have a meaningful 99th percentile. When even the widest window cannot assemble enough, you get no percentile and the reason, not a number computed from too little.

So the claim is always shown with the population it was taken from — “top 1% of 3,412 gardening long-form videos indexed in the last 90 days” — and never as “top 1% of gardening on YouTube”. We have not seen all of gardening on YouTube. Neither has anybody else, which is the part the shorter sentence leaves out.

Topic signal, brand effect, and emerging topics

Did the topic do the work, or the channel? When a video beats its channel's median we look for other videos we have already indexed with materially similar titles — sharing at least two of that title's distinctive words, and at least 40% of them, matched on stemmed titles. Among those we count how many separate channels also produced an outlier. Three or more distinct other channels is evidence the topic is doing the work; similar outliers on the originating channel alone, or none at all, points at that channel's audience instead. Between one and two we say so rather than pick a side. Every answer prints its counts and lists the similar titles it found, because the counts are the claim and the label is only shorthand.

Emerging topics counts how often each stemmed title word appears in videos first fetched in the last fortnight against the fortnight before, within one niche, and requires it on at least three distinct channels. Two limits there are definitions rather than caveats. The dates are when this platform's research first fetched a video, so it measures what is emerging in what we have seen and never what is emerging on YouTube. And the corpus grows every time somebody runs a search, so both windows' totals are shown beside every topic — a word rising more slowly than the corpus around it is not rising. Below fifty indexed videos in a niche we decline to report topics at all, because word frequencies need a corpus to be frequencies of.

Video teardowns

A teardown is deliberately split in two, and the split is structural — the measured facts and the interpretation are stored in separate columns and rendered in separate panels.

The measured half is API statistics plus arithmetic on them, including the same outlier factor described above. The interpreted half is a language model reading how the video is constructed: what the title withholds, how the opening is built, how the material is paced.

What the captions add. When a video has captions, we read them through a third-party provider — YouTube itself stopped serving public caption downloads in 2026 — and the timings come with them. That is what lets a teardown give you a timestamped outline of the video, quote its opening line, and place the moments where it opens a loop or asks you to subscribe. Sections built this way carry a From the transcript badge, and every timestamp is a link: you can check any of it in one click, which is the point. It is a third register, sitting between measured and interpreted — a caption track can be wrong, but it is not our opinion.

Long videos are read in samples: the opening, a passage from the middle and the ending. The model is told where the gaps are, in the text it reads, and so are you — the missing stretch is listed under what could not be measured. A video with no captions is analysed exactly as it was before, from the title, thumbnail and figures, and the sections that need a transcript simply do not appear rather than being filled in with guesses.

Pattern labels. Every teardown also files a video under four labels: hook type, story format, pacing, and where it asks the viewer for something. They come from fixed lists defined in our code rather than from whatever words the model reaches for that day, because the entire point is that two videos labelled “open loop” are saying the same thing — free-form labels fragment into synonyms within a week and stop being comparable. Each list carries a “something else” for a construction we can see clearly but have no name for, and a “not determinable” for one the evidence does not settle. Every label is shown with one sentence saying what settled it, so you can disagree with it rather than only accept it. Without captions, pacing and hook type are usually not determinable, and we would rather say so than pick the plausible answer. A label describes how a video is built. It is not a reason the video performed as it did.

It will not tell you why a video performed, because nobody can. The ranking system is not published, the audience is unobserved, and the video that was never made cannot be compared against. Any tool that answers that question confidently is guessing. No teardown claims a watch time, a click-through rate or a retention curve — none of those are in our input. Every teardown ends with an explicit list of what it could not assess.

Channel discovery

YouTube has no usable “channels in this niche” endpoint, so discovery works the way a person would: it searches the topic's videos and looks at who keeps appearing. A channel that ranks repeatedly for a topic is, by definition, a channel that ranks in that topic.

The limitation is real and is printed above every result set: this finds channels that rank for the phrasing you typed, not every channel in the niche. Ordering is by frequency of appearance, then median views within the sample — a ranking heuristic, not a judgement of quality. There is no revenue, RPM or “channel value” figure anywhere, because none of those can be measured from outside a channel.

Where the underlying data comes from

  • YouTube Data API v3 — search results, video statistics, channel statistics and your own channel's uploads. Every call's documented quota cost is recorded so we can tell you honestly when the daily allowance is exhausted rather than serving stale results.
  • YouTube search suggestions — the same autocomplete that powers the YouTube search box. It gives us the phrasings YouTube itself surfaces for a term. It is a list of queries, not a volume signal, and we never present it as one.
  • Captions — three routes, and each transcript records which one it was. For videos on a channel you have connected, we retrieve them through YouTube's official captions endpoints with your own authorisation. For any other public video, we ask Supadata, a third-party service, for the same transcript YouTube's own transcript panel shows — we ask it never to generate one by machine, so what you get is the caption track or nothing. We do not scrape youtube.com and we do not use unofficial endpoints to reach captions. A video with no caption track has nothing to retrieve, so there the text has to come from you: paste it, or upload the audio and we will transcribe it.
  • Exa web search — used in one place, and never for a number. Measuring a proposed sub-niche costs a YouTube search out of an allowance shared by everyone using PublishBench, so the niche explorer first checks whether the open web has anything to say about each proposal. One the language model invented, that nothing has ever been written about, is skipped and named as skipped rather than measured into a meaningless row. It chooses which queries are worth measuring; YouTube still supplies every figure you see. Where this is not configured, every proposal is measured.

Caching

Upstream responses are cached for three to twelve hours depending on the tool. A cached result is labelled as such and costs you no credits, because it cost us no upstream call. The fetch time is always displayed so you know how fresh the numbers are.

Originality checks

Scripts derived from an article or a video are compared against their source using three independent signals: 5-gram Jaccard similarity for wholesale phrasing reuse, the longest run of identical consecutive words for a single lifted passage, and — when an embedding provider is configured — cosine similarity for close paraphrase. The overall score is the maximum of the available lexical signals, so one strong overlap cannot be averaged away. A run of more than fifteen identical words fails regardless of the ratio.

Questions about any of this? Ask us — we will answer specifically.