If you built a YouTube script from something that already exists — a video you liked, an article, one of your own old uploads — two measurements check how original it really is, and between them they tell you almost everything worth knowing:
- Overlapping phrasing, as a proportion of the script
- The longest identical passage, in consecutive words
They catch different problems, which is why you need both.
Why two numbers instead of one
A script can be 5% overlapping and still have a problem, if that 5% is one forty-word passage lifted verbatim. A script can be 30% overlapping and be completely fine, if the overlap is unavoidable vocabulary — the terms, product names and stock phrases any video on the subject would use.
Proportional overlap tells you whether you rewrote the source or reorganised it. Longest identical run tells you whether any single stretch was copied. A high proportion with a short longest run is usually shared vocabulary. A low proportion with a long run is the one to worry about.
Neither is a verdict on its own, which is why any tool giving you a single pass-or-fail is discarding the information you actually needed.
| What it catches | What it is blind to | |
|---|---|---|
| Proportional overlap | Whether you rewrote the source or merely reorganised it | One lifted passage inside a script that is otherwise yours |
| Longest identical run | Whether any single stretch was reproduced word for word | A thorough paraphrase that keeps the source's order and argument |
What counts as too much
There is no universal threshold, and anyone quoting one is picking a number. What is defensible:
- A long identical run is the clearer signal. Somewhere past roughly a dozen consecutive words, coincidence stops being a plausible explanation for a sentence neither of you had to phrase that way.
- Proportional overlap has to be read against the subject. A script about a named product will share more vocabulary with every other script about that product than one about an abstract technique. The number means different things in those two cases.
- The direction matters more than the level. A repurposing workflow producing steadily rising overlap is drifting toward reproduction, and that trend is more informative than any single reading.
How to check a YouTube script by hand
You do not need a tool for one script.
- Paste both texts into a document side by side.
- Search your script for any six-word run from the source. If a six-word string matches, look at the sentence around it.
- Read your version without the source in front of you. If you cannot explain a sentence in your own words, you did not rewrite it — you rearranged it.
- Check the structure, not only the wording. Identical section order with different sentences is still derivative, and it is the form that a word-level check will not catch at all.
That fourth one matters more than people expect. Two scripts can share almost no phrasing and still be the same video.
The failure mode nobody checks for
Invented specifics. When a model rewrites a source, the thing most likely to go wrong is not copying — it is that a number, a name or a date drifts. The source says one figure, the rewrite says another, and the rewrite is confident about it.
This is worse than overlap, because overlap is visible and a drifted figure is not. Anything factual in a derived script should be checked against the source before it is spoken, not because the rewrite is dishonest but because it has no way of knowing it changed something.
Where this stops working
These numbers are not a legal or copyright judgement. They describe textual similarity. Whether a particular use is permissible is a question about jurisdiction, licence, purpose and amount, and no similarity measurement answers it.
They cannot see ideas. A script that takes another video's argument, structure and examples while sharing no phrasing at all will measure as original and will not be. Structural similarity is a judgement you have to make yourself.
They cannot check a source you did not give them. These measure your script against the specific text you built it from. They are not a search of everything ever published, and nothing that runs on your laptop is.
Running it automatically
In PublishBench, every script built from a source carries this check by default. It reports both figures — the proportion of overlapping phrasing and the longest identical passage — as numbers rather than hiding them behind a pass or fail, because whether that overlap is acceptable depends on the subject and the intent, and that is a judgement it should not be making for you.
It is included on every plan and never priced separately. There is also a repair pass that rewrites the passages driving the overlap, which is the thing to reach for when the number is high but the script is otherwise the one you wanted.
The script writer covers the rest of what it does, and the derived-script method covers how the rewrite is structured.