Original research pilot
I Read the First Five Words of 25 Shorts. Most Did Not Name the Subject
An exploratory first-five-words study of 25 recovered Shorts transcripts, with a transparent codebook, multilingual ASR limits, and no claim that wording caused performance.
A creator can spend three seconds talking and still withhold the subject. The opening sounds energetic, the captions move, and the sentence is technically underway. A cold viewer is left holding five words like 'so you guys won't believe' with no object, situation, or useful promise yet.
I took the first five recovered tokens from 25 auto-caption tracks and coded the job they performed. Did they explicitly name a subject or object? Did they establish an action or role? Or were they generic, incomplete, or too dependent on what came next?
Fourteen of 25 landed in the generic-or-incomplete bucket. That is an interesting pipeline result, not a formula. The videos came from one view-count-ordered convenience frame, languages were mixed, and automatic speech recognition made errors. The study did not compare successful and unsuccessful Shorts, so the words cannot be credited for performance.
The first-five-words pilot results
| Opening code | Result | 95% Wilson interval | Interpretation |
|---|---|---|---|
| Explicit subject or object | 5/25 (20.0%) | 8.9%–39.1% | The recovered tokens directly named who or what the opening concerned. |
| Situational action | 6/25 (24.0%) | 11.5%–43.4% | The words established an action, role, or situation without necessarily naming the central object. |
| Generic or incomplete | 14/25 (56.0%) | 37.1%–73.3% | The five-token window depended on later language or visual context under the pilot rule. |
What counts as naming the subject in five words
An explicit-subject opening had to name the person, object, topic, or question the viewer would use to identify the premise. 'This phone battery failed' qualifies because the object and event are present. 'You won't believe what happened' does not, even though it signals surprise.
A situational-action opening could establish enough activity to orient the viewer without naming the central object. 'I walked into the meeting' gives a role and action, but the actual conflict may still be missing. The category is descriptive, not automatically weaker than an explicit noun.
Generic or incomplete language included greetings, discourse markers, pronouns without a visible referent, suspense shells, and fragments whose meaning arrived after token five. A visual-first Short with no speech can still be strong; the category says nothing about the images unless they were studied separately.
The coding difference in plain language
These are recreated examples, not quotations from sampled creators.
'Okay, so you need to…' The sentence has started, but the viewer still cannot name the topic from the words alone.
'Your Reel caption is late.' The object and problem arrive inside the same five-word window.
'You won't believe what happened…' Curiosity is requested before evidence appears.
'I deleted the wrong Short.' The action creates a concrete scene even before the consequence is explained.
Five words are not a stable multilingual measuring stick
Languages package meaning differently. A pronoun can carry information that English expresses with a noun. Compounds, clitics, code-switching, honorifics, and tokenization rules change what counts as one word. A five-token window is convenient for a pilot, but convenience is not measurement equivalence.
The sampled captions included Hindi fragments and other multilingual signals. Automatic captions omitted or distorted some words. YouTube's own automatic-caption documentation warns that pronunciation, accents, dialects, background noise, overlapping speakers, and multiple languages can produce incorrect or missing text. Creators are told to review and correct it.
The full study should therefore begin with one validated language cohort and use human-verified first spoken clauses. Separate language replications can follow with fluent coders and documented tokenization. Combining them first and apologizing in the limitations later would produce a cleaner chart and a worse result.
The caption timestamp was not a speech-onset measurement
Nineteen of the 25 recovered tracks began within 250 milliseconds, and the median first cue timestamp was 160 milliseconds. That looks like a speaking-speed result until the measurement is examined. The timestamp belongs to an automatically aligned caption cue, not a laboratory annotation of the first audible phoneme.
Automatic captions can shift the start of a cue, group several words, or delay availability while processing. A long initial silence can even prevent automatic captions from generating normally. The correct label is 'first recovered caption-cue time,' not 'time the creator began speaking.'
If speech onset matters, trained reviewers should inspect the waveform under a frozen acoustic rule. If semantic onset matters, they should mark the point where the clause becomes specific. Those two timestamps can differ, and neither should be inferred from ASR metadata alone.
Raw Shorts views would be the wrong outcome
A view-count-ordered sample mixes channel scale, upload age, topic demand, geography, repeat plays, and discovery surface. Since 31 March 2025, YouTube counts a Shorts view when the video starts or replays, with no minimum watch time. The earlier measure remains as engaged views in Analytics, according to YouTube's current Shorts view explanation.
Comparing five-word categories against raw views would therefore reward older uploads, larger audiences, repeat starts, and popular topics alongside any wording effect. Even a matched public cohort would remain observational. The stronger outcome is creator-authorized early audience response inside niche, duration, upload-age, and channel-scale bands.
The useful claim is not 'nouns get views.' It is, at most, whether a verified opening-language category is associated with an early audience measure after the study accounts for obvious clustering and confounders. The causal version would need an experiment that changes wording while preserving the rest of the video.
The full study's opening-language codebook
The pilot's three buckets should expand before performance analysis.
| Category | Operational question | Common edge case |
|---|---|---|
| Explicit subject/object | Does the verified first clause name the central person, object, or topic? | A named object may still fail to state why it matters. |
| Explicit question | Does the opening ask a question with a specific subject and answer space? | A rhetorical question may perform a different job from information seeking. |
| Role/action setup | Does it establish who is acting and what is happening? | Pronouns require a visible or prior referent. |
| Greeting/filler | Could the words be removed without losing the premise? | A greeting can be culturally or narratively meaningful. |
| Other incomplete setup | Has speech begun without enough information to identify the promise? | The visual track may supply what the words omit. |
| No speech | Is the opening deliberately visual or silent? | No speech is not automatically a failure category. |
A first-five-words check you can use before posting
- Write the first spoken clause exactly as a cold viewer hears it; do not summarize what you meant.
- Circle the first concrete subject, object, action, conflict, or question.
- Remove greetings and discourse markers temporarily to see whether the premise becomes faster without changing your voice.
- Check the visual track separately; a deliberate visual-first opening may not need speech to carry the subject.
- Review automatic captions against the audio, especially for names, dialect, code-switching, and overlapping speakers.
- Test two versions with target viewers if the opening choice matters. Do not infer success from word category alone.
The validated first-five-words study I would run next
Start with 300 English-language clips and human-verify the complete first spoken clause. Use two trained coders and preregister the six mutually exclusive categories above. Audit transcript verification against the audio and stop publication if accuracy falls below the registered 95% threshold.
To test association, use creator-authorized analytics or tightly matched public cohorts within niche, upload-age band, channel-size band, duration, and discovery source. Treat channel as a random effect when repeated channels are allowed. Publish the distribution of wording categories before any outcome model.
A filler leaderboard can be a secondary descriptive feature after recognition accuracy is established. It should never become a shame list or a claim that one phrase kills a Short. The primary research question is whether specific opening information predicts a defined response inside a valid comparison—not which creator used 'so' most often.
Claims this pilot cannot support
- High-performing Shorts usually name the subject in the first five words.
- Generic openings cause viewers to swipe.
- Creators should always speak within 160 milliseconds.
- Five English-style tokens measure equivalent information across languages.
- No-speech openings are weaker than spoken openings.
- The first five words predict exact views, retention, or recommendation.
Research and accessibility sources
YouTube's caption creation guide explains that caption files contain text and timing, with some formats also carrying position and style. That makes captions useful research material, but it does not make automatic tracks a verified transcript.
Research with Deaf and hard-of-hearing viewers documents the difference between caption availability and caption quality. The YouTube creator captioning study found that manually added caption signals can affect confidence and trust. This pilot did not audit accessibility; the source strengthens the case for human verification and context-rich labels.
No sampled transcript excerpt is reproduced on this page. Aggregate categories preserve the research point without turning creators' speech into a searchable quotation archive. The internal source frame remains subject to the documented platform-data refresh or deletion deadline.
Frequently asked questions
What were the most common first five words in the sampled Shorts?
The public pilot reports categories rather than a phrase leaderboard. Fourteen of 25 openings were generic or incomplete, six established a situational action, and five explicitly named a subject or object.
Should every YouTube Short name the subject immediately?
No. Visual-first, comedy, and narrative openings can establish a premise without an explicit noun. The practical test is whether the target viewer can identify the situation and reason to continue, not whether one grammatical pattern appears.
Do the first five words affect YouTube Shorts views?
This pilot did not test that. It had no lower-performing comparison group and no creator-authorized retention data. Any performance association requires matched cohorts or a controlled wording experiment.
Can automatic captions be used as exact transcripts?
Not without review. YouTube warns that accents, dialects, noise, multiple speakers, and multiple languages can cause errors. A research transcript should be verified against the audio by a language-qualified reviewer.