Original research feasibility pilot

I Screened 471 Comments for Viewer Confusion. The Keyword Test Failed

A feasibility study showing why question words, rhetorical reactions, and information requests cannot be collapsed into a viewer-confusion metric without human context.

A comment containing 'why' is not automatically confused. It can be a joke, a criticism, a quote, a request for the song name, or a perfectly informed reaction to a character's decision. A keyword screen that ignores intent will label ordinary conversation as a failure of the video.

I tested that failure mode on 471 newest top-level comments from ten sampled Shorts. A multilingual lexicon flagged nine candidates. After reading each candidate in context, I confirmed one as premise confusion. The other eight were rhetorical questions, quoted dialogue, criticism, or normal information requests.

That poor precision is useful. It stops a more damaging report from publishing a 'confusion density' percentage or claiming late captions produce a multiplier. The pilot did not code all 462 non-candidates, so it cannot even tell us how much confusion the lexicon missed.

The 471-comment feasibility result

Feasibility counts only; 1/9 candidate precision has a very wide 95% Wilson interval of 2.0%–43.5%.
StageCountRateInterpretation
Comments screened471100%Newest top-level comments from ten videos; no replies.
Lexical candidates91.9%Matched one or more multilingual confusion markers.
Confirmed premise confusion111.1% of candidatesManual review found one comment expressing failure to understand the premise.
False-positive candidates888.9% of candidatesRhetorical, quoted, critical, conversational, or information-seeking uses.
Non-candidates fully human-coded0 of 462Not measuredRecall and true prevalence remain unknown.

Question form and question intent are different variables

A sentence can end with a question mark while functioning as a statement, challenge, joke, or emotional reaction. 'Why would he do that?' may show that the viewer understood the story perfectly and is judging the character. 'What song is this?' requests metadata, not the premise. 'Who is that?' might be referent confusion, fandom curiosity, or a request for an actor's name.

Natural-language research treats this as an intent-classification problem. Work on non-information-seeking question types separates rhetorical and other pragmatic uses, while a study of question intentions shows that surface form alone does not reveal what a speaker is doing.

Social comments add abbreviations, code-switching, emoji, quoted dialogue, missing punctuation, and shared context. A dictionary is useful for candidate retrieval. It is not a gold-standard labeler, and its candidate rate should never be renamed as viewer confusion prevalence.

The six comment intents the study must separate

IntentOperational definitionWhy the distinction matters
Premise confusionThe viewer cannot tell what happened or what the video is about.Closest to a failure of opening or narrative clarity.
Referent confusionThe viewer asks who or what a person, object, or term is.May reflect missing labels, assumed fandom knowledge, or ordinary curiosity.
Information requestThe viewer wants a song, product, location, recipe, source, or other metadata.Often signals interest rather than confusion.
Rhetorical reactionA question-shaped statement expresses disbelief, judgment, or emotion.The premise may be fully understood.
Quoted dialogue or joke participationThe comment repeats or extends language from the clip.Lexical overlap is participation, not diagnosis.
Off-topic or noisePromotion, emoji-only response, unrelated thread, or context-free text.Should not enter the substantive-text denominator.

Precision was measurable; recall was not

Precision asks: of the comments the lexicon flagged, how many were true premise-confusion cases? Here the answer was one of nine. Recall asks: of all true confusion comments, how many did the lexicon find? The pilot did not manually code the 462 non-candidates, so the denominator required for recall does not exist.

That missing audit is not a footnote. Viewers can express confusion without a question word: 'I have no idea what's happening,' 'context?', a language-specific idiom, or a reply that only becomes clear with the parent comment. A low candidate count can coexist with many missed cases.

The correct publication decision is therefore to stop. Reporting one confirmed comment out of 471 as a 0.2% confusion rate would assume perfect sensitivity, identical comment visibility, and a valid sampling window. None was established.

Newest comments are not early viewer reactions

The connector returned newest available top-level comments, not the first comments after publication. Older videos may have accumulated an audience that understands recurring characters and inside jokes. Moderation, deleted comments, ranking, and reply structure can also change what remains visible.

A study about early confusion needs a fixed window such as the first 24 or 48 hours, a collection timestamp, and a rule for delayed comment loading. It should store the position and thread structure needed for analysis without publishing usernames or raw text unnecessarily.

The denominator also needs two versions: all comments and substantive-text comments. Emoji-only responses cannot contain lexical evidence, but excluding them silently can exaggerate the apparent density. Both counts should be visible.

Caption timing cannot be linked with ten videos

The original idea was tempting: compare caption onset with confusion-comment density and estimate whether late text produces more confusion. Ten videos, one confirmed comment, and an unvalidated classifier cannot support that model. The unit of analysis would be the video, not the individual comment, leaving almost no independent observations.

Even with a larger sample, language, niche, comment volume, audience size, upload age, creator familiarity, text presence, and discovery source could affect both comments and editing style. A raw ratio would mistake those conditions for an effect of caption timing.

The publishable design is a mixed-effects model at the video level after a human-coded gold standard establishes language-specific error rates. The result would remain correlational. A caption-delay multiplier needs stronger evidence than a regression coefficient beside a weak classifier.

How to read your own comments without fooling yourself

  • Read the full sentence and nearby thread before labeling a question as confusion.
  • Separate premise confusion from requests for names, songs, products, recipes, or sources.
  • Distinguish rhetorical reaction from genuine information seeking.
  • Record the post-publication window; ten comments from day one are not comparable with ten comments from month six.
  • Look for confusion without question words, including 'context?' and declarative statements of uncertainty.
  • Treat comments as one evidence source alongside retention and direct viewer tasks, not a ground-truth diagnosis of the edit.
  • Preserve privacy by aggregating or paraphrasing unless a user has given permission to quote them.

The gold-standard comment study I would run next

Build a stratified set of at least 1,000 comments across languages, niches, audience sizes, and fixed post-publication windows. Two fluent coders should independently label all six intent categories, with a separate training set and preserved adjudication. Replies should be sampled deliberately rather than mixed accidentally with top-level comments.

Measure precision, recall, and F1 by language and category. Krippendorff's alpha is appropriate for multiple coders and can handle missing ratings; a recent open methods paper explains the coefficient, bootstrap uncertainty, coder training, and the common 0.80 reliability target.

Only after the classifier meets preregistered thresholds should it scale to a larger video sample. The public report must show the confusion matrix, hard examples, subgroup errors, excluded comments, and both comment denominators. An aggregate score without its errors would hide the exact problem this pilot found.

Claims this feasibility pilot cannot support

  • Only 0.2% of viewers were confused.
  • Question words accurately detect viewer confusion.
  • Late captions caused more confusion comments.
  • The classifier works equally across English, Hindi, Spanish, and Indonesian comments.
  • Newest comments represent the first audience response.
  • A comment classifier can replace direct comprehension testing.

Privacy and reproducibility boundary

The internal pilot stored aggregate counts and coding outcomes without retaining comment authors or republishing raw comments. Public accessibility is not permission to build a permanent identity-linked dataset. Any future corpus needs a retention plan, platform-policy review, and a decision about whether paraphrase is sufficient.

The codebook, collection date, candidate lexicon, language coverage, and reviewer decisions should be versioned. Model-generated labels remain proposals until a human audit establishes error rates. If the lexicon changes, old and new results should not be combined without recoding.

No relevant ViralJury customer win exists for this topic. The failed screen is itself the useful proof: the research process rejected a convenient metric before it became a product claim.

Frequently asked questions

How many of the 471 comments showed viewer confusion?

One of nine lexical candidates was confirmed as premise confusion, but the other 462 comments were not fully human-coded. The study therefore cannot estimate the true number or prevalence of confusion comments.

Why did the question-word classifier perform poorly?

Question words served many intents: rhetorical reaction, quoted dialogue, criticism, information requests, and normal conversation. The words identified candidates but could not determine intent without context.

Can comments reveal where a video is unclear?

They can supply useful clues, especially when several independent viewers describe the same missing context. They remain a selected, conversational sample and should be checked against direct comprehension tasks and creator analytics.

Did late captions cause more confusion comments?

No such result was established. The pilot had only ten videos, one confirmed premise-confusion candidate, and an unvalidated classifier. A larger video-level study with confounder controls is required.