Original research protocol

I Hid the Outcome Before Testing First-Frame Clarity

A proposed teardown of 100 rights-cleared openings that measures reconstruction and reviewer agreement without treating public views as proof of clarity.

Pause a Short on its first encoded frame and you may see a clear subject, a black transition, a half-drawn caption, or a motion blur that was never meant to be inspected. That makes 'first-frame clarity' easy to discuss and surprisingly hard to measure.

The useful question is not whether the frame looks polished. It is what a first-time viewer can state after a controlled exposure. Can they name the subject? Can they describe the topic? Is there a promise or unresolved tension? Which visible word or object supports that answer?

This study keeps public view counts, creator identity, and the research team's preferred edit out of the reviewer task. Clarity gets measured as reconstruction, not as a feeling that one version looks more viral.

Define the frame before scoring it

Store two starting points. The literal first encoded frame preserves what the export actually contains. The first contentful frame is the earliest frame that passes a registered visibility rule. Both matter: silently skipping a black frame can improve a video's score after the fact, while treating a single decoder artifact as the intended opening can misrepresent the edit.

The study also stores the first 0.25, 0.5, and 1.0 seconds as exact frame ranges from the same source timebase. The first three stages remain muted. Audio is introduced only at the registered final stage. That sequence lets the report separate what the image establishes from what speech or sound supplies.

Reviewers cannot revisit a later stage while filling an earlier response. Otherwise the answer at one second leaks backward into the frame-zero score. The interface should lock each stage before the next begins.

Four staged exposures

A static image of the literal first encoded frame, muted. Reviewers record subject, context, topic, promise or tension, evidence, and uncertainty.

The first quarter-second loops without sound. Reviewers may update only fields that new motion or text onset changes.

The first half-second loops without sound. Short text and the direction of an action may now become available.

The first second plays with the registered audio condition. Reviewers record what audio adds and whether the opening depends on it.

Clarity is reconstruction, not attractiveness

A low-budget kitchen clip can be clear. A beautifully graded close-up can be unclear. Production value, attractiveness, emotional intensity, novelty, and personal preference belong outside the primary score.

The reviewer task asks for a one-sentence reconstruction, not a 22-item aesthetic rating. A preregistered answer key identifies necessary topic elements, acceptable alternatives, intended promise or tension, and misleading interpretations. Two coders score each response without seeing the public outcome.

Deliberate ambiguity can pass. If a clean image of a strange object establishes a specific question, the opening contains tension even if the answer is unknown. Accidental confusion is different: the reviewer cannot identify the object, the question, or what evidence would resolve it.

A smaller, testable clarity rubric

The raw memo proposed 22 dimensions. This protocol keeps the primary task narrow enough to train and audit.

DimensionQuestionResponse
SubjectCan the reviewer identify the central person, object, action, or scene?Specific, broad, absent, or cannot determine.
TopicCan the reviewer state what the video is about in one sentence?Reconstruction scored 0, 1, or 2 against a frozen key.
PromiseDoes the opening establish what will be shown, explained, tested, or changed?Specific, broad, absent, not applicable, or cannot determine.
TensionIf no promise is needed, is there a fair unresolved question or conflict?Specific, broad, absent, not applicable, or cannot determine.
EvidenceWhich visible object, action, word, or sound supports the answer?Short literal note linked to a frame or audio interval.
Channel dependenceWhat becomes clear only after audio, caption, motion, or prior knowledge?Independent flags with confidence.
ObstructionDoes interface-risk placement or another element cover information needed for the reconstruction?None, possible, essential, or cannot determine.

How the 100 openings would be selected

Use owned, synthetic, licensed, or explicitly consented media. The first study is a measurement teardown, so the sample should be built for variation rather than marketed as a random portrait of the internet. One workable frame is 20 openings each from commentary, tutorial, story, product demonstration, and visually led entertainment, with language and caption conditions declared.

Limit repeated creators, templates, and source concepts. Freeze the screening list before scoring and publish exclusion reasons. Do not select only examples that the ViralJury analyzer already considers clear; that would build the conclusion into the sample.

A top-performing-only version may still describe the openings of a defined successful set, but it cannot show that clarity caused success. Public views mix creator size, topic demand, distribution, timing, audience history, and many unobserved variables. A later matched study could compare openings from the same creator or source concept, while still avoiding causal language unless the design earns it.

Blinding and reviewer controls

ControlImplementationFailure it limits
Outcome blindingHide views, likes, comments, retention, upload date, and any top-performing label.Halo effects from knowing the result.
Identity blindingUse neutral IDs and remove handles when rights permit and the edit remains intact.Familiarity filling in missing context.
Order controlRandomize clip order per reviewer and counterbalance paired versions.Fatigue, learning, and first-version preference.
Independent scoringLock ratings before showing model flags, preferred edits, or peer scores.Consensus produced by social influence.
Fatigue limitsCap sessions, require breaks, and record completion time without using it as a quality shortcut.Drift during dense micro-timing review.
Held-out qualificationTrain on synthetic examples, then qualify on separate items.Testing memorization instead of rubric use.

Agreement and the prevalence problem

Report the raw response distribution and percent agreement before any chance-corrected coefficient. If nearly every opening has a visible person, some agreement statistics can behave oddly because the category prevalence is extreme. The report needs enough detail for a reader to see that pattern.

Use a preregistered coefficient suited to the response scale and multiple raters, with uncertainty intervals. Gwet's AC1 can be included as a sensitivity analysis for highly imbalanced nominal ratings; Krippendorff's alpha can support several measurement levels and missing data when its assumptions and distance functions are declared.

Agreement does not make the answer key correct. It only shows whether reviewers applied the frozen task consistently. Validity needs a separate study and a named target such as topic comprehension, not a broad claim about engagement.

What the final report must publish

  • The preregistration, sampling frame, screening flow, frozen rubric, answer-key procedure, and every deviation.
  • The rights basis for the sample and exactly which frames or clips may be reproduced.
  • Category distributions, reconstruction scores, missingness, uncertainty intervals, order effects, and reviewer agreement.
  • Results at every staged exposure, including literal frame zero and first contentful frame as separate values.
  • Null, ambiguous, and not-applicable outcomes instead of forcing every opening into a hook formula.
  • A limitation statement that separates clarity in this task from retention, recommendation, satisfaction, and creative quality.

Sources and protocol foundations

The preregistration boundary follows the Open Science Framework's definition of a timestamped, read-only study plan.

The response codebook follows the emphasis on explicit categories, independent coding, and revision history in Roberts and colleagues' codebook-development study.

Agreement reporting is informed by Hayes and Krippendorff's multi-rater framework and Gwet's discussion of nominal-data agreement.

Participant and media withdrawal procedures should be specified in advance, consistent with HHS guidance on withdrawal from research.

Frequently asked questions

Does a clear first frame need text?

No. A recognizable action, subject, setting, result, or tension can establish the opening visually. Text-assisted and purely visual clarity should be recorded separately so one does not masquerade as the other.

Why score both literal frame zero and first contentful frame?

Exports can begin on black, a transition, or a decoder artifact. The literal frame documents the file. The first-contentful rule documents intended visible content. Reporting both prevents silent improvement of the source.

Would 100 videos prove first-frame clarity causes performance?

No. One hundred rights-cleared openings can support a descriptive measurement pilot. A performance claim would require outcome data, a defensible comparison design, controls for creator and topic effects, and cautious interpretation.

Can an ambiguous opening still be clear?

Yes. Reviewers may clearly identify the subject and the unresolved question while not knowing the answer. The protocol separates fair tension from confusion with no identifiable question.