Original research protocol

I Built the Weak Second Index Around a Codebook, Not a Headline

A proposed annual benchmark for the first reviewable structural problem in a short-form video, with the definition, validation gates, and limits published before data collection.

A retention graph can show where people stopped watching after a video was published. It cannot tell an editor, by itself, what changed in the cut at that moment. The Weak Second Index asks a smaller question: when does the first observable interval appear that a trained reviewer would flag for another look?

That distinction matters. A reviewable interval is not a swipe. It may be a repeated promise, a missing referent, a caption attached to the wrong phrase, a continuity break, or a stretch that delays the promised result without adding progress. A viewer may enjoy it. Another may leave for an unrelated reason. The index describes the edit; it does not read the viewer's mind.

The study can become citable only if the rules exist before the result. This page freezes the proposed unit, labels, annotation record, validation gates, and claims that remain out of bounds.

The unit is an interval, not one clock second

The phrase weak second is useful editorial shorthand, but the unit cannot be forced into a one-second box. A caption mismatch may last six frames. A repeated setup may span three seconds. Reviewers should mark the smallest interval that contains enough evidence to apply a code, then store both boundaries.

The primary outcome is time from the first available frame to the start of the first qualifying interval. A video with no qualifying interval is right-censored at its end. That makes a time-to-event summary possible, but it does not turn the annotation into audience behavior. The event is 'first coded interval,' not 'viewer abandonment.'

Every record must retain the source export version and taxonomy version. If the codebook changes halfway through a year, the two versions cannot be quietly pooled. A moving definition can manufacture a trend even when the videos did not change.

The six proposed codes

Version 1 narrows the original draft to six observable categories. A reviewer may add up to two secondary codes, but must choose one primary explanation.

These labels are proposed editorial categories. They have not yet demonstrated reliability or validity.
CodeUse it whenDo not use it when
RepetitionA later beat restates an established point without a new distinction, escalation, or function.The repetition supplies emphasis, instruction, rhythm, a callback, or visible progress.
Missing contextA cold viewer lacks a referent, goal, relationship, or piece of context needed for the local beat.The ambiguity poses a fair question and the cut supplies timely resolution.
Visual inactivityThe image supplies no useful evidence or progression while the current idea needs visual support.A stable frame protects reading time, proof inspection, emotion, or deliberate restraint.
Caption lagCaption timing changes the meaning, speaker, joke, warning, or reference it is meant to support.A small offset stays within the registered tolerance and preserves meaning.
Continuity breakAdjacent beats lose the intended spatial, temporal, causal, or identity relationship.The discontinuity is a legible montage, ellipsis, surprise, or formal choice.
Payoff delayAn open promise remains unresolved while the interval adds no progress, evidence, complication, or preparation.The process itself is the value, or the interval earns suspense, prediction, or emotional response.

What one annotation must contain

A label without evidence is still vague feedback. Each annotation needs a start and end timestamp, primary and secondary codes, a literal observation, the possible friction, one revision hypothesis, one preserve note, reviewer confidence, reviewer ID, and codebook version.

The literal observation comes first. 'The speaker repeats the same promise over an unchanged close-up from 0:04.2 to 0:06.1' is observable. 'Everyone will swipe here' is not. The revision can then be framed as a hypothesis: remove the repeated line and test whether the promise remains clear.

Confidence describes the clarity of the evidence. It does not describe the size of an assumed audience effect. Low confidence is a useful result because it shows where the codebook or source context is inadequate.

A phased study instead of one giant scrape

The original memo proposed very large samples before the construct had been tested. This protocol reverses the order: prove the measurement can be applied before scaling it.

PhaseMaterialDecision
Dry runOwned or synthetic clips chosen to include clear, negative, and ambiguous examples.Does the annotation interface preserve frames, timebase, missingness, and original ratings?
Agreement pilotA diverse rights-cleared sample, independently rated under a frozen codebook.Which codes and boundaries produce usable agreement, and which require revision or removal?
Validation studyHeld-out media plus separate viewer tasks designed for comprehension or preference.Do the labels relate to the specific construct being tested? Agreement alone is not validity.
BenchmarkA preregistered sampling frame large enough for the promised comparisons.Only now estimate distributions by format, niche, language, or platform, with uncertainty and minimum subgroup rules.
Annual updateNew material annotated with a locked or bridged codebook version.Report version effects and never present a model update as a creator trend.

How the index would be analysed

The descriptive output starts with the distribution of first coded intervals and the share of videos with no coded interval under the frozen rules. A Kaplan-Meier curve can retain those right-censored videos instead of deleting the strongest cuts from the sample. The curve would estimate survival without a coded event, not survival of viewer attention.

Platform, duration, format, language, and niche comparisons belong in secondary analysis. If a proportional-hazards model is used, its assumptions must be tested and its effect described as the relative rate of receiving a code. If the assumptions fail, the analysis needs a registered alternative rather than a prettier chart.

Report raw category distributions, missingness, boundary distance, percent agreement, and a preregistered chance-corrected agreement measure with intervals. A single pooled coefficient can hide one unusable category, so every code needs its own diagnostic.

Publication gates

The first public benchmark should stop if any of these conditions fail.

  • Media rights and consent cover annotation, storage, reviewer access, and every clip or frame intended for publication.
  • The sampling frame, codebook, exclusions, primary outcomes, and analysis have been preregistered before confirmatory annotation.
  • Independent ratings remain intact; adjudication has not overwritten disagreement.
  • Agreement and boundary precision meet thresholds chosen before the data were seen, or the report is published as a failed measurement pilot.
  • No raw identity key, private transcript, creator handle, or restricted media appears in the research dataset.
  • The headline describes coded structure and does not substitute retention, views, satisfaction, or recommendation for the measured event.

What this report will never prove by itself

A weak-second annotation is not a platform penalty. It cannot reveal why a recommendation system stopped showing a video. It cannot identify the exact person who left, and it cannot prove that removing the interval improves the work.

A later experiment could compare comprehension or preference across controlled edits. Native analytics could show audience response after publication. Both would answer different questions and carry different confounds. The index should stay useful by refusing to swallow those questions whole.

Creators can use the proposed labels today as a review vocabulary. They should not quote an index statistic until the study actually produces one.

Sources and protocol foundations

The codebook procedure follows the emphasis on explicit definitions and independent coding in Roberts and colleagues' codebook-development case study.

Agreement is treated as a property of a defined process, not proof of truth, following Krippendorff's reliability guidance and the multi-rater framework discussed by Hayes and Krippendorff.

The event-boundary logic is informed by research on continuity editing and event segmentation, while the caption code is bounded by the current WCAG 2.2 standard.

Any confirmatory study should post a timestamped plan before collection using a repository such as OSF Registrations.

Frequently asked questions

Is a weak second the moment viewers swipe away?

No. It is a reviewer-coded structural interval under a published definition. Viewer departure is a separate behavioral measure available only from the relevant platform analytics or a purpose-built study.

Why use time-to-event analysis?

Some videos may end without a qualifying code. Time-to-event methods can retain those right-censored cases instead of pretending every video fails. The event still means first coded interval, not viewer loss.

Can a deliberate pause receive a weak-second code?

Only if it meets the frozen inclusion criteria. A pause used for reading, emotion, proof, suspense, or rhythm should be excluded when that function is legible. Reviewer uncertainty must remain in the data.

Has ViralJury calculated the median weak second?

No. This page publishes the protocol before data collection. There is no median, platform comparison, or annual index value to cite yet.