Original research protocol

I Defined a Caption Onset Benchmark Without Chasing a Magic Millisecond

A proposed caption-event benchmark that records timing against three reference points and refuses to turn one median into a universal editing rule.

The first text in a video can name the subject, transcribe speech, label an object, establish a question, or merely display a watermark. Treating all five as 'caption onset' produces a precise timestamp for the wrong thing.

The useful unit is a caption event with a job. The study has to identify what the text represents, when that speech or visual becomes available, whether the caption blocks essential information, and whether the reading window is workable. Only then does the timestamp mean anything.

This protocol is deliberately descriptive. It can show how captions are timed in the sampled material. It cannot infer that the earliest caption wins, that viewers watched on mute, or that a platform rewarded the edit.

One caption, three reference points

A subtitle can be perfectly synchronized with the first phoneme and still arrive before the sentence means anything. Hook text can appear before speech on purpose. A product label can be timely in audio terms but cover the object it names. That is why one zero point is insufficient.

Reviewers should mark acoustic, semantic, and visual onset independently. Differences are calculated only after the boundaries are locked. A negative value means the caption leads the reference; a positive value means it follows. Neither sign is automatically good or bad.

The unit is one continuous displayed text state before its words, speaker assignment, or meaningful sound description changes. Word-by-word animation can be retained as sub-events so the study does not pretend a sentence appeared all at once.

Reference points for every caption event

ReferenceBoundaryWhy it is separate
Acoustic onsetThe first detectable point of the represented speech or meaningful sound under a frozen audio-inspection rule.Measures synchronization with the sound source.
Semantic onsetThe earliest point at which the represented idea is specific enough to identify.Fillers and lead-ins may begin before the caption's full meaning exists.
Visual onsetThe first frame containing the object, action, proof, label, or state referenced by the text.A caption can be synchronized to speech while leading or trailing the image it explains.
Caption onsetThe first source frame in which the caption event is visibly present under the registered opacity rule.This is the measured event, compared with the three references above.
Caption exitThe last source frame before the text state changes or disappears.Late persistence can cross speakers, propositions, shots, or visual states.
Not applicableA registered state when a caption has no acoustic, semantic, or visual counterpart of that type.Missing reference types should not be invented to complete a spreadsheet.

Meaningful text needs a taxonomy

Speech subtitles, sound descriptions, hook text, persistent headlines, topic labels, product labels, speaker labels, calls to action, environmental text, usernames, watermarks, and platform interface text should not share one field. The primary benchmark includes text that carries the video's accessible or narrative meaning. It stores the rest separately.

Environmental text becomes meaningful only when the edit recruits it into the story, for example by pointing to a sign or zooming into a label. A logo does not trigger narrative onset merely because OCR can read it. Platform interface text should be masked or classified before creator-intended placement is analysed.

Automated OCR can propose boxes and strings, but semantic type, obstruction, and ambiguous visual references need human review. Any model confidence belongs beside, not in place of, reviewer confidence.

The minimum event record

The public dataset should contain timing and coded conditions, not creator identity or unrestricted source media.

Field groupStoreReason
TimingSource timebase, onset and exit frames, reference frames, derived signed offsets.Frame coordinates remain reproducible when a millisecond conversion changes.
TextDisplayed text, language, speaker or sound identity, text type.Supports segmentation and meaning review.
PlacementBounding box, line count, motion, placement region, obstruction code.A timely caption can still hide essential evidence.
Reading conditionsCharacter count, display duration, contrast review, background motion, concurrent visual demand.Readability cannot be reduced to one words-per-minute threshold.
ReviewConfidence, short reason, reviewer ID, codebook version, adjudication link.Preserves uncertainty and version history.
RightsA separate rights ID and permitted-use record.The identity key and consent documents stay outside public analytical data.

Sampling starts with permission

Use media ViralJury owns, fully synthetic clips, material licensed for research, or participant submissions with explicit permission for annotation, reviewer access, storage, and intended publication. A video being publicly viewable does not automatically grant permission to download it, reproduce frames, redistribute transcripts, or publish a derived dataset.

The pilot should vary speech rate, language, speaker count, caption style, visual motion, and content type only to the extent the design can support those comparisons. A convenience sample can test the instrument. It cannot stand in for all Shorts, Reels, and TikToks.

Deaf and hard-of-hearing reviewers and captioning specialists belong in the protocol design, paid annotation, and interpretation. Their role is not to bless one timer setting. It is to help define what the study observes and what a laboratory record misses.

Reviewer procedure

The order keeps later interpretation from silently moving earlier boundaries.

Confirm the source export, timebase, calibrated display size, player, audio setup, and permitted use.

Watch the full clip once, then mark each caption event without seeing model flags or other reviewers' work.

Mark acoustic references under the frozen audio rule, then semantic and visual references with reasons and confidence.

Code speaker, segmentation, obstruction, placement, contrast, line breaks, background motion, and reading conditions.

Lock the independent submission. Calculate offsets, agreement, and tolerance summaries afterward.

Use adjudication to improve the codebook while preserving every original rating and disagreement type.

How the benchmark would be reported

Start with distributions of caption onset against each reference, not a single overall mean. Report medians, quantiles, uncertainty, missing references, and the share of events that lead, align with, or follow under every preregistered tolerance.

Obstruction and difficult reading conditions are separate outcomes. A caption can have a zero-millisecond offset and still cover a chart value. Another can lead speech and improve orientation. The report should show the combinations instead of crowning the earliest style.

Platform, language, and format comparisons are secondary. They need enough independent material, a minimum reporting group, and explicit control for repeated templates or creators. No regression should translate structural payoff detection into actual viewer retention without a study that measures viewers.

Claims this benchmark can support

If the protocol is completed, the report may say:

  • In the registered sample, the median meaningful caption onset relative to a specified reference was a stated value with an uncertainty interval.
  • A stated share of caption events obstructed information reviewers coded as essential under the published rubric.
  • Agreement was stronger for acoustic boundaries than semantic or visual boundaries, if that pattern appears in the data.
  • Results differed across the sampled languages, formats, or caption types, with subgroup sizes and corrections disclosed.
  • Changing the frame tolerance changed the classification distribution by a reported amount.

Claims this benchmark cannot support

  • All social video is watched muted.
  • Captions must appear on frame zero.
  • A specific offset causes viewers to leave.
  • One caption speed is accessible for every language, viewer, device, and visual task.
  • The platform recommends a video because its captions fit the benchmark.

Sources and accessibility foundations

WCAG 2.2 is the current normative web accessibility reference used here. It requires captions in defined synchronized-media contexts but does not prescribe a universal short-form onset millisecond.

The W3C's Synchronization Accessibility User Requirements discusses synchronization needs but explicitly is not a finished conformance standard. FCC caption rules supply another context-specific reference without creating one timer for every social clip.

Research on subtitle segmentation and reading supports recording segmentation and presentation conditions rather than treating timing alone as readability.

Section508.gov caption guidance informs the accessibility-review checklist. It does not replace the legal or accessibility review applicable to a specific publication.

Frequently asked questions

Should captions start before speech?

Sometimes. Hook text may intentionally lead speech, while subtitles usually represent speech or sound. The benchmark records the relationship and its function before judging it. A negative offset is not automatically an error.

Does WCAG require captions on the first frame?

No universal first-frame rule appears in WCAG 2.2. The standard addresses caption availability and accessibility outcomes in defined contexts. Synchronization, accuracy, placement, and readability still require context-specific review.

Why not use OCR alone?

OCR can find text boxes and propose strings. It cannot reliably decide whether a sign is environmental, whether hook text carries the premise, whether a label blocks essential evidence, or which visual event the words explain.

Has ViralJury measured caption onset across platforms?

No. This page is the protocol. There is no platform median, cross-platform ranking, or retention effect to cite yet.