Original research protocol
I Designed a Creator Video Self-Assessment Study That Does Not Let the Tool Grade Itself
A protocol comparing creator self-diagnosis, blinded expert consensus, and ViralJury output without treating the analyzer's own label as ground truth.
After the twentieth replay, an editor no longer watches the opening like a stranger. They remember the missing setup, anticipate the payoff, and hear the sentence before the caption completes. That familiarity may create a diagnostic gap—or it may give the creator context an outsider lacks.
I designed a study that can test both possibilities without letting ViralJury declare itself correct. Fifty creators would diagnose one underperforming video of their own and one unfamiliar clip before seeing any report. Three experienced reviewers, blinded to the creator, analytics, and analyzer output, would code the same material.
The experts' adjudicated consensus becomes a human reference, not metaphysical truth. ViralJury is scored separately against it. The study asks where diagnoses agree, where they diverge, how confident each person was, and whether an outside report changes the edit the creator intends to make.
The analyzer cannot be its own ground truth
If ViralJury selects a weak timestamp and the creator selects another, calling the creator wrong would bake the product claim into the scoring rule. The same circularity appears if an AI explanation is graded only by another instance of the same system.
NIST's work on evaluation probes describes checking system claims against human-curated reference documents and producing an audit trail. Broader NIST guidance also emphasizes reference measures, independent assessors, domain experts, users, and representative human-subject evaluation.
For this study, the reference is three qualified human reviewers following a frozen taxonomy and adjudication procedure. ViralJury sees the same clips but not the expert labels. Agreement shows alignment with that reference; it does not establish a single objectively correct edit.
The creator diagnosis task
Every response is locked before outside feedback appears.
The creator supplies one eligible underperforming video, intended audience, intended promise, and the outcome evidence used to call it underperforming.
Before seeing analytics curves or any report, the creator selects the weakest timestamp or interval.
The creator chooses one primary failure category, ranks up to three suspected problems, and states the edit they would make.
The creator records confidence from 0 to 100 and explains the evidence behind the diagnosis.
The creator repeats the task on an unfamiliar, rights-cleared clip under the same conditions.
Only then does the creator see the structural report and record any changed timestamp, category, confidence, or edit intent.
The fixed diagnosis taxonomy
| Primary category | Operational question | Common confusion |
|---|---|---|
| Hook clarity | Does the opening identify the subject, situation, or reason to continue? | A visually active frame can still be vague. |
| First-frame clarity | Can a cold viewer identify what they are looking at before later context arrives? | A creator recognizes their own footage automatically. |
| Caption coordination | Does text arrive, remain, and sit where it can explain the relevant speech or visual? | Timing, accuracy, and placement are separate failures. |
| Pacing or weak second | Does an interval spend attention without adding information, stakes, proof, or emotional progress? | Slow is not always weak; buildup can be productive. |
| Promise mismatch | Does the opening imply a result the video does not deliver? | A clear hook can still promise the wrong thing. |
| Payoff distance | Is the primary proof or reveal delayed relative to what the opening asks viewers to wait for? | The nearest exciting moment may not be the promised payoff. |
| Audience fit | Does the edit assume context, vocabulary, or fandom knowledge the intended viewer lacks? | Experts may not share the target audience's knowledge. |
| No structural diagnosis | Is the available evidence insufficient or the likely issue outside the observable edit? | The taxonomy must allow 'unknown' rather than force a product-shaped answer. |
Three blinded experts create a reference, not a verdict
Reviewers should have documented short-form editing or audience-research experience and receive training on a separate set. They remain blind to the creator diagnosis, platform outcome, and ViralJury report while assigning timestamp, category, ranked alternatives, confidence, and a short rationale.
Their original ratings are preserved. Adjudication happens only after independent submission and follows written rules. The final reference may use majority category, tolerance-window timestamp overlap, and a documented resolution when all three disagree.
Inter-reviewer reliability must be published. If the experts cannot reach the preregistered threshold, the study has discovered taxonomy ambiguity—not creator failure. The correct response is to refine or split the category, not average the disagreement away.
Own-video and unfamiliar-video tasks test different explanations
Creators may struggle with their own work because intention fills missing context and repeated viewing makes every beat familiar. They may also outperform outsiders because they know the target audience, source footage, brand constraints, and story the edit is trying to preserve.
An unfamiliar-video control separates self-attachment from general diagnostic skill. If creator-expert agreement is lower only on the creator's own video, familiarity or attachment becomes a plausible explanation. If agreement is equally low on unfamiliar clips, the issue may be taxonomy, training, or broad diagnostic difficulty.
The unfamiliar clips must be rights-cleared and matched for platform, niche, language, and complexity where possible. A simple recipe Short cannot serve as the control for a dense multilingual fan edit and then be used to explain the difference.
Primary and secondary outcomes
| Outcome | Measure | What it answers |
|---|---|---|
| Primary category agreement | Creator versus adjudicated expert top category. | Do they identify the same kind of structural problem? |
| Timestamp overlap | Registered tolerance-window overlap or distance between selected intervals. | Do they point to the same moment even when labels differ? |
| Ranked-category agreement | Agreement across up to three ordered suspected problems. | Is disagreement only about priority? |
| Confidence calibration | Confidence compared with reference agreement and ambiguity. | Are confident diagnoses more likely to align? |
| Edit-intent change | Before-versus-after timestamp, category, confidence, and proposed edit. | Does outside feedback change the planned revision? |
| Analyzer-versus-expert | ViralJury output scored independently against the same human reference. | How does the tool align without grading itself? |
| Own-versus-unfamiliar difference | Within-creator change across the two tasks. | Is the gap specific to one's own work? |
A changed mind is not a correct mind
After seeing a confident report, a creator may change their answer because the explanation is persuasive, because the interface looks authoritative, or because it names evidence they genuinely missed. The before-and-after diagnosis can measure influence, but not correctness by itself.
The study should therefore record the creator's reason for changing and whether the new answer moves toward expert consensus. It should also record cases where the creator rejects the report and later explains a constraint the reviewers missed.
Confidence can move in either direction. A useful report may reduce overconfidence by exposing ambiguity, or increase confidence by clarifying a specific edit. Treating every increase as success would reward certainty rather than calibration.
Self-assessment research supports the comparison, not the conclusion
A systematic review of video-based self-assessment defined accuracy through direct comparison with an external evaluator. Five of nine included studies found improvement after video interventions, while methods and outcomes varied too much for a simple pooled conclusion.
A related review of surgical self-assessment warns that the external score itself must be reliable: multiple observers, appropriate blinding, inter-rater reliability, and justified expert selection are necessary. Poor reference measurement can make the participant look inaccurate when the benchmark is unstable.
Those domains are not short-form editing. They provide design lessons: define self-assessment against an external standard, validate that standard, and avoid turning mixed evidence into a universal creator deficit.
Bias controls the study needs
- Freeze creator diagnoses before showing analytics curves, expert labels, or ViralJury output.
- Blind experts to creator identity, intent response, outcomes, and analyzer output during primary coding.
- Use a separate training set and a versioned taxonomy with edge examples.
- Randomize unfamiliar control clips and balance platform, niche, language, and complexity.
- Preserve every original expert rating before adjudication and publish inter-reviewer reliability.
- Score ViralJury separately; never use its category as the reference label.
- Pre-register exclusions, timestamp tolerance, missing data, agreement metrics, and multiplicity handling.
- Allow 'insufficient evidence' so neither people nor tool are forced to diagnose every video.
Report confusion matrices, not one flattering percentage
A single agreement percentage hides which categories are confused. Creators and experts may agree on most clear hooks while repeatedly split between hook clarity and payoff distance. ViralJury may locate the right timestamp but assign the wrong label. Those patterns matter more than one score.
Publish the full confusion matrix, per-category precision and recall against the human reference, timestamp distributions, confidence calibration, and own-versus-unfamiliar differences. Include adjudication examples using recreated or consented material.
Results should remain participant-level where possible. Experience, niche, editing role, language, and audience familiarity may explain heterogeneity, but subgroup analysis needs preregistered minimums and privacy protection. Fifty creators cannot support an unlimited set of slices.
Claims this protocol cannot support
- Creators miss the weak second a specific percentage of the time.
- Expert consensus is objective truth.
- A creator who disagrees with ViralJury is wrong.
- Changing an edit plan after the report proves the report was correct.
- High confidence means an accurate diagnosis.
- Performance outcomes can be inferred from diagnostic agreement alone.
Consent, publication gate, and product boundary
Participants must control whether their video, frames, niche, diagnosis, and quotes can be published. The consent form should define reviewer access, storage, de-identification, withdrawal deadline, and the possibility that expert or analyzer feedback conflicts with the creator's intent.
The publication gate is 50 completed creators, three blinded experts, preregistered reference rules, acceptable inter-reviewer reliability, an unfamiliar-video control, and a separate analyzer-versus-expert result. Missing cases and taxonomy failures must remain visible.
ViralJury is built around the idea that creators can stop seeing a repeated draft like a cold viewer. This protocol tests that philosophy rather than assuming it. Until participants complete the study, it remains a product hypothesis—not a measured customer win.
Frequently asked questions
Do creators misdiagnose their own short-form videos?
No ViralJury result exists yet. The protocol is designed to compare creator diagnoses with a blinded, reliability-checked expert reference and an unfamiliar-video control before making that claim.
Why not use ViralJury's diagnosis as the correct answer?
That would be circular. The tool is one system being evaluated. Three blinded experts create an independent human reference, and their disagreement is published rather than hidden.
Are the experts always right about a creator's video?
No. Expert consensus is a reproducible reference, not objective truth. Creators may know audience, rights, brand, or story constraints that outsiders lack, so rationales and unresolved disagreements remain part of the result.
What would count as a useful ViralJury result?
Useful evidence could include reliable alignment with the expert reference, correct timestamp localization, calibrated uncertainty, and creator edit changes supported by stated evidence—not simply making creators agree with the tool.