← The Audit Desk

Audience-as-Judge — the study, whole

THE AUDIENCE-AS-JUDGE STUDY (working title: the Battle Rap study)
The single public home of this desk’s first experiment: does adversarial debate style win because it is hostile, or because it relocates the verdict from the moderators to the room? This page holds everything the project has — the registered protocol, and the entire working state of the build, in the open.

Status: Phase 0 — the draft protocol is public (below); the frozen version will be posted to OSF before any data is collected. Fielding is not yet funded.
How the desk announced it, in its own voice: the Special Report dispatch.
Rendered 2026-08-09T14:13:04Z · repo commit d5b218e

On this page:
1. The registered protocol — frozen; its sha256 is its signature
2. The build, in the open — phase state, gates, and the deviation log
3. The coder instrument — how the six features get scored
4. The stimuli blocker — why no stimulus text exists yet, on purpose
5. The operator kit — the human steps, pre-assembled: OSF text, coder recruitment, pilot config

1 · The registered protocol — FROZEN. sha256 of this document: 129000f909aa3732cf882b0101b4adc2a162a1fa2ab4fdaebb8b370f90bfcf71. A desk that audits overclaiming registers its own hypotheses — including the ways this study can fail — before it collects a single row.

BATTLE RAP SPEC

Audience-as-Judge: Testing Whether Adversarial Debate Style Routes Authority Away From Institutions

Status: pre-registration draft + build order Target: Claude Code on the Spark (Track A), Prolific fielding (Track B) Editorial home: The Stochastic Parrot Version: 0.2 Changelog: v0.2 (2026-08-09) adds §2.6 (stopping rule), §2.7 (exclusion criteria), and the enumerated secondary family in §2.5 — the three items §5 Phase 0 required before the freeze. No changes to hypotheses, design, indices, or kill conditions. v0.1 is preserved in the repository history.


0. The claim being tested

The folk version is "millennials like candidates who attack like battle rappers." That version is untestable as stated: it confounds a rhetorical style with a candidate type, and it treats a cohort interaction as if it were a main effect. The version this spec tests:

Adversarial debate performance affects favorability primarily through audience-authority routing — features that relocate the verdict from institutional arbiters (moderators, fact-checkers, party elites) to the room — and not through aggression per se. The effect is moderated by institutional trust, and any apparent birth-cohort effect is substantially mediated by trust and media diet. This is the battle-rap analogy taken seriously. A rap battle has no referee. You win because the crowd says you won. If that structure is what's being imported into debate performance, then the payoff should concentrate on the crowd-routing features and be flat or negative on the merely hostile ones.

Primary hypotheses

Registered kill conditions

State these before any data is collected. They are what makes a null publishable rather than embarrassing. - If H2 fails — aggression alone carries the effect — the mechanism story is dead. The finding becomes "adversarial affect, not audience routing," and the piece is written that way. - If H1's contrast is not significant but both indices are, the features are not separable at this sample size. Report as underpowered, do not reinterpret. - If H3 fails and cohort survives trust and media diet, report that. "It really is generational" is a legitimate outcome and should not be argued away.


1. Construct definition: the six features

Do not score "aggression" as a single dimension. Score six binary features per exchange, each span-grounded to the transcript, then compose two pre-registered indices.

Feature codebook

Each feature is coded 0/1 at the exchange level (one candidate's contiguous speaking turn responding to or targeting another candidate). Every 1 requires a supporting character span. F1 — Direct second-person address, target present. Speaker addresses the opponent as "you" while that opponent is on stage. Positive: "You said that on this stage four years ago." Negative: Third-person reference to an absent figure. Second-person addressed to the moderator or the audience (that is F5). Rule-detectable: partially. spaCy dependency parse for 2nd-person pronoun subjects with a person-entity antecedent in the prior 3 turns. LLM adjudicates target identity. F2 — Setup/reveal structure. The turn contains a deliberate delay between premise and payoff — the point is withheld for at least one clause boundary and then landed. Positive: A statistic offered flatly, then the reversal that makes it damning. Negative: Assertion followed by elaboration. The test is whether removing the final clause destroys the meaning of the earlier ones. Rule-detectable: no. LLM-only, with an explicit "would the setup be inert without the reveal?" decision rule. F3 — Characterological targeting. The attack is on who the opponent is — their consistency, courage, authenticity, class position — rather than on a policy position or record outcome. Positive: "You'll say anything in this room." Negative: "Your tax plan adds four trillion." Record-based attacks with a policy referent are F3=0 even when hostile. Boundary: Hypocrisy attacks are F3=1 when the point is the hypocrisy, F3=0 when the point is the policy. F4 — The flip. The speaker reuses the opponent's own framing, phrasing, or attack line and redirects it. Requires prior-turn context in the coding window. Positive: Opponent says "career politician"; speaker returns it against them within the same exchange or the next. Negative: Independent coincidental phrasing. Requires lexical or structural overlap with a span in a prior turn, cited by the coder. Rule-detectable: partially — n-gram overlap with prior 5 turns as a candidate generator, LLM confirms intent. F5 — Audience-facing performance. The speaker plays to the room rather than to the moderator or the camera-as-neutral-record. Explicit audience address, pausing for reaction, soliciting response. Positive: "Ask the people in this room whether that's true." Negative: Rhetorical questions with no audience referent. Note: Transcript-only coding will underdetect this. See §3.3 on the audio/video secondary channel. F6 — Frame-breaking. The speaker names the format, the rules, the media, or the artifice while inside it — declaring the arbiter illegitimate or the process rigged. Positive: "This question is exactly what I'd expect from this network." "We all know what this format is for." Negative: Ordinary complaints about time limits. The test is whether the arbiter's legitimacy is at stake, not merely their timekeeping.

Composed indices


2. Track B is primary: the vignette experiment

Track A cannot identify this. Aggressive candidates are systematically younger, more online, more anti-institutional, and more often insurgents. Style and candidate type are collinear in every naturally occurring debate corpus. The experiment is the study; the corpus is the scaffolding. Build Track B first.

2.1 Design

2 × 2 between-subjects factorial, aggression × audience-authority. This is the design choice that isolates the mechanism — a four-level ordinal "intensity ladder" would not. | Condition | AGGR | AUTH | Description | |---|---|---|---| | C1 Baseline | low | low | Policy contrast, no personal targeting, no audience routing | | C2 Aggression-only | high | low | Personal/characterological attack, addressed to opponent and moderator, no crowd appeal, no frame-breaking | | C3 Authority-only | low | high | Mild policy content, but plays to the room and names the format as rigged | | C4 Full package | high | high | Both | Held constant across all four: policy content, candidate name, party cue, apparent age, race, gender, word count (±8%), reading level (±0.5 Flesch-Kincaid grade), and the opponent's preceding turn.

2.2 Partisan symmetry — non-negotiable

Every respondent sees the stimulus with a randomly assigned party cue (D / R / no-party-stated), crossed with condition. Eight-cell party-by-condition randomization. This exists because the finding is worthless — and worse, dishonest — if the design can only detect the effect for one side's style. Pre-register the party-symmetry test: if the style effect differs significantly by party cue, that difference is a reported finding, not a footnote. Any writeup that shows the effect for one party must show the other party's estimate at equal prominence. Two stimulus sets (a Democratic-primary-flavored issue frame and a Republican-primary-flavored one), each rendered in all four conditions, to avoid issue-content confounds with party cue.

2.3 Measures

Dependent variables 1. Candidate feeling thermometer, 0–100 (primary DV). 2. Vote intention vs. the opponent, forced choice. 3. "Who won this exchange?" — mechanism check. 4. Perceived authenticity, 5-item. 5. Perceived fairness of the format — manipulation check for AUTH. Moderators (measured, not assigned) - Institutional trust: 4–6 item battery covering media, elections, Congress, courts. Use an existing validated battery rather than writing one. - Media diet: short-form video hours/week; primary news source; podcast consumption; whether the respondent typically encounters debates as full broadcasts or as clips. That last item is the sleeper variable and should be worded carefully. - Birth year (continuous — do not collect cohort as a category and then treat it as ordinal). - Political interest, partisanship strength, education, gender, race. Manipulation checks: two items confirming perceived aggression and perceived audience-orientation, used to validate the stimuli in a pilot before full fielding.

2.4 Sample and power

Detecting a difference between two interaction slopes is expensive. Interaction effects generally require ~4× the N of the corresponding main effect, and H1 is a contrast between two of them. - Pilot: n = 200, purpose is manipulation-check validation only. No hypothesis testing. If the AGGR and AUTH manipulations do not separate cleanly on their checks, revise stimuli and re-pilot. - Main field: n = 2,000, quota-stratified on birth year (four strata, oversampling 1981–1996 and 1997–2005 to ~600 each), party ID, gender, and education. - Powered for a small-to-moderate interaction contrast (f² ≈ 0.02) at 80%. - Budget at roughly $1.50–2.00 per respondent for a 6–8 minute instrument: $3,000–4,000 for the main field, plus ~$400 for the pilot. If this is out of range, say so now — a 700-person version can test main effects honestly but cannot test H1, and running it while claiming otherwise is the exact overclaiming the Parrot audits.

2.5 Analysis plan

Primary model, OLS on the thermometer with robust SEs:

FT ~ AGGR * TRUST + AUTH * TRUST + FLIP
     + AGGR * AUTH
     + PARTY_CUE * AGGR + PARTY_CUE * AUTH
     + BIRTH_YEAR + covariates

H1 test: bootstrap the contrast β(AUTH × TRUST) − β(AGGR × TRUST), 10,000 resamples, report the CI on the difference. Pre-register this single contrast as the primary test. Everything else is secondary and labeled as such. H3 test: nested models. Fit cohort × style without trust and media diet, then with. Report the proportion of the cohort interaction coefficient that survives, with a bootstrap CI on the attenuation ratio. Do not use a Sobel test. Multiple comparisons: one primary test. Everything else is secondary — Benjamini-Hochberg at q = 0.05 across the following family, closed here and not amended after fielding:

Fifteen secondary tests, named before any data exists. A test not on this list is exploratory, is labeled exploratory in the writeup, and cannot be promoted after the fact (§7, post-hoc feature promotion).

2.6 Stopping rule

2.7 Exclusion criteria

Applied before any outcome column is read, using quality variables only; counts reported per condition; a sensitivity analysis with excluded respondents included ships in the appendix.

  1. Failed either of the two attention checks.
  2. Completion time under 40% of the pilot median duration.
  3. Duplicate participant ID or IP address (first submission kept).
  4. Straightlining: zero variance across the thermometer plus all Likert batteries combined.
  5. Self-reported non-serious responding on the end-of-survey commitment item.
  6. Missing primary DV (thermometer).

An exclusion rule not on this list may not be invented after fielding; if one proves necessary (e.g., a platform-level fraud event), it is applied symmetrically across conditions and disclosed in the deviation log with the piece.

3. Track A: the observational corpus

Scaffolding and descriptive color. It cannot carry a causal claim and the writeup must not imply that it does.

3.1 Corpus

3.2 The normalization trap, again

The 3.4× agentless-passive-by-category finding applies here in a new form: debate topic drives attack rates. Immigration and crime segments will run hot on F1/F3 regardless of speaker. Never rank speakers on raw feature rates. - Topic-code every exchange (fixed taxonomy, ~12 categories). - Report within-topic comparisons and direct-standardized rates. - Validate with a mixed-effects model: random intercepts for speaker, debate, and topic. - Calibration gate: before spending LLM tokens on the full corpus, confirm that topic explains a substantial share of raw feature-rate variance on a 300-exchange subsample. If it doesn't, the topic coder is broken — fix it before proceeding.

3.3 The clip-selection problem — read this before using engagement as a DV

Engagement-per-view on clipped debate moments is conditioned on the clip having been clipped. Clippers select for exactly the features being measured. This is not a nuisance parameter; it is a mechanism that can generate the entire predicted result out of nothing. Mitigations, in order of preference: 1. Preferred: sample exchanges from the full transcript, then check which ones were clipped. Model clipping probability itself as an outcome (CLIPPED ~ AUTH + AGGR + FLIP + topic + speaker). This turns the confound into a finding: it measures whether the media ecosystem selects for audience-authority routing, which is a Parrot piece on its own. 2. Restrict engagement analysis to official full-debate uploads where segment-level retention is available. 3. If neither is available, drop the engagement DV entirely. Report feature prevalence descriptively and stop. Platform age-skew is not a cohort measure. If it appears in the writeup at all, it appears with that sentence attached.

3.4 Dedup

Same-clip reuploads across accounts are the wire-copy problem in new clothes. Perceptual hash on video, plus transcript n-gram overlap ≥ 0.9 within a 14-day window. One canonical row per underlying exchange, engagement summed with the aggregation rule logged.

4. Reliability protocol

Matches the existing standard. No rankings, no rates, and no LLM-coded variable enters any model before this gate passes. - Two human coders independently code a 250-exchange stratified sample. - Krippendorff's α ≥ 0.70 per feature, computed separately for each of F1–F6. A pooled α is not acceptable — F2 and F4 are the hard ones and will hide behind F1's easy agreement. - Any feature below 0.70: revise the codebook decision rules, re-train, re-code a fresh sample. Two failed rounds on a feature means that feature is dropped from the indices and the drop is disclosed. - LLM coder is then validated against the human-consensus set. Report LLM-vs-human α per feature alongside human-vs-human. If the LLM underperforms the human floor, it does not get to code the corpus. - Every LLM code returns a character span. Codes without spans are discarded, not repaired.


5. Build order

Phase 0 — Pre-registration. Hypotheses, the single primary contrast, the enumerated secondary family, kill conditions, stopping rule, exclusion criteria. Timestamped and posted (OSF) before Phase 2 data collection. This is the whole credibility of the project; a study auditing overclaiming that pre-registers after peeking is not recoverable. Phase 1 — Codebook + reliability. §1 and §4. Human coding, α gate, LLM validation. Gate: all retained features ≥ 0.70. Phase 2 — Stimulus construction + pilot. Write eight stimuli (4 conditions × 2 issue frames), verify they differ on the intended dimensions and only on those, by human coding with the §1 instrument. Field the n=200 pilot. Gate: manipulation checks separate. Phase 3 — Main field. n = 2,000. Analysis per §2.5. Gate: none — the result is the result. Phase 4 — Corpus. Track A on the Spark. Feature-score the transcript corpus, topic-normalize, run the clipping-probability model. Descriptive only. Phase 5 — Writeup. Both tracks, experiment leading. Publish the pre-registration diff — every deviation from Phase 0, itemized.


6. Acceptance criteria

  1. Pre-registration is public and timestamped before Phase 2 fielding.
  2. All retained features clear α ≥ 0.70 for both human-human and LLM-human.
  3. Manipulation checks separate AGGR and AUTH in the pilot.
  4. Main field hits quota targets within 10% per cohort stratum.
  5. The primary contrast is reported with its CI regardless of direction, in the first three paragraphs of any writeup.
  6. Party-symmetry estimates are reported at equal prominence.
  7. Every observational claim is labeled observational in the sentence that makes it.
  8. Deviation log published with results.

7. What kills this study


8. Publication paths by outcome

Outcome Headline
H1 and H2 hold The debate stage is being scored like a rap battle by low-trust voters — and generation was never the variable
H2 fails, aggression carries it Voters like a fighter. The mechanism is simpler and worse than the elegant version
H1 null, both indices positive Style matters, the components don't separate — a replication call with the instrument attached
H3 fails, cohort survives The generational read holds up under controls that should have killed it
Everything null The instrument, the codebook, and the reliability data published as a tool — plus the clipping-probability model from Phase 4, which is a piece on its own
Every row is publishable. That is the point of registering it this way.

2 · The build, in the open — LIVING DOCUMENT: this section changes as the work moves, and is not covered by the hash above. Every deviation from the protocol lands in the log below, dated.

BATTLE RAP — builder handoff

Read SPEC.md first. The spec is the contract; this file is only the state of the build. Any builder session (Claude on the Mac, Claude on the Spark, DeepSeek scoping pass) works from this file and updates it when a gate is passed or a deviation is logged.

Ground rules carried over from the desk

State (update in place)

Money gates (Mike-only)

Deviation log

(append-only; every departure from SPEC.md v0.1, dated) - none yet


3 · The coder instrument — living document; frozen only when it clears its reliability gate.

Coding instrument — six features, exchange level

The construct definitions live in ../SPEC.md §1 and are NOT restated here (one definition, one place). This file adds only what a coder needs operationally.

Unit

One exchange = one candidate contiguous speaking turn responding to or targeting another candidate. Moderator-only turns are not units. An exchange keeps its prior 5 turns as coding context (F4 needs them).

Code record schema (one JSON object per exchange per coder)

{
  "exchange_id": "...", "coder": "...", "ts": "...",
  "F1": 0, "F2": 0, "F3": 0, "F4": 0, "F5": 0, "F6": 0,
  "spans": {"F3": {"start": 0, "end": 0, "quote": "..."}},
  "notes": ""
}

RULES: every feature scored 1 MUST have a span entry whose quote matches the transcript verbatim at [start,end). A 1 without a span is invalid — the harness rejects the record, it does not repair it (SPEC §4).

Decision rules coders trip on (from the spec, restated as tests)


4 · The stimuli blocker — the file whose job is to stop us writing test material too early.

Stimuli — DO NOT DRAFT YET

Eight stimuli (C1–C4 × two issue frames) per SPEC §2.1–2.2. Blocked until the codebook clears its α gate (see HANDOFF.md — the same instrument validates the stimuli; freezing it first prevents teaching to the test).

Hard constraints across all four conditions of a frame: policy content, candidate name, party cue slot, apparent age/race/gender, word count ±8%, Flesch-Kincaid grade ±0.5, identical opponent preceding turn. Party cue is a render-time variable (D / R / none), never baked into the text. Pilot includes a humor-rating item (SPEC §7, stimulus-artifact kill).


5 · The operator kit — the steps only humans can take, pre-assembled so each is one action from done.

Faster than email: the desk keeps a Slack room (#the-desk) — ask for an invite and say you want to code.

Second coder — recruitment one-pager

The reliability protocol (SPEC §4) requires TWO human coders working independently. Coder A can be Mike; Coder B is the open seat.

The job: code 250 debate exchanges (form_coder_B.md) on six binary features per the instrument (codebook/CODEBOOK.md). Every "1" requires a verbatim supporting quote. Realistic effort: 8–12 hours; can be split across a week. The two coders must not discuss items until both files are submitted.

Who fits: anyone careful with text — a grad student, a journalist friend, a retired lawyer. Political knowledge helps; strong priors about the people quoted do not (the instrument codes form, not politics).

Options: (a) someone Mike knows, paid or favor; (b) Prolific's vetted "expert" pool (post as a transcription/annotation task, ~$12–15/hr, budget ~$150); (c) both coders external and Mike codes nothing (cleanest optics — the desk's operator not coding his own study).

When both code files exist: run python3 scripts/reliability.py codes_A.jsonl codes_B.jsonl — the α gate report is automatic. Features under 0.70 trigger the SPEC §4 revise-retrain-recode loop.

OSF Registration — ready to paste

This file is the freeze, pre-assembled. Mike: create a registration at osf.io (template: "OSF Preregistration"), paste the sections below into the matching fields, and submit. The moment it has a DOI, put the DOI in STATUS.json and the study page header; Phase 2 unblocks.


Title: Audience-as-Judge: Testing Whether Adversarial Debate Style Routes Authority Away From Institutions

Authors: The Stochastic Parrot desk (automated), M. Portney (operator).

Description: We test whether adversarial debate performance affects candidate favorability through audience-authority routing — features that relocate the verdict from institutional arbiters to the audience — rather than through aggression per se, with institutional trust as moderator and birth cohort deliberately demoted to a covariate. Full protocol, version-controlled with hash: https://thestochasticparrot.com/research/battle-rap-prereg/

Hypotheses: SPEC.md §0 verbatim — H1 (concentration: β(AUTH×TRUST) > β(AGGR×TRUST), pre-registered as a single coefficient contrast), H2 (aggression null-to-negative in all strata), H3 (cohort × style attenuates ≥50% when trust and media diet enter). Registered kill conditions: if H2 fails, the mechanism story is dead and is reported that way; if H1's contrast is non-significant while both indices are positive, report as underpowered without reinterpretation; if H3 fails, "it really is generational" is reported as the finding.

Design: 2×2 between-subjects factorial (aggression × audience-authority), crossed with randomized party cue (D/R/none) — eight cells — two issue-frame stimulus sets. Held constant across conditions: policy content, candidate identity cues, word count (±8%), Flesch-Kincaid grade (±0.5), opponent's preceding turn. (SPEC §2.1–2.2.)

Sampling plan: Prolific. Pilot n=200 (manipulation-check validation only, no hypothesis tests). Main field n=2,000, quota-stratified on birth year (four strata, 1981–1996 and 1997–2005 oversampled to ~600 each), party ID, gender, education. Stopping rule per SPEC §2.6: 2,000 approved completes or 30 days; no interim analyses; shortfalls reported per stratum.

Variables: DVs — 0–100 feeling thermometer (primary), forced-choice vote intention, "who won this exchange," 5-item perceived authenticity, perceived format fairness (AUTH manipulation check). Moderators — validated institutional-trust battery; media diet including clip-vs-broadcast item; birth year (continuous); political interest, partisanship strength, education, gender, race. (SPEC §2.3.)

Analysis plan: OLS on the thermometer with robust SEs: FT ~ AGGRTRUST + AUTHTRUST + FLIP + AGGRAUTH + PARTY_CUEAGGR + PARTY_CUE*AUTH + BIRTH_YEAR + covariates. Primary test: bootstrap contrast β(AUTH×TRUST) − β(AGGR×TRUST), 10,000 resamples, CI on the difference. H3 via nested models with bootstrap CI on the attenuation ratio (no Sobel). Secondary family: the 15 tests enumerated in SPEC §2.5, Benjamini-Hochberg q=0.05, family closed at registration. Exclusions per SPEC §2.7 (attention checks, speed, duplicates, straightlining, commitment item, missing DV), applied before outcome columns are read.

Other: Track A (observational debate-transcript corpus) is scaffolding and descriptive only; it cannot carry causal claims (SPEC §3). Reliability gate: Krippendorff's α ≥ 0.70 per feature, human-human and LLM-human, before any coded variable enters any model (SPEC §4). Every deviation from this registration will be published as an itemized diff with the results.

Prolific pilot — config outline (n=200, ~$400)

BLOCKED until: OSF freeze posted + stimuli written and validated (stimuli are themselves blocked until the codebook clears its α gate). This file exists so the fielding day is configuration, not design.