Audience-as-Judge — the study, whole
1 · The registered protocol — FROZEN. sha256 of this document: 129000f909aa3732cf882b0101b4adc2a162a1fa2ab4fdaebb8b370f90bfcf71. A desk that audits overclaiming registers its own hypotheses — including the ways this study can fail — before it collects a single row.
BATTLE RAP SPEC
Audience-as-Judge: Testing Whether Adversarial Debate Style Routes Authority Away From Institutions
Status: pre-registration draft + build order Target: Claude Code on the Spark (Track A), Prolific fielding (Track B) Editorial home: The Stochastic Parrot Version: 0.2 Changelog: v0.2 (2026-08-09) adds §2.6 (stopping rule), §2.7 (exclusion criteria), and the enumerated secondary family in §2.5 — the three items §5 Phase 0 required before the freeze. No changes to hypotheses, design, indices, or kill conditions. v0.1 is preserved in the repository history.
0. The claim being tested
The folk version is "millennials like candidates who attack like battle rappers." That version is untestable as stated: it confounds a rhetorical style with a candidate type, and it treats a cohort interaction as if it were a main effect. The version this spec tests:
Adversarial debate performance affects favorability primarily through audience-authority routing — features that relocate the verdict from institutional arbiters (moderators, fact-checkers, party elites) to the room — and not through aggression per se. The effect is moderated by institutional trust, and any apparent birth-cohort effect is substantially mediated by trust and media diet. This is the battle-rap analogy taken seriously. A rap battle has no referee. You win because the crowd says you won. If that structure is what's being imported into debate performance, then the payoff should concentrate on the crowd-routing features and be flat or negative on the merely hostile ones.
Primary hypotheses
- H1 (concentration). The interaction between audience-authority style and low institutional trust is significantly larger than the interaction between aggression and low institutional trust. Pre-registered as a coefficient contrast, not a fishing expedition across six features.
- H2 (aggression null). Aggression without audience-authority routing has a null-to-negative effect on favorability in all strata. Ordinary attack politics is a solved literature; this replicates it as an internal validity check.
- H3 (cohort demotion). Birth cohort's interaction with style attenuates by ≥50% when institutional trust and media diet enter the model. Cohort is a covariate here, not a headline.
Registered kill conditions
State these before any data is collected. They are what makes a null publishable rather than embarrassing. - If H2 fails — aggression alone carries the effect — the mechanism story is dead. The finding becomes "adversarial affect, not audience routing," and the piece is written that way. - If H1's contrast is not significant but both indices are, the features are not separable at this sample size. Report as underpowered, do not reinterpret. - If H3 fails and cohort survives trust and media diet, report that. "It really is generational" is a legitimate outcome and should not be argued away.
1. Construct definition: the six features
Do not score "aggression" as a single dimension. Score six binary features per exchange, each span-grounded to the transcript, then compose two pre-registered indices.
Feature codebook
Each feature is coded 0/1 at the exchange level (one candidate's contiguous speaking turn responding to or targeting another candidate). Every 1 requires a supporting character span. F1 — Direct second-person address, target present. Speaker addresses the opponent as "you" while that opponent is on stage. Positive: "You said that on this stage four years ago." Negative: Third-person reference to an absent figure. Second-person addressed to the moderator or the audience (that is F5). Rule-detectable: partially. spaCy dependency parse for 2nd-person pronoun subjects with a person-entity antecedent in the prior 3 turns. LLM adjudicates target identity. F2 — Setup/reveal structure. The turn contains a deliberate delay between premise and payoff — the point is withheld for at least one clause boundary and then landed. Positive: A statistic offered flatly, then the reversal that makes it damning. Negative: Assertion followed by elaboration. The test is whether removing the final clause destroys the meaning of the earlier ones. Rule-detectable: no. LLM-only, with an explicit "would the setup be inert without the reveal?" decision rule. F3 — Characterological targeting. The attack is on who the opponent is — their consistency, courage, authenticity, class position — rather than on a policy position or record outcome. Positive: "You'll say anything in this room." Negative: "Your tax plan adds four trillion." Record-based attacks with a policy referent are F3=0 even when hostile. Boundary: Hypocrisy attacks are F3=1 when the point is the hypocrisy, F3=0 when the point is the policy. F4 — The flip. The speaker reuses the opponent's own framing, phrasing, or attack line and redirects it. Requires prior-turn context in the coding window. Positive: Opponent says "career politician"; speaker returns it against them within the same exchange or the next. Negative: Independent coincidental phrasing. Requires lexical or structural overlap with a span in a prior turn, cited by the coder. Rule-detectable: partially — n-gram overlap with prior 5 turns as a candidate generator, LLM confirms intent. F5 — Audience-facing performance. The speaker plays to the room rather than to the moderator or the camera-as-neutral-record. Explicit audience address, pausing for reaction, soliciting response. Positive: "Ask the people in this room whether that's true." Negative: Rhetorical questions with no audience referent. Note: Transcript-only coding will underdetect this. See §3.3 on the audio/video secondary channel. F6 — Frame-breaking. The speaker names the format, the rules, the media, or the artifice while inside it — declaring the arbiter illegitimate or the process rigged. Positive: "This question is exactly what I'd expect from this network." "We all know what this format is for." Negative: Ordinary complaints about time limits. The test is whether the arbiter's legitimacy is at stake, not merely their timekeeping.
Composed indices
- AUTH = mean(F5, F6) — audience-authority routing.
- AGGR = mean(F1, F2, F3) — adversarial performance without authority relocation.
- F4 (flip) is scored and modeled separately. It loads on both theoretically and is the most likely single-feature confound. Do not fold it into either index. If the whole effect turns out to be F4, the honest headline is "voters like counterpunching," which is old news and should be reported as such.
2. Track B is primary: the vignette experiment
Track A cannot identify this. Aggressive candidates are systematically younger, more online, more anti-institutional, and more often insurgents. Style and candidate type are collinear in every naturally occurring debate corpus. The experiment is the study; the corpus is the scaffolding. Build Track B first.
2.1 Design
2 × 2 between-subjects factorial, aggression × audience-authority. This is the design choice that isolates the mechanism — a four-level ordinal "intensity ladder" would not. | Condition | AGGR | AUTH | Description | |---|---|---|---| | C1 Baseline | low | low | Policy contrast, no personal targeting, no audience routing | | C2 Aggression-only | high | low | Personal/characterological attack, addressed to opponent and moderator, no crowd appeal, no frame-breaking | | C3 Authority-only | low | high | Mild policy content, but plays to the room and names the format as rigged | | C4 Full package | high | high | Both | Held constant across all four: policy content, candidate name, party cue, apparent age, race, gender, word count (±8%), reading level (±0.5 Flesch-Kincaid grade), and the opponent's preceding turn.
2.2 Partisan symmetry — non-negotiable
Every respondent sees the stimulus with a randomly assigned party cue (D / R / no-party-stated), crossed with condition. Eight-cell party-by-condition randomization. This exists because the finding is worthless — and worse, dishonest — if the design can only detect the effect for one side's style. Pre-register the party-symmetry test: if the style effect differs significantly by party cue, that difference is a reported finding, not a footnote. Any writeup that shows the effect for one party must show the other party's estimate at equal prominence. Two stimulus sets (a Democratic-primary-flavored issue frame and a Republican-primary-flavored one), each rendered in all four conditions, to avoid issue-content confounds with party cue.
2.3 Measures
Dependent variables 1. Candidate feeling thermometer, 0–100 (primary DV). 2. Vote intention vs. the opponent, forced choice. 3. "Who won this exchange?" — mechanism check. 4. Perceived authenticity, 5-item. 5. Perceived fairness of the format — manipulation check for AUTH. Moderators (measured, not assigned) - Institutional trust: 4–6 item battery covering media, elections, Congress, courts. Use an existing validated battery rather than writing one. - Media diet: short-form video hours/week; primary news source; podcast consumption; whether the respondent typically encounters debates as full broadcasts or as clips. That last item is the sleeper variable and should be worded carefully. - Birth year (continuous — do not collect cohort as a category and then treat it as ordinal). - Political interest, partisanship strength, education, gender, race. Manipulation checks: two items confirming perceived aggression and perceived audience-orientation, used to validate the stimuli in a pilot before full fielding.
2.4 Sample and power
Detecting a difference between two interaction slopes is expensive. Interaction effects generally require ~4× the N of the corresponding main effect, and H1 is a contrast between two of them. - Pilot: n = 200, purpose is manipulation-check validation only. No hypothesis testing. If the AGGR and AUTH manipulations do not separate cleanly on their checks, revise stimuli and re-pilot. - Main field: n = 2,000, quota-stratified on birth year (four strata, oversampling 1981–1996 and 1997–2005 to ~600 each), party ID, gender, and education. - Powered for a small-to-moderate interaction contrast (f² ≈ 0.02) at 80%. - Budget at roughly $1.50–2.00 per respondent for a 6–8 minute instrument: $3,000–4,000 for the main field, plus ~$400 for the pilot. If this is out of range, say so now — a 700-person version can test main effects honestly but cannot test H1, and running it while claiming otherwise is the exact overclaiming the Parrot audits.
2.5 Analysis plan
Primary model, OLS on the thermometer with robust SEs:
FT ~ AGGR * TRUST + AUTH * TRUST + FLIP
+ AGGR * AUTH
+ PARTY_CUE * AGGR + PARTY_CUE * AUTH
+ BIRTH_YEAR + covariates
H1 test: bootstrap the contrast β(AUTH × TRUST) − β(AGGR × TRUST), 10,000 resamples, report the CI on the difference. Pre-register this single contrast as the primary test. Everything else is secondary and labeled as such. H3 test: nested models. Fit cohort × style without trust and media diet, then with. Report the proportion of the cohort interaction coefficient that survives, with a bootstrap CI on the attenuation ratio. Do not use a Sobel test. Multiple comparisons: one primary test. Everything else is secondary — Benjamini-Hochberg at q = 0.05 across the following family, closed here and not amended after fielding:
- S1–S6: each single feature (F1…F6) × TRUST interaction on the thermometer.
- S7: FLIP main effect on the thermometer.
- S8: FLIP × TRUST interaction.
- S9: AUTH × PARTY_CUE (style effect split by D / R / no-party cue).
- S10: AGGR × PARTY_CUE.
- S11: AGGR × AUTH interaction.
- S12: AUTH × BIRTH_YEAR (continuous).
- S13: AUTH × clip-vs-broadcast media-diet item (the sleeper variable).
- S14: AUTH and AGGR main effects on the "who won this exchange?" item.
- S15: AUTH main effect on the authenticity index.
Fifteen secondary tests, named before any data exists. A test not on this list is exploratory, is labeled exploratory in the writeup, and cannot be promoted after the fact (§7, post-hoc feature promotion).
2.6 Stopping rule
- Pilot: field until 200 approved completes or 14 days, whichever comes first. The manipulation-check gate is evaluated once, after the pilot closes. No hypothesis tests on pilot data, ever.
- Main field: field until 2,000 approved completes or 30 days, whichever comes first. If quota strata are unfilled at day 30, close anyway, report the shortfall per stratum, and recompute achieved power before running any test — a shortfall is reported, never quietly absorbed.
- No interim analyses. The primary contrast is computed exactly once, on the closed dataset, after exclusions. Nobody — human or machine — opens an outcome column before the field closes; data-quality columns only until then.
2.7 Exclusion criteria
Applied before any outcome column is read, using quality variables only; counts reported per condition; a sensitivity analysis with excluded respondents included ships in the appendix.
- Failed either of the two attention checks.
- Completion time under 40% of the pilot median duration.
- Duplicate participant ID or IP address (first submission kept).
- Straightlining: zero variance across the thermometer plus all Likert batteries combined.
- Self-reported non-serious responding on the end-of-survey commitment item.
- Missing primary DV (thermometer).
An exclusion rule not on this list may not be invented after fielding; if one proves necessary (e.g., a platform-level fraud event), it is applied symmetrically across conditions and disclosed in the deviation log with the piece.
3. Track A: the observational corpus
Scaffolding and descriptive color. It cannot carry a causal claim and the writeup must not imply that it does.
3.1 Corpus
- Primary-debate transcripts, both parties, 2016–2026. Prefer official/network transcripts over auto-captions where available; log the source per document.
- Unit of analysis: the exchange, not the article or the full debate.
- Speaker metadata: party, incumbency, polling position at debate date, age, gender, race, insurgent/establishment coding (pre-register the coding rule).
- Debate metadata: network, moderator, stage size, primary vs. general, date.
3.2 The normalization trap, again
The 3.4× agentless-passive-by-category finding applies here in a new form: debate topic drives attack rates. Immigration and crime segments will run hot on F1/F3 regardless of speaker. Never rank speakers on raw feature rates. - Topic-code every exchange (fixed taxonomy, ~12 categories). - Report within-topic comparisons and direct-standardized rates. - Validate with a mixed-effects model: random intercepts for speaker, debate, and topic. - Calibration gate: before spending LLM tokens on the full corpus, confirm that topic explains a substantial share of raw feature-rate variance on a 300-exchange subsample. If it doesn't, the topic coder is broken — fix it before proceeding.
3.3 The clip-selection problem — read this before using engagement as a DV
Engagement-per-view on clipped debate moments is conditioned on the clip having been clipped. Clippers select for exactly the features being measured. This is not a nuisance parameter; it is a mechanism that can generate the entire predicted result out of nothing.
Mitigations, in order of preference:
1. Preferred: sample exchanges from the full transcript, then check which ones were clipped. Model clipping probability itself as an outcome (CLIPPED ~ AUTH + AGGR + FLIP + topic + speaker). This turns the confound into a finding: it measures whether the media ecosystem selects for audience-authority routing, which is a Parrot piece on its own.
2. Restrict engagement analysis to official full-debate uploads where segment-level retention is available.
3. If neither is available, drop the engagement DV entirely. Report feature prevalence descriptively and stop.
Platform age-skew is not a cohort measure. If it appears in the writeup at all, it appears with that sentence attached.
3.4 Dedup
Same-clip reuploads across accounts are the wire-copy problem in new clothes. Perceptual hash on video, plus transcript n-gram overlap ≥ 0.9 within a 14-day window. One canonical row per underlying exchange, engagement summed with the aggregation rule logged.
4. Reliability protocol
4. Reliability protocol
Matches the existing standard. No rankings, no rates, and no LLM-coded variable enters any model before this gate passes. - Two human coders independently code a 250-exchange stratified sample. - Krippendorff's α ≥ 0.70 per feature, computed separately for each of F1–F6. A pooled α is not acceptable — F2 and F4 are the hard ones and will hide behind F1's easy agreement. - Any feature below 0.70: revise the codebook decision rules, re-train, re-code a fresh sample. Two failed rounds on a feature means that feature is dropped from the indices and the drop is disclosed. - LLM coder is then validated against the human-consensus set. Report LLM-vs-human α per feature alongside human-vs-human. If the LLM underperforms the human floor, it does not get to code the corpus. - Every LLM code returns a character span. Codes without spans are discarded, not repaired.
5. Build order
Phase 0 — Pre-registration. Hypotheses, the single primary contrast, the enumerated secondary family, kill conditions, stopping rule, exclusion criteria. Timestamped and posted (OSF) before Phase 2 data collection. This is the whole credibility of the project; a study auditing overclaiming that pre-registers after peeking is not recoverable. Phase 1 — Codebook + reliability. §1 and §4. Human coding, α gate, LLM validation. Gate: all retained features ≥ 0.70. Phase 2 — Stimulus construction + pilot. Write eight stimuli (4 conditions × 2 issue frames), verify they differ on the intended dimensions and only on those, by human coding with the §1 instrument. Field the n=200 pilot. Gate: manipulation checks separate. Phase 3 — Main field. n = 2,000. Analysis per §2.5. Gate: none — the result is the result. Phase 4 — Corpus. Track A on the Spark. Feature-score the transcript corpus, topic-normalize, run the clipping-probability model. Descriptive only. Phase 5 — Writeup. Both tracks, experiment leading. Publish the pre-registration diff — every deviation from Phase 0, itemized.
6. Acceptance criteria
- Pre-registration is public and timestamped before Phase 2 fielding.
- All retained features clear α ≥ 0.70 for both human-human and LLM-human.
- Manipulation checks separate AGGR and AUTH in the pilot.
- Main field hits quota targets within 10% per cohort stratum.
- The primary contrast is reported with its CI regardless of direction, in the first three paragraphs of any writeup.
- Party-symmetry estimates are reported at equal prominence.
- Every observational claim is labeled observational in the sentence that makes it.
- Deviation log published with results.
7. What kills this study
- Stimulus artifact. If C4 is simply written better or funnier than C1, the study measures writing quality. The pilot's job is to catch this; a humor-rating item in the pilot is cheap insurance.
- Underpowered contrast. Running n=700 and reporting H1 anyway.
- Clip-selection laundering. Using viral-clip engagement as a proxy for public preference. §3.3 exists to prevent it.
- Cohort reification. Reporting "millennials" as an explanation after trust and media diet have eaten the coefficient.
- Asymmetric framing. Publishing the party split in only one direction.
- Post-hoc feature promotion. Discovering F4 carries everything and rewriting the theory around the flip. If that happens, report it as an exploratory finding requiring independent replication.
8. Publication paths by outcome
- Stimulus artifact. If C4 is simply written better or funnier than C1, the study measures writing quality. The pilot's job is to catch this; a humor-rating item in the pilot is cheap insurance.
- Underpowered contrast. Running n=700 and reporting H1 anyway.
- Clip-selection laundering. Using viral-clip engagement as a proxy for public preference. §3.3 exists to prevent it.
- Cohort reification. Reporting "millennials" as an explanation after trust and media diet have eaten the coefficient.
- Asymmetric framing. Publishing the party split in only one direction.
- Post-hoc feature promotion. Discovering F4 carries everything and rewriting the theory around the flip. If that happens, report it as an exploratory finding requiring independent replication.
8. Publication paths by outcome
| Outcome | Headline |
|---|---|
| H1 and H2 hold | The debate stage is being scored like a rap battle by low-trust voters — and generation was never the variable |
| H2 fails, aggression carries it | Voters like a fighter. The mechanism is simpler and worse than the elegant version |
| H1 null, both indices positive | Style matters, the components don't separate — a replication call with the instrument attached |
| H3 fails, cohort survives | The generational read holds up under controls that should have killed it |
| Everything null | The instrument, the codebook, and the reliability data published as a tool — plus the clipping-probability model from Phase 4, which is a piece on its own |
| Every row is publishable. That is the point of registering it this way. |
2 · The build, in the open — LIVING DOCUMENT: this section changes as the work moves, and is not covered by the hash above. Every deviation from the protocol lands in the log below, dated.
BATTLE RAP — builder handoff
Read SPEC.md first. The spec is the contract; this file is only the state of the build. Any builder session (Claude on the Mac, Claude on the Spark, DeepSeek scoping pass) works from this file and updates it when a gate is passed or a deviation is logged.
Ground rules carried over from the desk
- The experiment (Track B) is the study; the corpus (Track A) is scaffolding. Do not
invert the build order to chase the fun part.
- No LLM-coded variable enters any model before the §4 reliability gate passes
(α ≥ 0.70 per feature, human-human AND llm-human).
- Every LLM code carries a character span; span-less codes are discarded, not repaired.
- Heavy LLM passes on the Spark go through the governed job queue, under the GPU lease
rules. Local models never judge; they run mechanical passes only (n-gram overlap
candidates for F4, dedup hashing) — routing doctrine applies.
- Nothing observational gets written up with causal language. Acceptance criterion 7.
State (update in place)
- Phase 0 — pre-registration: WORDING COMPLETE (v0.2, 2026-08-09: stopping
rule §2.6, exclusion criteria §2.7, enumerated 15-test secondary family in
§2.5). Public at https://thestochasticparrot.com/research/battle-rap-prereg/
with sha. NOT YET FROZEN — the freeze is the OSF posting (Mike's account;
kit/OSF_REGISTRATION.md is ready to paste). After posting, put the DOI in
STATUS.json and the page header.
- Phase 1 — codebook + reliability: MACHINE SIDE COMPLETE (overnight
2026-08-09). In the repo:
- corpus/raw/ — 56 debate transcripts, 2015-08 → 2024-10, American
Presidency Project, per-document ledger in corpus/sources.jsonl.
- corpus/exchanges.jsonl — 6,685 exchanges (segmenter operationalization
documented in scripts/segment_exchanges.py: candidate turn ≥20 words,
prior-3-turn other-speaker requirement, prior-5-turn context).
- codebook/sample.jsonl — 250-exchange stratified sample (party × cycle ×
stage band, seed 20260809, strata table in sample_strata.json).
- codebook/form_coder_A.md / form_coder_B.md — same 250, independent orders.
- scripts/reliability.py — span validation + per-feature Krippendorff α +
the 0.70 gate report.
SEGMENTER v2 (later same night): the PARTICIPANTS header now outranks every
heuristic — v1's moderator guesser had DROPPED sitting candidates from five
general-election debates (e.g. Philadelphia 2024 lost one of Harris/Trump),
deleting their exchanges. Corpus re-cut 5,581 -> 5,834 exchanges, all
generals >= 2 candidates; SAMPLE REDRAWN (v2, same seed) before any coding
began, forms regenerated.
RESIDUAL LIMITATION: stage_size can overcount when a chatty moderator passes
the candidate heuristic (e.g. a 2-candidate debate reading 3); the banding
(2-4 / 5-8 / 9+) absorbs it. Fix before Phase 4 corpus scoring.
WAITING ON HUMANS: two coders (kit/CODER_RECRUITMENT.md). The α gate cannot
be machine-substituted — SPEC §4 requires human-human agreement first.
- Phase 2 — stimuli + pilot: BLOCKED (OSF freeze + codebook α gate +
~$425). Fielding-day config pre-written: kit/PROLIFIC_PILOT.md. Stimulus
text remains deliberately unwritten (stimuli/README.md).
- Phase 3 — main field: BLOCKED on money ($3,000–4,000) + Phase 2 gate.
- Phase 4 — corpus scoring on the Spark: BLOCKED on Phase 1 α gate;
calibration gate (topic-variance on 300 exchanges) before any full-corpus
token spend. Exchanges are ready for it the day the gate clears.
- Phase 5 — writeup: blocked on everything above.
Money gates (Mike-only)
- Pilot: ~$400 (n=200). Main field: $3,000–4,000 (n=2,000, quota-stratified).
- Spec §2.4 verbatim: a 700-person version cannot test H1 and must not claim to.
Deviation log
- Phase 0 — pre-registration: WORDING COMPLETE (v0.2, 2026-08-09: stopping rule §2.6, exclusion criteria §2.7, enumerated 15-test secondary family in §2.5). Public at https://thestochasticparrot.com/research/battle-rap-prereg/ with sha. NOT YET FROZEN — the freeze is the OSF posting (Mike's account; kit/OSF_REGISTRATION.md is ready to paste). After posting, put the DOI in STATUS.json and the page header.
- Phase 1 — codebook + reliability: MACHINE SIDE COMPLETE (overnight 2026-08-09). In the repo:
- corpus/raw/ — 56 debate transcripts, 2015-08 → 2024-10, American Presidency Project, per-document ledger in corpus/sources.jsonl.
- corpus/exchanges.jsonl — 6,685 exchanges (segmenter operationalization documented in scripts/segment_exchanges.py: candidate turn ≥20 words, prior-3-turn other-speaker requirement, prior-5-turn context).
- codebook/sample.jsonl — 250-exchange stratified sample (party × cycle × stage band, seed 20260809, strata table in sample_strata.json).
- codebook/form_coder_A.md / form_coder_B.md — same 250, independent orders.
- scripts/reliability.py — span validation + per-feature Krippendorff α + the 0.70 gate report. SEGMENTER v2 (later same night): the PARTICIPANTS header now outranks every heuristic — v1's moderator guesser had DROPPED sitting candidates from five general-election debates (e.g. Philadelphia 2024 lost one of Harris/Trump), deleting their exchanges. Corpus re-cut 5,581 -> 5,834 exchanges, all generals >= 2 candidates; SAMPLE REDRAWN (v2, same seed) before any coding began, forms regenerated. RESIDUAL LIMITATION: stage_size can overcount when a chatty moderator passes the candidate heuristic (e.g. a 2-candidate debate reading 3); the banding (2-4 / 5-8 / 9+) absorbs it. Fix before Phase 4 corpus scoring. WAITING ON HUMANS: two coders (kit/CODER_RECRUITMENT.md). The α gate cannot be machine-substituted — SPEC §4 requires human-human agreement first.
- Phase 2 — stimuli + pilot: BLOCKED (OSF freeze + codebook α gate + ~$425). Fielding-day config pre-written: kit/PROLIFIC_PILOT.md. Stimulus text remains deliberately unwritten (stimuli/README.md).
- Phase 3 — main field: BLOCKED on money ($3,000–4,000) + Phase 2 gate.
- Phase 4 — corpus scoring on the Spark: BLOCKED on Phase 1 α gate; calibration gate (topic-variance on 300 exchanges) before any full-corpus token spend. Exchanges are ready for it the day the gate clears.
- Phase 5 — writeup: blocked on everything above.
Money gates (Mike-only)
- Pilot: ~$400 (n=200). Main field: $3,000–4,000 (n=2,000, quota-stratified).
- Spec §2.4 verbatim: a 700-person version cannot test H1 and must not claim to.
Deviation log
(append-only; every departure from SPEC.md v0.1, dated) - none yet
3 · The coder instrument — living document; frozen only when it clears its reliability gate.
Coding instrument — six features, exchange level
The construct definitions live in ../SPEC.md §1 and are NOT restated here (one definition, one place). This file adds only what a coder needs operationally.
Unit
One exchange = one candidate contiguous speaking turn responding to or targeting another candidate. Moderator-only turns are not units. An exchange keeps its prior 5 turns as coding context (F4 needs them).
Code record schema (one JSON object per exchange per coder)
{
"exchange_id": "...", "coder": "...", "ts": "...",
"F1": 0, "F2": 0, "F3": 0, "F4": 0, "F5": 0, "F6": 0,
"spans": {"F3": {"start": 0, "end": 0, "quote": "..."}},
"notes": ""
}
{
"exchange_id": "...", "coder": "...", "ts": "...",
"F1": 0, "F2": 0, "F3": 0, "F4": 0, "F5": 0, "F6": 0,
"spans": {"F3": {"start": 0, "end": 0, "quote": "..."}},
"notes": ""
}
RULES: every feature scored 1 MUST have a span entry whose quote matches the transcript verbatim at [start,end). A 1 without a span is invalid — the harness rejects the record, it does not repair it (SPEC §4).
Decision rules coders trip on (from the spec, restated as tests)
- F2: delete the final clause. If the earlier clauses become inert, F2=1.
- F3 vs policy attack: is the POINT the person or the policy? Hypocrisy → person.
- F4: cite the prior-turn span being flipped, or code 0.
- F5 vs F1: "you" to the opponent is F1; playing to the room is F5; both can be 1.
- F6: is the ARBITER's legitimacy at stake, or just the clock? Clock complaints are 0.
4 · The stimuli blocker — the file whose job is to stop us writing test material too early.
Stimuli — DO NOT DRAFT YET
Eight stimuli (C1–C4 × two issue frames) per SPEC §2.1–2.2. Blocked until the codebook clears its α gate (see HANDOFF.md — the same instrument validates the stimuli; freezing it first prevents teaching to the test).
Hard constraints across all four conditions of a frame: policy content, candidate name, party cue slot, apparent age/race/gender, word count ±8%, Flesch-Kincaid grade ±0.5, identical opponent preceding turn. Party cue is a render-time variable (D / R / none), never baked into the text. Pilot includes a humor-rating item (SPEC §7, stimulus-artifact kill).
5 · The operator kit — the steps only humans can take, pre-assembled so each is one action from done.
Faster than email: the desk keeps a Slack room (#the-desk) — ask for an invite and say you want to code.
Second coder — recruitment one-pager
The reliability protocol (SPEC §4) requires TWO human coders working independently. Coder A can be Mike; Coder B is the open seat.
The job: code 250 debate exchanges (form_coder_B.md) on six binary features per the instrument (codebook/CODEBOOK.md). Every "1" requires a verbatim supporting quote. Realistic effort: 8–12 hours; can be split across a week. The two coders must not discuss items until both files are submitted.
Who fits: anyone careful with text — a grad student, a journalist friend, a retired lawyer. Political knowledge helps; strong priors about the people quoted do not (the instrument codes form, not politics).
Options: (a) someone Mike knows, paid or favor; (b) Prolific's vetted "expert" pool (post as a transcription/annotation task, ~$12–15/hr, budget ~$150); (c) both coders external and Mike codes nothing (cleanest optics — the desk's operator not coding his own study).
When both code files exist: run
python3 scripts/reliability.py codes_A.jsonl codes_B.jsonl
— the α gate report is automatic. Features under 0.70 trigger the SPEC §4
revise-retrain-recode loop.
OSF Registration — ready to paste
This file is the freeze, pre-assembled. Mike: create a registration at osf.io (template: "OSF Preregistration"), paste the sections below into the matching fields, and submit. The moment it has a DOI, put the DOI in STATUS.json and the study page header; Phase 2 unblocks.
Title: Audience-as-Judge: Testing Whether Adversarial Debate Style Routes Authority Away From Institutions
Authors: The Stochastic Parrot desk (automated), M. Portney (operator).
Description: We test whether adversarial debate performance affects candidate favorability through audience-authority routing — features that relocate the verdict from institutional arbiters to the audience — rather than through aggression per se, with institutional trust as moderator and birth cohort deliberately demoted to a covariate. Full protocol, version-controlled with hash: https://thestochasticparrot.com/research/battle-rap-prereg/
Hypotheses: SPEC.md §0 verbatim — H1 (concentration: β(AUTH×TRUST) > β(AGGR×TRUST), pre-registered as a single coefficient contrast), H2 (aggression null-to-negative in all strata), H3 (cohort × style attenuates ≥50% when trust and media diet enter). Registered kill conditions: if H2 fails, the mechanism story is dead and is reported that way; if H1's contrast is non-significant while both indices are positive, report as underpowered without reinterpretation; if H3 fails, "it really is generational" is reported as the finding.
Design: 2×2 between-subjects factorial (aggression × audience-authority), crossed with randomized party cue (D/R/none) — eight cells — two issue-frame stimulus sets. Held constant across conditions: policy content, candidate identity cues, word count (±8%), Flesch-Kincaid grade (±0.5), opponent's preceding turn. (SPEC §2.1–2.2.)
Sampling plan: Prolific. Pilot n=200 (manipulation-check validation only, no hypothesis tests). Main field n=2,000, quota-stratified on birth year (four strata, 1981–1996 and 1997–2005 oversampled to ~600 each), party ID, gender, education. Stopping rule per SPEC §2.6: 2,000 approved completes or 30 days; no interim analyses; shortfalls reported per stratum.
Variables: DVs — 0–100 feeling thermometer (primary), forced-choice vote intention, "who won this exchange," 5-item perceived authenticity, perceived format fairness (AUTH manipulation check). Moderators — validated institutional-trust battery; media diet including clip-vs-broadcast item; birth year (continuous); political interest, partisanship strength, education, gender, race. (SPEC §2.3.)
Analysis plan: OLS on the thermometer with robust SEs: FT ~ AGGRTRUST + AUTHTRUST + FLIP + AGGRAUTH + PARTY_CUEAGGR + PARTY_CUE*AUTH + BIRTH_YEAR + covariates. Primary test: bootstrap contrast β(AUTH×TRUST) − β(AGGR×TRUST), 10,000 resamples, CI on the difference. H3 via nested models with bootstrap CI on the attenuation ratio (no Sobel). Secondary family: the 15 tests enumerated in SPEC §2.5, Benjamini-Hochberg q=0.05, family closed at registration. Exclusions per SPEC §2.7 (attention checks, speed, duplicates, straightlining, commitment item, missing DV), applied before outcome columns are read.
Other: Track A (observational debate-transcript corpus) is scaffolding and descriptive only; it cannot carry causal claims (SPEC §3). Reliability gate: Krippendorff's α ≥ 0.70 per feature, human-human and LLM-human, before any coded variable enters any model (SPEC §4). Every deviation from this registration will be published as an itemized diff with the results.
Prolific pilot — config outline (n=200, ~$400)
BLOCKED until: OSF freeze posted + stimuli written and validated (stimuli are themselves blocked until the codebook clears its α gate). This file exists so the fielding day is configuration, not design.
- Sample: n=200, US adults, balanced sex, no quota strata (pilot validates manipulations, not effects). Exclude prior participation in our studies.
- Instrument length: 6–8 minutes → pay $1.60 (≥$12/hr equivalent). Budget: 200 × $1.60 × 1.33 (Prolific fee) ≈ $425.
- Randomization: 8 cells (2×2 condition × D/R/none party cue) × 2 issue frames, balanced assignment.
- Measures: manipulation checks (perceived aggression, perceived audience-orientation — the pilot's entire point: SPEC gate is that AGGR and AUTH manipulations separate cleanly), humor rating (stimulus-artifact check, SPEC §7), thermometer + full battery for timing calibration only.
- Attention checks (2): one instructed-response item mid-battery; one content check about the stimulus ("which issue did the candidates discuss?").
- Exclusions: per SPEC §2.7, using pilot's own median duration.
- No hypothesis tests on pilot data. The gate is manipulation-check separation; if it fails, revise stimuli and re-pilot (SPEC §5 Phase 2).