An integrity referee for public debates. We score how each side argued, never who won.
Scoring integrity, not victory. Persuasion is never scored. A debater can lose the crowd and post the highest figure in the library; a debater can win the crowd and post one of the lowest.
Conduct, not correctness. We never score who is factually right. A person can be wrong in good faith and score high. Fact checking is someone else's job; honest engagement is what we measure.
The whole debate is the record. A point answered thirty minutes later counts as answered, in full. Tangents are not dodges. No adverse verdict is issued until the entire debate has been searched for a response.
Total auditability. Score → components → verdicts → events → quotes. Every layer is one click deeper, and the bottom layer is always the debaters' own words.
A DIS is a weighted blend of four rates: rates, not raw counts, so a 30-second exchange and a 3-hour debate are scored on the same scale:
Confronted the point (45%). Of the challenges aimed at you, weighted by how central they were: how many did you actually confront?
Argued clean (30%). Logical fallacies and record denial per 1,000 words spoken, weighted by how central the point they touched was. The score decays as fouls pile up but never reaches zero: zero would claim a performance was maximallydishonest, and that isn't something we can measure.
Dealt straight (15%).The moments where the honest move cost something and the debater made it anyway. Conceding a point, naming common ground, owning an error, but equally: stating the opponent's case at its strongest before attacking it, answering the hard question head-on instead of dodging it, returning to something left hanging, and saying "I don't know". Half of these require agreeing with nobody. A debater who thinks their opponent is comprehensively wrong can still score well here, by taking them seriously enough to state their case fairly and then answer it.
Let them talk (10%). Interruptions, per 1,000 words spoken by everyone else while this debater was in the room. You can only cut in while someone else holds the floor, so that is the denominator. Moderators are tracked but never scored.
Every tracked point gets a crux: the specific proposition the point demands be confronted: not the topic, the proposition. Cruxes must be position-neutral: each one offers a genuine branch to rebut or defend, never "concede X" as the only satisfying move. A partisan of either side should read every crux as a fair question.
The iron rule: a "confronted" verdict must quote the exact sentence that confronts the crux. If no such sentence can be pasted, the verdict is "talked around it": engagement theater, scored barely above outright evasion. Saying "that's a good question and we need to address it" and then not addressing it earns exactly what it deserves.
| Confronted | 1.0 | Engaged the actual proposition: agreed, refuted with argument, or gave a reasoned refusal the crux allowed |
| Conceded | 1.0 | Gave the point (also earns a credit) |
| Partially confronted | 0.5 | Real engagement with part of the crux |
| Talked around it | 0.15 | Responded to the topic, gave context, answered a weaker version, never the proposition |
| Evaded | 0 | Changed the subject or refused without reasons |
| Mutually dropped / out of time / moderator cut / not addressed / covered by ally | n/a | No fault; excluded from the denominator |
We retired the Engagement Standard in August 2026. Until then every card carried a mark ("met the standard" at 65, "with distinction" at 85), and we no longer print any of it. Two reasons, and the second is why it had to go rather than merely move.
One: a threshold is a chosen cut wearing the grammar of a certificate. "Met the standard" is a claim about a person's conduct, and defending it means defending the number 65 as morally meaningful. Nothing does. It was our cut point, picked by us, and printing it beside someone's face made it look like a fact about them rather than a decision by us.
Two, and decisively: it could not survive our own error bars. We publish a re-run band of ±5–8 points. Measured against the corpus at the time, 48% of every performance we had scored sat within ±8 of the 85 line and 31% within ±8 of 65, meaning for roughly half the library, whether someone got a medal was decided by measurement noise rather than by how they argued. And 86% cleared 65 anyway, which makes it a participation floor, not a standard. A bar that most people clear and that noise can flip is not measuring what it claims to.
What replaced it is an anchor, not a bar. A raw number needs something to sit against or the reader anchors on the only other number in view, the person beside them, which is the league table we refuse to print. So instead of asking whether a score cleared our bar, we publish where it sits among everything we have measured:
| Lowest scored | 41 | The full range is published; nothing is withheld |
| Lower quartile | 70 | A quarter of performances sit below this |
| Median | 76 | Half above, half below |
| Upper quartile | 87 | A quarter sit above this |
| Highest scored | 94 | Across 29 published performances, rubric v0.11 |
A quartile is a fact about a distribution. A grade is a claim about a person. We will print the first and not the second.
Four of those 22 performances are under correction. In August 2026 we found that our scoring pipeline had, on two debates, accepted events from the judge with their rubric values missing, so conduct we had recorded, and printed, scored nothing. Both debates carry a correction note saying so on their own cards, naming which components moved and in which direction.
They will be re-scored, and this row will not change. Two separate commitments, and both are load-bearing. A debate re-run under a later rubric cannot be set beside the ones published with it (that is the whole reason these rows are pinned per version), so the repaired debates will seed a new row rather than edit this one. And we are not quietly correcting the numbers inside a row that published pages already carry: pinning a row and then editing it defeats the reason we pin. A published page must never shift meaning under a reader, and that has to hold when the news is bad or it was never a commitment. So this row stands, four of its 22 values are known to be wrong, and this paragraph is where we say so.
The honest cost of the change. The old bar was absolute (65 meant the same thing forever), and a distribution shifts as the library grows. That is a real loss and we are not going to pretend otherwise. We think it is worth paying because the bar was only pretending to be absolute: 65 was as arbitrary as any percentile, it just hid it better. The distribution is pinned per rubric version, exactly the way the thresholds were, so a published page never shifts meaning under a reader.
There is still no failing mark, and there never will be. We do not print a negative judgement next to a named person. A score is a measurement with error bars; stamping a verdict on a human being asserts a precision this instrument does not have. Where the sample behind a figure is thin, the card says so: every component prints the denominator it was computed from, and an asterisk on the total means look closer.
Career figures are weighted by challenges faced, not by words spoken. Word count is the denominator of the "argued clean" component only: weighting a career average by it would let whoever talked most buy a larger say in their own average. Challenges faced is the denominator of the 45% component: how hard the performance was actually tested.
Base values below are multiplied by the importance of the point they touched (×3 central, ×2 supporting, ×1 peripheral; jokes and asides are never penalized). Citing evidence within an argument is not an appeal to authority: the fallacy is authority offered instead of an argument. Defending or contextualizing your own record is not record denial, only misrepresenting what the record shows.
| Strawman | 2.5 | Rebutting a position the opponent does not hold. The tell is a rebuttal that would not land against what was actually said. |
| Quote or context stripping | 2.5 | Using the opponent's own words to mean something they did not mean, by removing the qualification or context that fixed their sense. |
| Record denial | 3.0 | Denying one's own documented words or actions when confronted with the record. Requires both receipts: the record and the denial. Defending or contextualising a record is not denial. |
| Whataboutism | 2.0 | Answering a charge by raising a different charge, so the original is never addressed ("how much did your department overspend?" met with "why has nobody asked what the other departments spend?"). Comparison used as argument is fine; comparison used as escape is this. |
| Ad hominem | 2.0 | Attacking the person, their motives or their manner instead of the point. Includes contesting an opponent's identity or sincerity in place of their argument ("you're not a real engineer if you think that"), and mockery or ventriloquism of a position ("so bridges hold themselves up by magic"). |
| Appeal to authority | 2.5 | "The experts say so" or "the law says so" offered instead of an argument. Citing evidence or expertise within an argument is not this rule. |
| False equivalence | 2.5 | Treating two things as equivalent when the feature that matters differs. Includes motte-and-bailey: defending a modest claim, then resuming the bold one. |
| False dilemma | 2.0 | Framing two options as the only options when others exist. |
| Unfalsifiable dismissal | 2.0 | Rejecting evidence by impugning its source in a way no evidence could survive. Arguing a source is unreliable with reasons is not this rule. |
| Manufactured novelty | 2.0 | Treating a mainstream, well-documented position within the opponent's own tradition as bizarre or self-refuting, to avoid engaging it. Requires showing the position is in fact standard. Genuine unfamiliarity is not this rule. |
| Rejecting the shared standard | 2.5 | Declaring mid-argument that the evaluative framework both sides were using no longer applies, without defending the exit (both sides argue from the inspection report until it cuts one way, then "reports can't capture what really goes on in a building"). Argued-for disagreement about what the standard should be is not this rule. |
| Unfalsifiable sourcing | 2.0 | Offering evidence that cannot in principle be checked (anonymous, undisclosed, "trust me") as satisfying an evidentiary burden, while declining every verification path offered. Anonymous sourcing itself is never the foul: disclosing to a body that can verify, or owning the limit (conceding the claim rides on trust), defeats the charge. |
| Resolution capture | 4.0 | Unilaterally redefining a load-bearing term of the agreed resolution so one's own burden becomes subjective, private or otherwise unfalsifiable, then leaning on the redefinition when the original claim is attacked. Announcing the redefinition up front does not immunise it: stipulation the opponent never accepted is capture with a courtesy note. |
| Asymmetric rigor | 2.5 | Holding the opponent's case to an evidentiary bar one's own case on the same question does not meet and is not offered to meet. Includes atomisation: disassembling a case that was in fact cumulative and requiring each fragment to clear the bar alone, while defending one's own side by totality-of-evidence inference. Includes procedural asymmetry: invoking the format's rules against the opponent while exempting oneself. |
| Gotcha quiz | 2.0 | Demanding a recall performance from the opponent (a name, a list, a number recited from memory) as a test of the person rather than a question in an argument, and treating hesitation, error or assistance as evidence against their case. Ad hominem machinery in the grammatical form of a question. |
| Unfalsifiable accusation | 2.5 | Accusing the opponent of in-debate misconduct (being fed answers, cheating, being coached), and when the accusation is checked and fails, expanding it to a form no check could clear rather than retracting. A retraction costs nothing and can earn C-4.2. An accusation supported by receipts is cross-examination, not this rule. |
| Withheld exhibit | 2.5 | Deploying evidence for its rhetorical force while denying the opponent examination of it: reading an exhibit into the record, then refusing to display it when asked. Evidence a debate cannot examine is testimony wearing an exhibit's authority. |
| Pre-discounting | 2.0 | Installing a discrediting lens before the opponent has argued, then re-applying it after they argue as if the intervening answers had not happened. Well poisoning run as a frame: each half reads as ordinary rhetoric, and the foul is the pair with the middle skipped. |
| Non-answer filibuster | 2.0 | Talking at length under a direct challenge while answering nothing in it. |
| Dodge or subject change | 2.5 | Changing the subject while a direct challenge is live. |
| Refused a direct question | 3.0 | Declining to answer at all, with no reason given that the question allows. |
| Non-mapping analogy | 2.0 | An analogy offered in place of engagement whose structure does not correspond to the argument, so it does no work on the proposition. The judge must be able to state the mismatch. A working analogy earns nothing adverse. |
| Analogy substitution | 1.5 | An analogy is explicitly challenged on its mapping and the speaker neither defends the correspondence nor concedes it: a different analogy arrives, or the original claim is merely restated. Each analogy may be individually sound; the foul is refusing to defend the correspondence once in question. |
| Asked and answered | 1.5 | Re-demanding an answer already responsively given: re-asking the same question without engaging the answer's content, or refusing every phrasing of the answer but the asker's own. It manufactures the appearance of evasion where none occurred. |
| Interruption | 0.5 flat | Cutting in during another speaker's answer to rebut, score a point, or take the floor. |
| Talking over | 1.0 flat | Continuing to speak through another speaker's turn rather than a single cut-in. |
| Procedural cut-off | 0.0 | Interrupting to run the exchange: calling time, moving on, handing over the floor, breaking up a crossfire. Recorded, never scored against anyone. Applies to debaters too: "let him finish" is not a foul. Where a cut-off is both procedural and argumentative, the argumentative rule is charged. |
Good faith earns points. Credits carry the same importance multipliers as deductions, and the scorecard shows them with the same prominence: a debater can carry a high fallacy rate and a high credit rate in the same debate, and both facts appear on the card.
| Clean concession | 3.0 | Granting a point without taking it back in the same breath. |
| Partial concession | 1.5 | Giving real ground while holding the rest. |
| Common ground | 1.5 | Naming a genuine agreement and building on it rather than passing over it. |
| Direct answer to a hard question | 2.0 | Taking the difficult question head-on, at its strongest form, without preamble or escape. An answer that arrives only after the questioner protests a non-answer is not head-on: the escape is on the tape, and is charged; the eventual answer may satisfy the verdict but earns no credit. |
| Returned to an open point | 2.0 | Coming back unprompted to answer something left open earlier. |
| Acknowledged uncertainty | 1.0 | "I don't know" where a confident answer was available and cheaper. |
| Owned an error | 3.0 | Admitting being wrong, including about a past position. |
| Steelman | 3.0 | Stating the opponent's case at its strongest, including where it tells against one's own side. |
1 · Transcription & attribution. The debate becomes a timestamped transcript with every utterance attributed to a speaker. Uncertain attributions are flagged, and anything built on them is excluded from the score.
2 · The point ledger.An AI judge extracts every claim, challenge, question, and thought experiment that shaped the debate, writes each one's crux, and weights its importance. Points raised by moderators or played clips count as challenges, but their authors are never scored.
3 · Event detection. Judges walk the transcript hunting deductions andcredits under a strict discipline: when in doubt between two rules, pick the cheaper one; when in doubt whether an event exists, it doesn't.
4 · Whole-debate reconciliation. Only after the entire debate is read does any point receive a verdict, so the answer given thirty minutes later, unprompted, counts in full.
5 · The hostile second judge.An independent adversarial pass attacks every "confronted" verdict, checking each confronting quote in context. Its only power is to demote: it can never promote. Verdicts that survive have been through two independent judges, one of whom was trying to kill them.
6 · Rollup.The four component rates are computed and blended into each debater's DIS, alongside the evidence badge (challenges faced, words spoken) that tells you how much data the score rests on.
An integrity referee will be accused of bias by whoever its scores embarrass. We designed for that day:
The rubric is published and versioned.Disputes are about whether a rule was correctly applied to a quote, never about an AI's opinion. Scorecards name the rubric version that scored them.
Categories are locked. The judging pipeline is technically barred from inventing rule categories: every event must cite one of the rules on this page.
Uncertainty is excluded, not hidden. Low-confidence findings appear on the scorecard but never count toward the score.
Scores are honest about their precision. Independent re-runs of the same debate agree within roughly ±5–8 points. A DIS is a measurement with error bars, not a verdict from on high.
We publish our own failures.During calibration, our checks caught the pipeline writing biased crux questions and inventing rule categories. Both were fixed structurally, with neutrality requirements and schema-enforced categories, and both incidents are part of this public record, because a referee that hides its corrections doesn't deserve your trust.
What we will never do: score who is right, score motives or character, or accept a request to tilt a scorecard. Every complaint about a score should arrive as: this quote, under this rule, was judged wrongly. Those we want to hear.
Every finding on a scorecard cites one of these. A rule may only be charged when the conduct matches its definition, and where two could apply the lower value is charged. Debates judged before v0.9 show the description without a citation: the rubric then defined six of these twenty-seven, so the id beside an older finding cannot be stood behind.
Rebutting a position the opponent does not hold. The tell is a rebuttal that would not land against what was actually said.
Using the opponent's own words to mean something they did not mean, by removing the qualification or context that fixed their sense.
Denying one's own documented words or actions when confronted with the record. Requires both receipts: the record and the denial. Defending or contextualising a record is not denial.
Answering a charge by raising a different charge, so the original is never addressed ("how much did your department overspend?" met with "why has nobody asked what the other departments spend?"). Comparison used as argument is fine; comparison used as escape is this.
Attacking the person, their motives or their manner instead of the point. Includes contesting an opponent's identity or sincerity in place of their argument ("you're not a real engineer if you think that"), and mockery or ventriloquism of a position ("so bridges hold themselves up by magic").
"The experts say so" or "the law says so" offered instead of an argument. Citing evidence or expertise within an argument is not this rule.
Treating two things as equivalent when the feature that matters differs. Includes motte-and-bailey: defending a modest claim, then resuming the bold one.
Framing two options as the only options when others exist.
Rejecting evidence by impugning its source in a way no evidence could survive. Arguing a source is unreliable with reasons is not this rule.
Treating a mainstream, well-documented position within the opponent's own tradition as bizarre or self-refuting, to avoid engaging it. Requires showing the position is in fact standard. Genuine unfamiliarity is not this rule.
Declaring mid-argument that the evaluative framework both sides were using no longer applies, without defending the exit (both sides argue from the inspection report until it cuts one way, then "reports can't capture what really goes on in a building"). Argued-for disagreement about what the standard should be is not this rule.
Offering evidence that cannot in principle be checked (anonymous, undisclosed, "trust me") as satisfying an evidentiary burden, while declining every verification path offered. Anonymous sourcing itself is never the foul: disclosing to a body that can verify, or owning the limit (conceding the claim rides on trust), defeats the charge.
Unilaterally redefining a load-bearing term of the agreed resolution so one's own burden becomes subjective, private or otherwise unfalsifiable, then leaning on the redefinition when the original claim is attacked. Announcing the redefinition up front does not immunise it: stipulation the opponent never accepted is capture with a courtesy note.
Holding the opponent's case to an evidentiary bar one's own case on the same question does not meet and is not offered to meet. Includes atomisation: disassembling a case that was in fact cumulative and requiring each fragment to clear the bar alone, while defending one's own side by totality-of-evidence inference. Includes procedural asymmetry: invoking the format's rules against the opponent while exempting oneself.
Demanding a recall performance from the opponent (a name, a list, a number recited from memory) as a test of the person rather than a question in an argument, and treating hesitation, error or assistance as evidence against their case. Ad hominem machinery in the grammatical form of a question.
Accusing the opponent of in-debate misconduct (being fed answers, cheating, being coached), and when the accusation is checked and fails, expanding it to a form no check could clear rather than retracting. A retraction costs nothing and can earn C-4.2. An accusation supported by receipts is cross-examination, not this rule.
Deploying evidence for its rhetorical force while denying the opponent examination of it: reading an exhibit into the record, then refusing to display it when asked. Evidence a debate cannot examine is testimony wearing an exhibit's authority.
Installing a discrediting lens before the opponent has argued, then re-applying it after they argue as if the intervening answers had not happened. Well poisoning run as a frame: each half reads as ordinary rhetoric, and the foul is the pair with the middle skipped.
Talking at length under a direct challenge while answering nothing in it.
Changing the subject while a direct challenge is live.
Declining to answer at all, with no reason given that the question allows.
An analogy offered in place of engagement whose structure does not correspond to the argument, so it does no work on the proposition. The judge must be able to state the mismatch. A working analogy earns nothing adverse.
An analogy is explicitly challenged on its mapping and the speaker neither defends the correspondence nor concedes it: a different analogy arrives, or the original claim is merely restated. Each analogy may be individually sound; the foul is refusing to defend the correspondence once in question.
Re-demanding an answer already responsively given: re-asking the same question without engaging the answer's content, or refusing every phrasing of the answer but the asker's own. It manufactures the appearance of evasion where none occurred.
Cutting in during another speaker's answer to rebut, score a point, or take the floor.
Continuing to speak through another speaker's turn rather than a single cut-in.
Interrupting to run the exchange: calling time, moving on, handing over the floor, breaking up a crossfire. Recorded, never scored against anyone. Applies to debaters too: "let him finish" is not a foul. Where a cut-off is both procedural and argumentative, the argumentative rule is charged.
Granting a point without taking it back in the same breath.
Giving real ground while holding the rest.
Naming a genuine agreement and building on it rather than passing over it.
Taking the difficult question head-on, at its strongest form, without preamble or escape. An answer that arrives only after the questioner protests a non-answer is not head-on: the escape is on the tape, and is charged; the eventual answer may satisfy the verdict but earns no credit.
Coming back unprompted to answer something left open earlier.
"I don't know" where a confident answer was available and cheaper.
Admitting being wrong, including about a past position.
Stating the opponent's case at its strongest, including where it tells against one's own side.
v0.11 (current). The standards layer. Everything before this priced what a debater said about the topic. Nothing priced what they did to the argument itself. Eight rules now do: unilaterally redefining a load-bearing term of the agreed resolution so your own burden becomes unfalsifiable, holding an opponent to an evidentiary bar your own case does not meet, demanding a recall performance as a test of the person rather than a question in an argument, accusing an opponent of in-debate misconduct and widening the accusation when the check fails instead of retracting it, reading an exhibit into the record and refusing to show it, installing a discrediting frame before an opponent argues and re-applying it afterwards as if the answers had not happened, re-demanding an answer already given, and offering evidence that cannot in principle be checked while declining every path that would check it. Three of the eight are recorded but not scored while we gather evidence that they fire the same way on debates we have not read. The scorecard shows them, and they cost nothing.
Three further changes came with it. Where a debate had a publicly agreed resolution, it is now recorded before judging, in the debaters' own words, with the utterance that shows each of them taking their burden up. A redefinition is then measured against what was agreed rather than against whichever debater stated the topic first. A participation gate withholds a score from anyone who faced fewer than four challenges or spoke fewer than 1,000 words: a rate computed over almost nothing was a number with no measurement behind it. Their record is still published, without a figure. And the two smaller components got the denominators they were always missing. Dealt straight is now credits per point the debater actually engaged, so the speech that can earn a credit is the speech it is divided by. Let them talk is interruptions per 1,000 words spoken by everyone else, because you can only cut in while someone else holds the floor: being interrupted no longer improves your own conduct figure, and a panel is not penalised for having more people in the room.
A fault we found while doing it.A review pass sits at the end of the pipeline and re-reads every positive verdict against the quote that is supposed to support it, demoting the ones the receipt does not carry. It recorded its decision in a field that accepted any text, and the code applied only the exact word it expected. Across the library that pass wrote 334 rulings, 267 of them a paragraph of reasoning in the field meant for the decision. Sixteen verdicts were marked for demotion and four were applied. The other twelve stayed at the higher verdict, across thirteen of the fourteen debates, lifting those debaters' confrontation figures by a small amount each. It had never worked, from the first debate we published. The field now accepts two words and nothing else, the code refuses anything it does not recognise instead of ignoring it, and every figure on the site was re-judged with the fixed pass. We are saying this out loud because the alternative is asking you to trust that we find our own faults.
v0.10.The third component was called “Gave credit”, and the name was doing damage. It told readers the component measured conceding, but four of its eight rules never involve agreeing with anyone: steelmanning, answering a hard question directly, returning to an open point, and admitting uncertainty. So a low score read as “refused to give an inch” when it may have meant something else entirely, and the component looked unanswerable: what is a debater supposed to do if there is genuinely nothing to concede? There was always an answer, and the label hid it. It is now Dealt straight, because what the eight rules share is not generosity: it is that in each one the honest move cost something strategically and the debater made it anyway. A label promising an opponent-independent route also has to price that route properly, so steelmanning rose from 2.0 to 3.0 and a direct answer to a hard question from 1.5 to 2.0. Both 3.0-point credits had previously required conceding, while steelmanning, the hardest thing anyone does in a debate, sat below them.
v0.9.The rule register. Every score here carried a rule id (R-2.1, C-3.1), and the rubric the judge reads defined six of the twenty-seven. The base table had been written once, early, and never carried forward into the working document. The judge was therefore choosing an id from the permitted list and writing an accurate description beside it, with nothing tying the two together. Audited against published debates the descriptions hold up and the ids frequently do not: the id meant for “false dilemma” was used for whataboutism, the one meant for “ad hominem” for strawman. No score is affected: every finding carries its own value and quotes, and those were assigned from the conduct, not the label, but the label was decorative and we were printing it as though it meant something. All twenty-seven rules are now published above, fed to the judge, and binding: a rule may only be charged when the conduct matches its definition. Findings from debates judged before v0.9 show their description without a citation, because a citation leading to a definition the finding does not match is worse than no citation. Nothing was rescored.
v0.8.“Gave credit” was reading the format instead of the debater. With enough debates to measure, panels averaged 0.367 on that component against 0.714 for head-to-head debates, a gap that tracked how much airtime each point got (r = 0.62), not how generous anyone was. Credits are earned in the back-and-forth around a point: in a one-on-one each point gets an extended exchange and many chances to concede, find common ground or steelman, while on a panel your point gets one exchange before the floor moves on. Measuring credit per challenge faced counted the openings a debater was given rather than the ones they had. Credits moved to per speaking turn(every time you take the floor is a chance to be generous), which closed the gap to 0.009. v0.11 moved the denominator again, to the weighted points a debater actually engaged. The constant was recalibrated so the library average is unchanged: this corrects a bias, it does not inflate scores. Generosity is still part of integrity, not a bonus on top of it; a debater who confronts everything cleanly and concedes nothing still forfeits the whole 15% the component is worth. Separately, every scoring verdict now shows what it cost in DIS points, because “Confronted the point” is 45% of the score and is where most debaters lose it, yet it appeared on the page as a small marker while a single foul drew a visible badge.
v0.7. Role is measured, not declared. On the programs people actually watch, the host argues, and until now the most argumentative voice on some shows was the one voice we never scored. Asking a hard question, however hostile, is refereeing and is captured as a point the debaters must answer. A host whoadvances or defends a position of their ownhas stepped onto the floor, and from that moment their claims can be challenged and their conduct is scored. The judge decides this from the transcript; the submitter's label is only a hint. Two denominators changed to make that fair: “Argued clean” now counts only words spoken while arguing, because procedural speech would dilute a host's foul rate toward zero and let them out-score every debater on cleanliness by doing their job; and “Let them talk” counts opponents' turns only while you were in the room, which also fixes guests who join late or leave early. A new rule separates the procedural cut-off (calling time, moving on, handing over the floor) from argumentative talk-over; only talk-over is a foul, for hosts and debaters alike. Nothing already published was rescored: with role findings absent, v0.7 reproduces v0.6 exactly and not one score moved.
v0.6. The Panel Standard. Most debates people actually watch are panels (three to six voices, sometimes all talking at once), and a rubric built for two chairs would convict panelists of dodges they were never offered. In a 1v1, silence toward a point is evidence; on a panel it is usually just not having the floor. Every point now names who owes it an answer, and an adverse verdict requires both that the point was yours and that you had a realistic opportunity to respond. Two no-fault verdicts complete the vocabulary: not-addressed (the point was never yours) and covered-by-ally(a co-aligned panelist answered it and you ceded the floor: division of labor is teamwork, not evasion). Interruption rate is now measured against every other debater's floor time, not one arbitrary opponent's: for two debaters the math is identical, so no published score moved. Conduct stays absolute in every format: on a show where everyone screams over everyone, everyone scores badly on Let Them Talk, and that is the finding. A debater who faced fewer than three challenges carries a “limited evidence” disclosure wherever their score appears. There are no team scores, no aggregates, and no ordering by result: four people on a card is exactly where a reader most expects a leaderboard, and exactly where we still refuse to print one.
v0.5. Added one rule, for the analogy nobody will defend. The existing rule judges a single analogy against a single proposition, so it cannot reach a speaker who offers an analogy, gets challenged on whether the comparison actually holds, and answers with a different analogy instead of defending the first. Each comparison in that sequence can be individually sound, which is exactly why it survived: the challenge to the mapping was a live question, and substituting a fresh illustration changes the subject while looking like sustained engagement. Charging it requires three quoted receipts (the analogy, the challenge to the mapping, and the substitution), and defending a mapping badly is not the foul, since being wrong is never scored. The penalty is deliberately small: one instance costs about a point, a habit of it costs several. The change came out of a debate that produced no fouls at all, where the substitution was the only thing the register could not describe.
v0.4. Closed the engagement-theater gap. An audit of the library found that credit scores ran opposite to integrity: the two highest-scoring performances carried the two lowest credit scores, while the two lowest carried near-perfect ones. The cause was the denominator: credits were measured per 1,000 words, so a short debate manufactured credit rate out of nothing (one debater scored 15.8 credits per 1,000 words in a 34-minute debate against 7.6 in a two-hour one, same person, same behaviour). Credits are now measured per challenge faced. We also tightened the iron rule (a confronting quote must now do argumentative workon the proposition, not merely name it), and added three rules for bad-faith moves that had nowhere to land: treating a mainstream position as absurd to avoid engaging it, abandoning a shared standard of evaluation mid-argument without defending the exit, and offering an analogy whose structure doesn't map onto the argument. Rescoring moved every published performance by 3 points or less. Deliberately notadded: a second penalty for "talked around it": that verdict is already priced at 0.15 against 1.0 on the heaviest component, and charging it twice for one act is double jeopardy.
v0.3.No component can bottom out. In v0.2, "argued clean" and "let them talk" ran on a linear scale with a hard floor at zero, so a debater with 5.4 weighted fallacy points per 1,000 words and one with 15 both scored zero, and near the floor a single extra fallacy could swing a 30%-weighted category by more than half. Both now decay exponentially, like the credits component always had. Rescoring every published debate moved no one by more than 2 points, inside the stated ±5–8 band, except one debater whose zeroed component resolved to 26.
v0.2.The crux standard: neutral, quotable cruxes; the "talked around it" verdict for engagement theater; the iron rule (no confronted verdict without a confronting quote); the hostile second judge; record denial as a scored rule; rate-based scoring comparable across debate lengths.
v0.1.The original point ledger, fallacy and credit rules, and whole-debate reconciliation. Retired after calibration showed additive scoring rewarded whoever faced the most questions, and that responsiveness judged "did they reply" when it should have judged "did they confront."