← Steelmanned 🛡️ · all debates
Published methodology · DIS Rubric v0.2

How we score

Steelmanned does not score who won. It scores debate integrity: who argued in good faith, who confronted the other side's actual points, and who played games instead. The result is a DIS — Debate Integrity Score— for each debater, and every point of it traces to an exact quote and a rule on this page. If a score can't show you its receipts, we don't publish it.
The creed

Integrity, not victory. Persuasion is never scored. A debater can lose the crowd and score 95; a debater can win the crowd and score 40.

Conduct, not correctness. We never score who is factually right. A person can be wrong in good faith and score high. Fact checking is someone else's job; honest engagement is what we measure.

The whole debate is the record. A point answered thirty minutes later counts as answered, in full. Tangents are not dodges. No adverse verdict is issued until the entire debate has been searched for a response.

Total auditability. Score → components → verdicts → events → quotes. Every layer is one click deeper, and the bottom layer is always the debaters' own words.

What the score is made of

A DIS is a weighted blend of four rates — rates, not raw counts, so a 30-second exchange and a 3-hour debate are scored on the same scale:

Confronted the point (45%). Of the challenges aimed at you, weighted by how central they were: how many did you actually confront?

Argued clean (30%). Logical fallacies and record denial per 1,000 words spoken, weighted by how central the point they touched was.

Gave credit (15%).Concessions, common ground, "I don't know," steelmanning — with diminishing returns, so volume can't farm it.

Let them talk (10%).Interruptions of the opponent's answers, per opponent turn. Moderators are tracked but never scored.

The crux standard — our core idea

Every tracked point gets a crux: the specific proposition the point demands be confronted — not the topic, the proposition. Cruxes must be position-neutral: each one offers a genuine branch to rebut or defend, never "concede X" as the only satisfying move. A partisan of either side should read every crux as a fair question.

The iron rule: a "confronted" verdict must quote the exact sentence that confronts the crux. If no such sentence can be pasted, the verdict is "talked around it"— engagement theater, scored barely above outright evasion. Saying "that's a good question and we need to address it" and then not addressing it earns exactly what it deserves.

Confronted1.0Engaged the actual proposition: agreed, refuted with argument, or gave a reasoned refusal the crux allowed
Conceded1.0Gave the point (also earns a credit)
Partially confronted0.5Real engagement with part of the crux
Talked around it0.15Engagement theater: responded to the topic, gave context, answered a weaker version — never the proposition
Evaded0Changed the subject or refused without reasons
Mutually dropped / out of time / moderator cutNo fault; excluded from the denominator
The rules — deductions

Base values below are multiplied by the importance of the point they touched (×3 central, ×2 supporting, ×1 peripheral; jokes and asides are never penalized). Citing evidence within an argument is not an appeal to authority — the fallacy is authority offered instead of an argument. Defending or contextualizing your own record is not record denial — only misrepresenting what the record shows.

Strawman2.5Rebutted a position the opponent didn't hold
Quote / context stripping2.5Used opponent's words to mean what they didn't mean
Record denial3.0Confronted with their own documented words, denied or reframed them instead of owning or defending them
Whataboutism2.0Answered a charge with a different charge
Ad hominem2.0Attacked the person instead of the point
Appeal to authority2.5“The experts / the law say so” offered instead of an argument
Motte-and-bailey2.5Defended a modest claim, then resumed the bold one
False dilemma2.0Framed two options as the only options
Gish gallop2.0Flooded claims faster than any could be defended
Non-answer filibuster2.0Talked long, answered nothing
Dodge / subject change2.5Changed the subject under direct challenge
Refused direct question3.0Declined to answer at all, no reason given
Interruption0.5Cut in during the opponent's answer (moderator exempt)
Talking over1.0Kept talking through the opponent's turn
The rules — credits

Good faith earns points. Credits carry the same importance multipliers as deductions, and the scorecard shows them with the same prominence — several debaters have scored poorly on fallacies while posting the best credit rate on the panel. Both facts appear on the card.

Clean concession3.0“You're right about that” — full stop
Partial concession1.5Gave real ground while holding the rest
Common ground1.5Named agreement and built on it
Direct answer to a hard question1.5Took the hard question head-on
Returned to an open point2.0Came back unprompted to answer something left open
Acknowledged uncertainty1.0“I don't know” — when they didn't
Owned an error3.0Admitted being wrong or apologized for a past position
Steelman2.0Stated the opponent's case at its strongest before rebutting
How a debate gets scored

1 · Transcription & attribution. The debate becomes a timestamped transcript with every utterance attributed to a speaker. Uncertain attributions are flagged — and anything built on them is excluded from the score.

2 · The point ledger.An AI judge extracts every claim, challenge, question, and thought experiment that shaped the debate, writes each one's crux, and weights its importance. Points raised by moderators or played clips count as challenges — but their authors are never scored.

3 · Event detection. Judges walk the transcript hunting deductions andcredits under a strict discipline: when in doubt between two rules, pick the cheaper one; when in doubt whether an event exists, it doesn't.

4 · Whole-debate reconciliation. Only after the entire debate is read does any point receive a verdict — so the answer given thirty minutes later, unprompted, counts in full.

5 · The hostile second judge.An independent adversarial pass attacks every "confronted" verdict, checking each confronting quote in context. Its only power is to demote — it can never promote. Verdicts that survive have been through two independent judges, one of whom was trying to kill them.

6 · Rollup.The four component rates are computed and blended into each debater's DIS, alongside the evidence badge (challenges faced, words spoken) that tells you how much data the score rests on.

Built to be attacked

An integrity referee will be accused of bias by whoever its scores embarrass. We designed for that day:

The rubric is published and versioned.Disputes are about whether a rule was correctly applied to a quote — never about an AI's opinion. Scorecards name the rubric version that scored them.

Categories are locked. The judging pipeline is technically barred from inventing rule categories — every event must cite one of the rules on this page.

Uncertainty is excluded, not hidden. Low-confidence findings appear on the scorecard but never count toward the score.

Scores are honest about their precision. Independent re-runs of the same debate agree within roughly ±5–8 points. A DIS is a measurement with error bars, not a verdict from on high.

We publish our own failures.During calibration, our checks caught the pipeline writing biased crux questions and inventing rule categories. Both were fixed structurally — neutrality requirements and schema-enforced categories — and both incidents are part of this public record, because a referee that hides its corrections doesn't deserve your trust.

What we will never do: score who is right, score motives or character, or accept a request to tilt a scorecard. Every complaint about a score should arrive as: this quote, under this rule, was judged wrongly. Those we want to hear.

Version history

v0.2 (current).The crux standard: neutral, quotable cruxes; the "talked around it" verdict for engagement theater; the iron rule (no confronted verdict without a confronting quote); the hostile second judge; record denial as a scored rule; rate-based scoring comparable across debate lengths.

v0.1.The original point ledger, fallacy and credit rules, and whole-debate reconciliation. Retired after calibration showed additive scoring rewarded whoever faced the most questions, and that responsiveness judged "did they reply" when it should have judged "did they confront."