The scientific basis of Dialectica
Most tools that grade student work do not explain where their criteria come from. We publish ours. This page lists every design decision that has published backing — the source, the finding, and what we do because of it — and then lists the opposite: what the literature does not support, and which we therefore do not claim.
Why this page exists
Most tools that score student work do not explain where their criteria come from. We publish ours.
Dialectica did not invent a debate rubric. The weighting is the World Schools marking standard, the one used by the international championships; the sheet we ship first is its Italian curricular adaptation, published by the Società Nazionale Debate Italia. The argument-quality dimensions come from the computational argumentation literature, and so do the measurement cautions.
This page also lists the opposite: what the literature does not support and which we therefore do not claim. That second list is what makes the first one credible.
How to read each section: the decision → the source → the finding → what we do because of it. No reference appears here without the specific finding that makes it relevant; a citation without a finding is decorative bibliography.
01Where the weights come from, and they are not ours
The decision. The assessment is split across content 40 % · style 40 % · strategy 20 %, out of 100, which then converts to the local grading scale.
The sources, in this order. The weighting is the World Schools marking standard, the format of the international championships since 1988. The sheet we ship first is its Italian curricular adaptation: Presutti, G. / SN-DI, Griglia di valutazione per il Debate preparato (March 2020) — five blocks scored 0.4 to 2.0 out of 10: two for content (argumentative and documentary research competences; active listening and quality of rebuttal), two for style (para-verbal; non-verbal), one for strategy (strategic competences and use of POI).
The finding. Four tenths, four tenths, two tenths — the same vector in the international standard and in the national curricular sheet, because the second is an adaptation of the first. That fixes what can be claimed: the weighting is not a national choice we inherited, it is the tournament standard; the national sheet is what turns it into descriptors a teacher can mark against.
What we do. That vector is the default and it does not move. A school that needs it can switch on 50/30/20, published by Russo & Refrigeri (2023) for curricular debate by subject, which argues its own deviation: content weighs more because studying sources other than the textbook is the precondition for everything else. Every assessment records which weighting produced it — if the weights are configurable and the active one is not stored, two cohorts stop being comparable and nobody notices.
What travels, and what does not. The weighting travels: every school on the World Schools circuit already uses it. The descriptors, the bands and the grading scale do not — those are the national sheet, and ours is the Italian one because Italy is the first market. We will not ship a preset for a country whose official assessment sheet we have not read: a preset written from a secondary description is an invented rubric.
02What cannot be assessed on text — and we say it first
The decision. Dialectica assesses a student’s written turn. It does not assess voice, gesture or presence.
The source and the finding. In the SN-DI griglia, two of the five blocks — para-verbal (tone, rhythm, pauses, reading versus esporre a braccio) and non-verbal (eye contact, gesture, posture) — are entirely prosodic and bodily. They are 4 points out of 10: 40% of the rubric, and no tool that reads text can score them.
That this is a boundary of modality and not a shortcoming of the build has independent confirmation: Ruiz-Dolz et al. (2021), building the VivesDebate corpus from a real university tournament, explicitly excluded oral fluency, grammatical correctness and non-verbal language from their assessment file, because those aspects — which the jury did use — were not reflected in the textual corpus.
What we do. Both blocks appear on the student’s screen flagged, with the sentence that their teacher assesses them in class because the system reads text. On text, Dialectica covers the content block in full and the part of strategy that does not depend on live clash. Which makes us a between-debates training tool, complementary to running the debate in class and not a substitute for it. One caution on the 40 % itself: it is a property of this rubric family, not of debate assessment in general. Each market's sheet gets its own figure, stated with the preset rather than repeated as a slogan.
03The three dimensions come from the canonical systematisation
The decision. Contenuto, stile and strategia are not categories invented for the product.
The source. Wachsmuth, H. et al. (2017), Computational Argumentation Quality Assessment in Natural Language, EACL — the reference systematisation: 15 argument-quality dimensions organised into three families, derived from the earlier theories and validated on the Dagstuhl-15512 corpus.
The finding, and the mapping.
| Dialectica dimension | Wachsmuth et al. family |
|---|---|
| contenuto | cogency (local acceptability, relevance and sufficiency) + global sufficiency |
| stile | effectiveness (credibility, emotional appeal, clarity, appropriateness, arrangement) |
| strategia | reasonableness (global acceptability and relevance) |
One correspondence comes out especially clean: the authors define global sufficiency as the argumentation adequately rebutting the counterarguments that can be anticipated — which is exactly the strength Dialectica labels as anticipation of objections.
The nuance we also publish. The same work reports that the three families correlate strongly with one another. That is: three separate dimensions carry less independent information than their separation suggests. It does not invalidate the rubric, but it fixes what can be said: the separation is pedagogical, not orthogonal. It exists because a student needs to know where to improve, not because the three measure independent things.
What we do not claim: that Dialectica implements the 15 dimensions. The mapping is of the three high-level ones.
04Fallacies: how many can be detected, and how reliably
The decision. The product identifies a broad catalogue of fallacies with operational definitions, but any aggregate measurement is done over families, not over the full catalogue.
The sources and the finding. The performance of automatic fallacy classification falls with the number of classes:
| Work | Classes | Result |
|---|---|---|
| Goffredo et al. (2022) | 6 | F1 84% at sentence level |
| Goffredo et al. (2023), EMNLP | 6 | mean F1 0.74 (detection + classification) |
| Jin et al. (2022) | 13 | F1 58.8% |
| Alhindi et al. (2022) | 28 | F1 between 17% and 62% depending on domain |
| Habernal et al. (2018), German Argotario | 6 | accuracy 50.9%, macro-F1 42.1% |
Above six classes, performance collapses.
And the human figure, which is the important one. Habernal, Pauli & Gurevych (2018) measured how well a person identifies the type of fallacy: macro-F1 of 77%, with enormous variation by type — appeal to irrelevant authority 95%, red herring around 60%. Some fallacies are intrinsically harder to agree on than others, and that is a property of the phenomenon, not of the annotator. But the collective vote converges: with three or more judgements, the aggregated label matched the expert one in 90 of 92 cases (97.8%).
What we do. Three things.
- 01
The product catalogue is broad because its job is to explain to the student, and a precise explanation is worth more than a coarse label.
- 02
Any performance figure is reported broken down by family, never as a single global number that hides the fact that some are easy and others are not.
- 03
The machine’s label is not the last word; the teacher’s decision is.
A future line, not yet in the product. Goffredo et al. (2025) introduce fallacy repair: beyond detecting them, generating a reformulated version of the argument without the fallacy. For a fourteen-year-old, “this is a false cause” is a label and “here is your argument without it” is a lesson. It is plausibly the best formative feedback in this whole literature, and it is on our roadmap, not in today’s product.
05The teacher decides, and the design is built so that they can
The decision. The AI reads and proposes, the system structures, the teacher decides. And the proposal is presented in a deliberately unassertive way.
The source. A study published in Frontiers in Psychology (2026) with a 2×2 design and 214 teachers who assessed the same text, varying only the number shown by an automated report and its visual highlighting.
The finding. The effect of the numeric cue on the teacher’s judgement, measured as η²p:
| Aspect assessed | η²p |
|---|---|
| Score out of 100 | 0.745 |
| Linguistic expression | 0.697 |
| Originality | 0.648 |
| Overall judgement | 0.579 |
| Logical structure | 0.529 |
The text did not change. That is the magnitude of anchoring, measured. The study closes by calling for neutral report design, independent teacher judgement, human oversight and training on automation bias.
What we do.
- 01
The teacher’s field is never locked, and the value the AI proposed always stays in a separate read-only column: that column is the record of what was proposed, not the field the grade is typed into.
- 02
When the model declares low confidence, the proposal is shown as a band, not a point and the three fields start empty — the reliability gate that Al-Zawqari, Safa & Vandersteen (2026) introduce to soften strong recommendations when confidence is low. At high or medium confidence the fields arrive pre-filled, which is the anchoring risk we are choosing to run and measure.
- 03
The confidence level is frozen into the record of the assessment, so that it can later be checked whether the caution achieved anything.
And we state it falsifiably. If showing a proposal as uncertain does not increase the time a teacher spends deciding, nor the number of corrections they make, then the measure is not working and we will say so in those words.
An honest caveat about this source: the study is about AI-detection reports, not grade proposals; the experimental manipulation is extreme and the sample is university-level. We cite it as evidence that the phenomenon exists and is large, not as a measurement of Dialectica’s case.
06How we measure whether the system resembles a teacher
This is the part almost nobody publishes, and where it would be easiest to overstate.
The objective. Not to “get the grade right”, but to know in which direction and by how much the machine’s proposal deviates from the teacher’s judgement, per dimension and per grade band.
Anchor 1 — what agreement is realistic
Wachsmuth et al. (2017) report that, with seven expert annotators, Krippendorff’s α did not exceed 0.22 on any dimension of argument quality; among the three most concordant it ranged from 0.23 (clarity) to 0.60 (credibility); full agreement among the three was between 17.4% and 44.7%, though majority agreement reached 87.5–98%. Consequence: argument quality is intrinsically a low-agreement judgement, even among trained experts. Any tool promising very high agreement with teachers is either measuring badly or describing itself badly.
Anchor 2 — ranking and distribution are different things
Debatable Intelligence (2025), over more than 600 debate speeches with fifteen human annotations each, finds that strong models reproduce the relative ranking of speeches well while their score distributions do not align with the human ones: large models tend to score globally lower. Consequence: we measure agreement (kappa) and ranking (Kendall’s tau) in the same table, because a system can win on one and lose on the other.
Anchor 3 — the mean can hide everything
Education Sciences (2026) documents proportional bias: automatic systems tend to score weak work relatively higher and good work more strictly. It is range compression, not a constant shift — and an average close to zero is perfectly compatible with severe bias, because the two tails cancel out. Consequence: the teacher–AI difference is reported stratified by grade band and with the slope against the final grade, not only as a mean.
Anchor 4 — the agreement metric also misleads
With grades concentrated on few values (which is exactly the case of a school scale), the kappa coefficient can collapse even when observed agreement is very high: this is the known kappa paradox (Feinstein & Cicchetti, 1990; Byrt, Bishop & Carlin, 1993). Consequence: we report three figures together — raw observed agreement, weighted kappa, and Gwet’s AC2 coefficient (2008). When they diverge sharply, the divergence is the result.
Anchor 5 — the right reference is not a threshold, it is another teacher
The number is only interpretable against how much two teachers from the same school agree when assessing the same debates. It is planned into the design of our pilots.
Anchor 6 — is it quality being measured, or length?
The standard objection to automatic scoring is that a system can agree with humans while measuring the wrong thing. It is documented that automatic judges’ assessments track raw text length far more closely than human assessors do. Consequence: the check is specified — correlating the AI proposal and the teacher’s grade against turn length, side by side — and we do not have the figure yet. We will publish it when we do, whatever it says.
07Why practice is curricular and not extracurricular
The decision. Dialectica is assigned in class, as a compito. It is not an app a student signs up to on their own.
The sources and the finding. Russo & Refrigeri (2023) note that evening and extracurricular projects “mantengono intatto il principio di autoselezione”, which in practice selects the most extroverted and sociable personalities leaving out precisely those weakest in learning; and being extracurricular, the activity falls outside the teacher’s assessment (Refrigeri & Russo, 2020).
The empirical reinforcement is blunt. Habernal et al. (2018), deploying a voluntarily adopted serious game on fallacies: in five weeks and across three recruitment channels they gathered 58 players and 296 arguments; 69% played for less than an hour and only 9% came back after five days. A voluntarily adopted argumentation tool does not retain.
What we do. Curricular assignment by the teacher is not a limitation of our model: it is the only documented route to sustained practice, and the only one that reaches the student who would never sign up in the afternoon. Put as the authors put it: classroom debate remains the place where learning happens; asynchronous training is what lets a student arrive there with hours of practice behind them.
08Inclusion: voice assistance is logged, never scored
The decision. Use of dictation and other aids is recorded as a compensatory measure and is excluded by design from any score calculation.
The sources. De Conti (2025), Debate 4.0, develops inclusivity as one of the three legitimate axes of AI in debate: simplifying a text, creating tailored explanations, rendering a text in voice format, and equity of access to competitions for students with BES or disabilities. Sarcina (2024) argues that rubrics are not only instruments of summative assessment but of continuous feedback, and that the student knowing the indicators in advance is what enables the self-assessment dimension.
What we do. The student sees the rubric — the real descriptors of the griglia, written as instructions — before writing their turn, not only in the final report. And the second half of the accessibility sentence is the half that matters: not “we have dictation”, but voice assistance is logged as a compensatory measure and stays out of the grade.
09Where Dialectica sits in the international state of the art
There are three active groups in this terrain, and none of them works on school debate in Italian.
Munazarat (Arabic)
Khader et al. (2024) publish a corpus of 73 competitive debates; Al-Zawqari et al. (2025) extend it to 110 debates and over 500,000 words from 29 countries; Al-Zawqari, Safa & Vandersteen (2026) reach rubric-aligned feedback with a small open model. They work on recorded human debates, and the recipient of the feedback is the judge or the coach.
Villata and Cabrio (European)
Goffredo et al. (2023, 2025) annotate fallacies at token level over sixty years of presidential debates and build the repair task. They work on adult political discourse.
VivesDebate (Catalan)
Ruiz-Dolz et al. (2021) publish the annotated corpus of a real university tournament, with the per-argumentative-discourse-unit annotation scheme that we adopted as it stands, without inventing a format of our own.
The gap. We have found nothing equivalent for school debate in Italian. We phrase it that way — “we have not found” — and not as “it does not exist”, because it is more honest and harder to rebut.
And that is why we do not export calibrations across languages. Two findings rule it out, and it is worth saying before anyone asks:
Toledo-Ronen et al. (2020, IBM Research)
tested cross-lingual transfer in argument mining: the methods work well for classifying stance and detecting evidence, and clearly worse for assessing quality, because quality survives translation poorly.
Fu & Liu (2025)
evaluated five models as judges across 25 languages: mean cross-language consistency of judgement was a Fleiss’ kappa around 0.3, and neither training on multilingual data nor increasing model size improved it.
Consequence. Any reliability figure we publish will be tied to the language and the rubric it was measured with. We will not say that a calibration obtained in Italian holds for Spanish or German, because the literature says it does not.
10What this page does not claim
This section is deliberate and will not be deleted in future versions.
We do not claim to have measured results
The system is instrumented, not measured: the calibration instruments are built and described, and data collection with real schools is beginning. We do not publish a kappa obtained in a classroom because we do not have one yet.
We do not claim to replace the teacher
Neither pedagogically nor legally. The grade that goes on the record is set by a person.
We do not claim to assess oral style
The 40% para-verbal and non-verbal share of the Italian rubric is outside what a system that reads text can score.
We do not claim the three dimensions are independent
The literature says they correlate; the separation is pedagogical.
We do not claim high agreement with human judges as a goal
Argument quality is a low-agreement judgement even among experts: suspiciously high agreement would be a sign that something else is being measured.
We do not train models on student texts
Any annotated corpus containing student production exists only inside a signed research agreement, with its own legal basis and with the educational institution as a party to it.
We have not reviewed the whole literature
There are Italian works and Italian-language conference papers we have not read yet, and this page gets corrected when we read them.
References
Rubric and pedagogy of debate
- 1Presutti, G. / Società Nazionale Debate Italia (2020). Griglia di valutazione per il Debate preparato. sn-di.it.
- 2Russo, N., & Refrigeri, L. (2023). Il contributo della pedagogia alla valutazione degli apprendimenti attraverso il debate. RicercAzione, 15(2 bis), 221–236.doi.org/10.32076/RA15311
- 3Refrigeri, L., & Russo, N. (2020). Imparare a dibattere nella scuola primaria. Formazione & Insegnamento, 18(1), 349–361.
- 4Sarcina, F. P. (2024). Il debate come strumento innovativo per valutare le competenze degli studenti nella scuola primaria. IUL Research, 5(9), 180–193.doi.org/10.57568/iulresearch.v5i9.553
- 5De Conti, M., & Giangrande, M. (2017). Debate. Pratica, teoria e pedagogia. Milano: Pearson.
- 6De Conti, M. (2025). Debate 4.0: come l’intelligenza artificiale trasforma la didattica. Educare.it, 25(3). ISSN 2039-943X.
- 7Cinganotto, L., Mosa, E., & Panzavolta, S. (2021). Il Debate. Roma: Carocci.
- 8World Schools Debating Championships — Debating Rules, sezione Marking Standard.
- 9DM 742/2017 — Certificazione delle competenze al termine del primo ciclo di istruzione. MIUR.
Argument quality and argument mining
- 10Wachsmuth, H., et al. (2017). Computational Argumentation Quality Assessment in Natural Language. EACL 2017, 176–187. Corpus Dagstuhl-15512 ArgQuality.aclanthology.org/E17-1017/
- 11Ruiz-Dolz, R., Nofre, M., Taulé, M., Heras, S., & García-Fornes, A. (2021). VivesDebate: A New Annotated Multilingual Corpus of Argumentation in a Debate Tournament. Applied Sciences, 11(15), 7160.doi.org/10.3390/app11157160
- 12Goffredo, P., Chaves, M., Villata, S., & Cabrio, E. (2023). Argument-based Detection and Classification of Fallacies in Political Debates. EMNLP 2023, 11101–11112. Dataset ElecDeb60to20.aclanthology.org/2023.emnlp-main.684/
- 13Goffredo, P., Dore, D., Cabrio, E., & Villata, S. (2025). DISPUTool 3.0. ACL 2025 (system demonstrations). Dataset FallacyFix.
- 14Habernal, I., Pauli, P., & Gurevych, I. (2018). Adapting Serious Game for Fallacious Argumentation to German: Pitfalls, Insights, and Best Practices. LREC 2018.aclanthology.org/L18-1526/
- 15Khader, M. M., et al. (2024). Munazarat 1.0: A Corpus of Arabic Competitive Debates. OSACT 2024 @ LREC-COLING, 20–30.
- 16Al-Zawqari, A., et al. (2025). Neural Classification of Argument Elements and Styles in Arabic Competitive Debates. IEEE Access, 13, 115944–115959.
- 17Al-Zawqari, A., Safa, A., & Vandersteen, G. (2026). Structured Feedback for Arabic Competitive Debates Using a Small Language Model. SLM4ED.
- 18Al-Zawqari, A., Ahmed, M., & Vandersteen, G. (2024). Evaluation with Language Models in Non-formal Education. EvalLAC 2024, CEUR-WS Vol-3772.ceur-ws.org/Vol-3772/
Automatic assessment, bias and measurement
- 19Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation (2025). arXiv:2506.05062.arxiv.org/abs/2506.05062
- 20Algorithmic anchoring in teachers’ assessment of writing (2026). Frontiers in Psychology. Disegno 2×2, n = 214.doi.org/10.3389/fpsyg.2026.1889402
- 21Do AI Grading Systems Systematically Differ from Human Teachers’ Grading? (2026). Education Sciences.doi.org/10.3390/educsci16071147
- 22Toledo-Ronen, O., Orbach, M., Bilu, Y., Spector, A., & Slonim, N. (2020). Multilingual Argument Mining: Datasets and Analysis. Findings of EMNLP 2020.arxiv.org/abs/2010.06432
- 23Fu, X., & Liu, W. (2025). How Reliable is Multilingual LLM-as-a-Judge? Findings of EMNLP 2025, 11040–11053.
- 24Feinstein, A. R., & Cicchetti, D. V. (1990); Byrt, T., Bishop, J., & Carlin, J. B. (1993) — sul paradosso di kappa. Gwet, K. L. (2008) — coefficiente AC1/AC2.
How this page is kept
This page carries a date and it expires. It is reviewed before every school year and every time a design decision is closed with published backing. Three rules we impose on ourselves:
Every reference enters with its finding
If we cannot say what the work found, we do not cite it.
Corrections are not deleted
When we have claimed something and it was false, the correction stays on record. It is the proof that we correct against the source rather than defend what we wrote.
If it is not built, it does not count as built
No exceptions, including for things that are two weeks away.
Last reviewed: 2026-08-24.
If you work on any of this and we have got something wrong, we would rather hear it than not.
Write to us