How the new TOEFL is actually scored.
Band 1–6 in 0.5 steps, per section and overall. Most of the exam is deterministic; the open-response tasks are scored by ETS’s own AI engines, with humans checking a sample. Here is exactly which is which, and why every prediction should come back as a range.
1–6 per section, overall is the mean.
Each of the four sections reports a band from 1 to 6 in 0.5 steps. Overall is the mean of the four section bands, rounded to the nearest half band — no section is weighted more than another in that average.
| Band | CEFR |
|---|---|
| 6 | C2 |
| 5 – 5.5 | C1 |
| 4 – 4.5 | B2 |
| 3 – 3.5 | B1 |
| 2 – 2.5 | A2 |
| 1 – 1.5 | A1 |
A legacy 0–120 figure is also dual-reported until January 2028, as a labelled secondary overall conversion — never per section. Full table on the score conversion page.
Rule-based by default, AI on four tasks.
Reading, Listening and Build a Sentence never touch a model — every one of those items is scored by deterministic rules. The exception is four constructed-response tasks, scored by ETS’s own proprietary AI engines — human raters step in on low engine confidence and a QA sample, so it is neither “examiners grade you” nor “100% AI, no humans involved”.
Rule-based — no AI
| Section | Task types | How it’s scored |
|---|---|---|
| Reading | Complete the Words, Read in Daily Life, Read an Academic Passage | 1 point per correct answer, no negative marking |
| Listening | Listen and Choose a Response, Listen to a Conversation, Listen to an Announcement, Listen to an Academic Talk | 1 point per correct answer, no negative marking |
| Writing | Build a Sentence | Deterministic exact-match against an ordered key — not a report |
AI-scored — 13 items
| Section | Task | How it’s scored |
|---|---|---|
| Writing | Write an Email | 0–5, ETS’s published rubric |
| Writing | Write for an Academic Discussion | 0–5, ETS’s published rubric |
| Speaking | Listen and Repeat | AI-scored across the full 7-clip attempt |
| Speaking | Take an Interview | AI-scored across all four questions |
Never compute a band from percent correct: the items presented in Reading and Listening include unscored tryout items mixed in, so a raw percentage of what you saw is not the same as your scored percentage. ETS itself states that reported scores are not a linear transform of raw percentage — there is no public raw-to-band table, and ScoreFluent does not invent one.
| Section | Items presented | Items scored | Unscored tryout items |
|---|---|---|---|
| Reading | 50 | 35 | 15 |
| Listening | 47 | 35 | 12 |
Reading presents 50 items across the section but scores only 35 — up to 15 are unscored tryout items, interleaved among the scored ones so you cannot tell which is which. Listening presents 47 items and scores 35, with up to 12 unscored tryout items mixed in the same way.
Why a band is a range, not a point.
| Section | Reliability | SEM (band) |
|---|---|---|
| Reading | 0.86 | ±0.37 band |
| Listening | 0.88 | ±0.35 band |
| Writing | 0.87 | ±0.36 band |
| Speaking | 0.94 | ±0.22 band |
| Overall | 0.90 | ±0.32 band |
A section SEM of roughly 0.35 band means a real ±0.7-band 95% confidence interval around any predicted band. ScoreFluent renders every prediction as a range for exactly this reason, and does not celebrate a half-band move between attempts — it is inside the exam’s own measurement noise. Most tools show a single confident-looking number; a range is the statistically honest choice, and it is also what the real exam’s reliability data supports.
Your best section scores, combined.
TOEFL’s MyBest score reports the highest band you achieved in each section across any administration within the last two years, then averages those four best section scores into a MyBest overall — even if that exact combination never occurred on one test date. Many institutions accept MyBest alongside, or instead of, a single sitting’s overall score; check directly with your target programme, since acceptance of MyBest is not universal.
Check the requirement against both scores
A requirement is often set per section as well as overall — a programme might ask for band 4 overall but require band 4 in Writing specifically. Check whether a requirement is overall, per section, or both: a per-section floor can be the binding constraint even when your overall average looks fine, and MyBest does not change what a single sitting needs to clear if a programme requires one-sitting scores.
Scoring questions, answered.
Does AI grade the new TOEFL?
For the four open-response tasks, yes — and it is ETS’s own published design, not a workaround. ETS’s proprietary AI engines score Write an Email, Write for an Academic Discussion, Listen and Repeat and Take an Interview, with human raters stepping in on low engine confidence and a QA sample. Reading, Listening and Build a Sentence are 100% rule-based and never touch a model at all.
What is the overall TOEFL band calculated from?
Overall is the mean of the four section bands — Reading, Listening, Writing, Speaking — rounded to the nearest half band. There is no separate overall algorithm; it is a straight average of the four numbers you already see.
Why does ScoreFluent show a band range instead of one number?
Because that is what the exam’s own published reliability data supports. Section-level standard error of measurement runs from about 0.22 band (Speaking) to 0.37 (Reading), which works out to roughly a ±0.7-band 95% confidence interval overall. A single confident-looking number would overstate the precision the underlying statistics actually have — a half-band shift between two attempts is inside measurement noise, not real improvement.
Which TOEFL tasks are scored by rules vs AI?
All of Reading, all of Listening, and Build a Sentence are scored by deterministic rules — no AI touches them. The four open-response tasks — Write an Email, Write for an Academic Discussion, Listen and Repeat, Take an Interview — are scored by ETS’s AI engines, with human raters stepping in on a QA sample.
How reliable is TOEFL section scoring?
ETS’s field-test reliability figures: Reading 0.86, Listening 0.88, Writing 0.87, Speaking 0.94, overall 0.90. Human–machine agreement on the AI-scored tasks runs r = 0.86 for writing (against 0.85 human–human) and r = 0.89 for speaking (against 0.96 human–human) — writing is near the human ceiling, speaking has a real remaining gap.
ScoreFluent is not affiliated with or endorsed by ETS. Figures on this page come from ETS’s TOEFL iBT test content page and its TOEFL iBT Technical Manual (RR-25-12), which publishes the reliability, SEM and human–machine agreement tables cited above. Both are pre-launch documents ETS may revise with operational data — check the official source before a decision hangs on a number here.
See also what changed in 2026 and the full band conversion table.
See your band range, in minutes.
Every quick score comes back with the same machine-scored/AI-scored split described above.
Create free account