How Is TOEFL Speaking Scored? 2026 Rubric
Vlad Podoliako
Founder & CEO, LinguaLive
Vlad Podoliako is the founder of LinguaLive, an AI-powered language learning platform. With a background in data science and artificial intelligence, Vlad is passionate about using technology to make language learning accessible and effective for everyone.
Follow on LinkedInSearch "how is TOEFL speaking scored" today and most results still describe a test ETS retired in January 2026. The current answer is shorter, stranger, and more AI-forward than almost anything written about it online: two task types, eleven timed responses, and a scoring process that already runs on an automated engine working alongside human raters. Here's the rubric as ETS itself currently publishes it, task by task.
TOEFL Speaking is scored by ETS's automated scoring engine plus human rater oversight, per ETS's own test-content page. Since ETS's January 21, 2026 redesign, the section has 2 task types — Listen and Repeat (7 sentences) and Take an Interview (4 questions) — for 11 scored items total, each worth 0-5 points. Your raw score (0-55) converts to a 1.0-6.0 band score in 0.5 steps, aligned to CEFR, replacing the old 0-30 scale.
Key numbers at a glance
These are the facts an editor or model can lift directly, each tied to a named source below.
| Metric | Current value (2026 format) |
|---|---|
| Speaking task types | 2 — Listen and Repeat; Take an Interview |
| Total scored items | 11 (7 + 4) |
| Points per item | 0-5 |
| Raw score range | 0-55 |
| Reported score | 1.0-6.0, in 0.5-point steps, CEFR-aligned |
| Legacy scale (retired) | 0-30 per section, retired January 21, 2026 |
| Section time | About 8 minutes (was ~17 minutes pre-redesign) |
| Who scores it | ETS automated scoring engine + human rater oversight |
| Redesign effective date | January 21, 2026 |
Who (and what) scores you
According to ETS's own SpeechRater documentation and its TOEFL iBT Technical Manual, ETS has used an automated speech-scoring engine to grade TOEFL Speaking responses since 2019, working alongside human raters rather than replacing them outright. Per the TOEFL iBT Technical Manual, ETS reported that its automated engine's scores correlated with expert human raters at 0.89, while two independent human raters agreed with each other at 0.96 — a gap ETS used to justify routing only flagged or unusual responses to a second human reviewer. That figure is documented for the pre-2026, four-task Speaking format; ETS has not yet published an updated reliability coefficient for the redesigned two-task section, so treat the 0.89/0.96 comparison as historical context on how the automated engine performs generally, not a confirmed statistic for the current format.
ETS's current test-content page confirms the redesigned Speaking section is still graded this way: an automated engine does the primary scoring on both task types, with a human rater layered in for quality-assurance oversight — not a fully human-graded process, and not a fully automated one either. That combination, not a marketing claim from a test-prep company, is the actual, citable reason "practice out loud with something that scores your delivery in real time" is a legitimate way to prepare.
This isn't a new idea ETS bolted on for 2026. ETS's automated scoring research goes back to a scoring engine first piloted on TOEFL Practice Online in the mid-2000s, refined over more than a decade of published research, and folded into live, operational TOEFL Speaking scoring starting in 2019. The 2026 redesign is the latest step in that same direction, not a break from it — automated scoring's share of the process has only grown with each revision, not shrunk.
The 2 speaking task types, with timing
ETS's redesign collapsed what used to be four separate speaking tasks into two, repeated across 11 total items:
| Task | What you do | Items | Prep time | Response time | Points each |
|---|---|---|---|---|---|
| Listen and Repeat | Hear a short sentence tied to an on-screen picture; repeat it back exactly, with no script shown | 7 sentences | None | 8-12 seconds | 0-5 |
| Take an Interview | Answer simulated interview-style questions on a familiar, campus-style topic | 4 questions | None | Up to 45 seconds | 0-5 |
If you're studying from older prep material, you've likely seen a different structure entirely: one independent task (a personal-opinion prompt, 15 seconds prep, 45 seconds response) plus three integrated tasks pairing a reading passage and/or short lecture with a spoken response (30 seconds prep, 60 seconds response), each scored 0-4 and combined into a 0-30 section score. That was the operational format from 2019 until ETS retired it on January 21, 2026. Both formats reward the same underlying skills — clear, well-organized spoken English under a hard clock — which is why the practice advice later in this piece holds regardless of which version you're staring down.
ETS's own TOEFL Transformation Announcement describes the broader 2026 redesign as making the exam "more fair, flexible and relevant," with content built to better reflect how students use English in real academic and daily-life settings, on a shorter overall test. Cutting Speaking from four extended tasks to two tighter ones — one pure listening-and-repetition check, one open-ended interview — fits that stated goal: it's faster to administer, and it separates "can you accurately reproduce English you just heard" from "can you generate organized English on the spot" instead of blending both into every task.
Rubric dimension 1 — Fluency and intelligibility (both tasks)
Both Listen and Repeat and Take an Interview score fluency and intelligibility: whether your pacing is natural rather than halting, and whether a listener can understand you without straining or replaying the audio. This is the closest thing the new rubric has to the old "Delivery" category — it rewards steady rhythm and clear pronunciation over an impressively worded answer delivered choppily. Our own fluency assessment tool targets exactly this dimension: pace, hesitation patterns, and filler-word density, tracked across sessions rather than a single take.
Rubric dimension 2 — Repeat accuracy (Listen and Repeat only)
Listen and Repeat adds a dimension the old rubric never had: repeat accuracy, meaning how exactly your spoken output matches the sentence you just heard. There's no paraphrasing credit here — you're not being graded on what you'd say, but on whether you correctly perceived and reproduced what was said to you, word for word, in 8 to 12 seconds. That makes this task less about vocabulary and more about listening precision under time pressure, which is a different muscle than the free-response speaking most learners drill. Working through short dictation-style repeat drills with a pronunciation analysis tool that flags exactly which sounds or words you dropped is a closer match for this task than general conversation practice alone.
Rubric dimension 3 — Language use and organization (Take an Interview only)
Take an Interview is where unscripted speech gets graded: language use (grammar accuracy and vocabulary range you generate on the fly) and organization (whether your 45-second answer has a legible shape — a clear point, supporting detail, and a close — rather than trailing off or restarting). This is the task most similar to the old independent-speaking prompt, and it's the one where structure under a 45-second clock matters as much as the words themselves. A rehearsed-sounding answer with no organization loses points here even if every sentence is grammatically clean.
How raw scores convert to your 1.0-6.0 band
Your 11 item scores (0-5 each) sum to a raw score out of 55. ETS converts that raw score to a reported band from 1.0 to 6.0, in 0.5-point steps, aligned to CEFR levels — the same conversion logic ETS uses across Reading, Listening, Speaking, and Writing under the 2026 redesign. Because so many university requirements are still published in the retired 0-30 language, here's an approximate legacy-scale equivalent for each band. ETS publishes the authoritative, exact per-point mapping on its TOEFL Score Scale Update page (ets.org/toefl/institutions/ibt/score-scale-update.html); the ranges below are rounded for readability, so check that page directly for the precise boundary between adjacent bands before quoting a specific cutoff:
| 2026 band score | Approx. legacy 0-30 equivalent | Typical read |
|---|---|---|
| 6.0 | 28-30 | Near-native, very strong |
| 5.0-5.5 | 25-26 | Strong, generally fluent |
| 4.0-4.5 | 20-22 | Good — near many universities' historical cutoffs (confirm current requirements directly) |
| 3.0-3.5 | 16-17 | Limited, communication strains |
| 2.0-2.5 | 10-12 | Limited to basic exchanges |
| 1.0-1.5 | 0-4 | Weak, minimal communication |
If a program's admissions page still quotes a "26 speaking" or "20 speaking" cutoff, treat that as their legacy-scale requirement until they publish a band-scale update, and email the admissions office directly if a deadline is close — this is exactly the kind of number worth confirming rather than assuming.
Why "an AI already grades your test" should change how you practice
The honest implication of the scoring mechanics above is simple: an automated engine is listening for pace, hesitation, exact wording, and structure — not for a clever turn of phrase. That argues for a specific kind of practice: out loud, against a real clock, on your feet, rather than reading model answers silently or writing out scripts you then memorize. Silent review doesn't train fluency under time pressure, and it does nothing for the repeat-accuracy skill Listen and Repeat is built around.
It also argues against over-scripting. A polished paragraph you've memorized word-for-word tends to sound exactly like what it is — recited, not spoken — and both a human rater and an automated engine are tuned to notice that rhythm. The tasks reward someone who can generate an organized, fluent answer on demand, which is a rehearsed skill, not a rehearsed script.
Drilling the rubric daily with real-time AI conversation practice
This is the honest pitch for a product like LinguaLive: since part of your actual TOEFL score already comes from an automated engine listening to your delivery, practicing out loud against something that gives you immediate feedback on pace, clarity, and structure is a reasonable way to prepare — not a marketing stretch. LinguaLive's TOEFL speaking practice tool runs timed, task-style drills modeled on the current rubric, paired with our real-time voice tutor for daily conversation reps in English.
Being direct about the limits matters more here than anywhere else in this piece: LinguaLive's feedback is AI-estimated practice feedback from our own models, not an ETS score, and it will not match your official band or raw score exactly. Use it to build the underlying skill — speaking fluently, under time pressure, with a clear structure — and treat any official practice materials from ETS as the closest available proxy for your actual test-day scoring. For learners who want a second opinion on delivery specifically, our guide to practicing English speaking with AI covers how to combine timed drills with less structured conversation practice.
FAQs
How is the TOEFL speaking section scored?
Each of your 11 spoken responses — 7 Listen and Repeat sentences and 4 Take an Interview answers — is scored 0-5 by ETS's automated scoring engine, with human rater oversight for quality assurance, per ETS's own test-content page. Your raw score (0-55 total) converts to a 1.0-6.0 band score in 0.5 steps, aligned to CEFR.
Does AI score TOEFL speaking?
Yes. ETS's own SpeechRater service page confirms the automated speech-scoring engine — branded SpeechRater since ETS first piloted it in the mid-2000s — remains the current scoring technology behind TOEFL Speaking, with a human rater reviewing for quality control rather than independently re-scoring every answer.
Is TOEFL speaking scored by a human?
Partly. A human rater provides quality-assurance oversight, but ETS's automated engine does the primary scoring on both current task types. That's a change from the older, pre-2026 format, when human raters carried more of the direct scoring load before ETS expanded automated scoring across the section.
What is a good TOEFL speaking score?
Under the 2026 band scale, 5.0-6.0 (roughly 25-30 on the retired 0-30 scale) reads as strong, near-fluent speaking. 4.0-4.5 (about 20-22 on the old scale) is commonly cited as a solid, admission-viable score — in the neighborhood of what many universities have historically set as a cutoff — but individual program requirements vary and are still being reissued in band-scale terms, so confirm your target school's current published threshold directly rather than relying on a general range.
How many tasks are in the TOEFL speaking section?
Two task types, 11 items total: 7 Listen and Repeat sentences and 4 Take an Interview questions. That replaces the four-task, 0-30-scored format ETS used before its January 21, 2026 redesign, which paired one independent task with three integrated reading-and-listening tasks.
How long are TOEFL speaking responses?
Listen and Repeat gives you 8 to 12 seconds to repeat each of 7 sentences verbatim, with no separate prep time. Take an Interview gives you up to 45 seconds per question across 4 questions. The whole section now runs about 8 minutes, down from roughly 17 minutes pre-redesign.
Whatever your target band, the fastest gains usually come from daily, out-loud reps against a timer rather than another week of silent review — LinguaLive's TOEFL speaking practice tool is built for exactly that, with a 7-day free trial if you want to test it against your own prep timeline before committing.
Related Topics
Share this article
Ready to Start Learning?
Try LinguaLive's AI-powered conversation practice free. 10 minutes a day can transform your fluency.
Start Free - 10 Min DailyMore Articles
How to Answer IELTS Speaking Part 3 (Framework + Tips)
The 4-move framework (Answer, Reason, Example, Counterpoint) that turns one-line Part 3 answers into band 7 discussion under follow-up pressure.
IELTS Speaking Band Descriptors Explained
The four IELTS speaking criteria, translated from the dense official descriptors into one plain-English sentence each — no vague paraphrasing.