Skip to main content
Inside Recording Practice scoring

How Joe Speakingimproves scoring accuracy

Using Recording Practice, we show how Joe Speaking generates a band score—and how we test and improve its accuracy.

See the scoring process

Recording Practice score

An AI practice estimate, not an official result.

Recording Practice evidence checks on a phoneRecording Practice band score on a phone

One method, test-specific rules

IELTS and CELPIP both use explainable yes-or-no checks, but their criteria and score ranges differ. The diagram, prompt, and public comparison below are an IELTS Recording Practice worked example.

01How the score is generated

AI does not choose your band directly

AI completes yes-or-no checks for each IELTS criterion. The results produce four criterion bands and an overall practice score. Because Recording Practice sends text—not audio—Pronunciation only reflects transcript clarity.

How Recording Practice builds an AI practice bandJoe Speaking runs graduated yes-or-no checks for Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. The yes count maps to a band or narrow band range; because the scorer receives text rather than audio, Pronunciation can only use transcript clarity.1READ THE ANSWERAnswer transcript + questions2RUN YES / NO CHECKSExample: VocabularyAdequate vocabularyYESTopic vocabularyYESFlexible useYESWide rangeYESFull flexibilityNO4 YES → Vocabulary band 83GET FOUR SCORESFluency andCoherence9Lexical Resource8Grammatical Rangeand Accuracy8Pronunciation7YES count sets the band or range4GET THE OVERALL SCORE(9 + 8 + 8 + 7) ÷ 4Average 8.00OVERALL BAND8.0

Pronunciation is a text-based proxy

The scorer cannot hear sounds, stress, rhythm, or intonation. It can only use the clarity of the transcript.

02Evaluation results

How we test scoring accuracy

We use 10 anonymized text samples adapted by Joe Speaking from public candidate examples. We score every sample three times, average the results, and compare that average with its reference band. A smaller difference means a closer result.

Reference source

Public candidate performance examples used only as a reference. View the public reference source

Evaluation input

Ten anonymized Part 3 text samples adapted by Joe Speaking, with matching question context

Each result

Three runs per setup, then one average

How the comparison works

How the reference comparison worksOne sample, prompt and model produce three runs. Their average is compared with the published reference band. The gap is reviewed before the prompt is revised and evaluated again.ONE SETUPPart 3 sample + prompt + modelSame context each runTHREE RUNSRUN 15.0RUN 25.5RUN 35.5THREE-RUN AVERAGE5.33REFERENCE BAND 5.0Gap 0.33Review → revise → evaluate againThe full result table stays visible below

Previous gap

0.85

V33 gap

0.45

Within 0.5

7/10

Complete evaluation matrix

These results come from the v0.6.0 scoring update. We iterated on the scoring prompt 33 times before comparing the final version with the previous scorer. The table shows every run: 10 samples, five scoring setups, and three runs per setup.

150 model scores shown

Sample 01

Reference 5.0

Previous · Gemini 3 Flash
5.5 / 5.5 / 5.5
Gemini 3 Flash
5.0 / 5.5 / 5.5
Gemini 3.5 Flash
6.0 / 5.5 / 5.5
Gemini 3.1 Pro Preview
6.0 / 5.0 / 5.5
Gemini 3.1 Flash-Lite
5.5 / 5.5 / 5.5

Sample 02

Reference 6.0

Previous · Gemini 3 Flash
6.5 / 6.5 / 6.0
Gemini 3 Flash
6.0 / 6.5 / 6.0
Gemini 3.5 Flash
7.0 / 6.5 / 6.5
Gemini 3.1 Pro Preview
6.0 / 6.5 / 6.5
Gemini 3.1 Flash-Lite
6.0 / 6.0 / 6.0

Sample 03

Reference 6.0

Previous · Gemini 3 Flash
6.0 / 6.0 / 6.0
Gemini 3 Flash
6.5 / 6.5 / 6.5
Gemini 3.5 Flash
6.5 / 6.5 / 6.5
Gemini 3.1 Pro Preview
6.5 / 6.0 / 6.5
Gemini 3.1 Flash-Lite
6.0 / 6.0 / 6.0

Sample 04

Reference 7.0

Previous · Gemini 3 Flash
6.5 / 6.5 / 7.0
Gemini 3 Flash
7.0 / 7.0 / 7.0
Gemini 3.5 Flash
7.0 / 7.0 / 7.0
Gemini 3.1 Pro Preview
7.0 / 7.0 / 6.0
Gemini 3.1 Flash-Lite
6.0 / 6.0 / 6.0

Sample 05

Reference 7.0

Previous · Gemini 3 Flash
6.5 / 7.0 / 6.0
Gemini 3 Flash
7.0 / 6.5 / 7.0
Gemini 3.5 Flash
7.0 / 7.0 / 7.0
Gemini 3.1 Pro Preview
6.5 / 6.5 / 7.0
Gemini 3.1 Flash-Lite
6.0 / 7.0 / 7.0

Sample 06

Reference 7.5

Previous · Gemini 3 Flash
7.0 / 7.0 / 7.0
Gemini 3 Flash
6.0 / 7.0 / 7.0
Gemini 3.5 Flash
7.0 / 7.0 / 7.0
Gemini 3.1 Pro Preview
6.5 / 6.5 / 6.5
Gemini 3.1 Flash-Lite
6.0 / 6.0 / 6.5

Sample 07

Reference 8.0

Previous · Gemini 3 Flash
7.0 / 7.0 / 7.0
Gemini 3 Flash
8.0 / 8.0 / 8.0
Gemini 3.5 Flash
8.0 / 8.0 / 8.0
Gemini 3.1 Pro Preview
7.5 / 8.0 / 7.5
Gemini 3.1 Flash-Lite
7.0 / 7.0 / 7.0

Sample 08

Reference 8.0

Previous · Gemini 3 Flash
6.0 / 6.5 / 6.0
Gemini 3 Flash
7.0 / 7.0 / 7.0
Gemini 3.5 Flash
7.0 / 7.0 / 7.0
Gemini 3.1 Pro Preview
7.0 / 6.5 / 6.5
Gemini 3.1 Flash-Lite
6.0 / 6.5 / 6.5

Sample 09

Reference 8.5

Previous · Gemini 3 Flash
7.0 / 7.0 / 7.5
Gemini 3 Flash
8.0 / 8.0 / 8.0
Gemini 3.5 Flash
8.0 / 8.0 / 8.0
Gemini 3.1 Pro Preview
8.0 / 8.0 / 8.0
Gemini 3.1 Flash-Lite
7.0 / 7.0 / 7.0

Sample 10

Reference 9.0

Previous · Gemini 3 Flash
6.5 / 7.0 / 7.0
Gemini 3 Flash
8.0 / 8.0 / 8.0
Gemini 3.5 Flash
8.0 / 8.0 / 8.0
Gemini 3.1 Pro Preview
7.5 / 7.5 / 7.5
Gemini 3.1 Flash-Lite
7.0 / 7.0 / 6.5

These anonymous comparisons use the published candidate bands only as reference points. This is a Joe Speaking development comparison—not an official IELTS validation or official score. The same samples informed scorer development, and three higher-band samples remained more than 0.5 band low.

03Model evaluation

Why Gemini 3 Flash is the default model

Four models, three runs per sample. Gemini 3 Flash had the smallest average difference.

Gemini 3 Flash

Default
Average difference0.45
Within 0.57/10
Est. credits≈ 13

Default

Gemini 3.5 Flash

Average difference0.48
Within 0.56/10
Est. credits≈ 36

Similar result, more credits

Gemini 3.1 Pro Preview

Average difference0.65
Within 0.57/10
Est. credits≈ 45

More credits, no gain here

Gemini 3.1 Flash-Lite

Average difference0.95
Within 0.54/10
Est. credits≈ 5

Lowest credits, largest difference

Gemini 3 Flash was closest and used fewer estimated credits than Flash 3.5 or Pro. Results vary by response.

The complete three-run results are in the evaluation matrix above. Credit figures are estimates from completed Joe Speaking feedback records, not fixed prices; rankings may change with a larger or fresh dataset.

04Evaluation prompt

The prompt used for this evaluation

Switch between Parts 1–3. V33 was tested above; it is not the current production prompt.

We publish it so learners and teachers can inspect the rules, question the results, and help us improve.

Evaluated prompt
V33
Updated
July 17, 2026
Template
feedback-combined-plus-ielts-recording-v33-v2

Scoring profile

saved-recording-dedicated-language-calibration-v33-duration-aware-2026-07-17

You are scoring one saved IELTS Speaking recording from an ASR transcript.

IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}

Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}

Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}

## 2. IELTS Test-Based Feedback

Topic/Prompt: {{PART_1_TOPIC}}

### Evaluation Method
Use GRADUATED BINARY CHECKS for each criterion. Each check corresponds to a band threshold.
Count "yes" answers to determine the band. Checks are ordered from basic (Band 5) to advanced (Band 9).

**SPOKEN ENGLISH CONTEXT:**
This is an ASR (Automatic Speech Recognition) transcript of spoken English. Apply these tolerance guidelines:
- Pass checks if the core meaning is clear, even with minor grammar slips
- Natural speech patterns (repetitions for emphasis, filler words, self-corrections) are acceptable
- Do NOT penalize punctuation, capitalization, or formatting (ASR artifacts)
- Only fail a check if the issue genuinely impedes communication

**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; there is no minimum word, sentence, example, linker, or discourse-marker requirement. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.

**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate legacy transcript-based Pronunciation rubric exactly as written below.
- Use the dominant demonstrated language pattern across the full supplied recording. At Bands 8-9, one isolated slip, one brief answer, or one strong phrase is not a veto or an automatic pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the candidate's surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Apply the published descriptor allowance for occasional inaccuracies at Band 8 and rare slips at Band 9; do not require a literally perfect transcript.

### Official IELTS Speaking Criteria

---

**Criterion 1: Fluency and Coherence**
*Evaluates: Flow of speech, logical organization, and use of cohesive devices*

Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the speaker maintain flow of speech, even if using repetition or slow speech?
   - YES: Keeps going despite some repetition, self-correction, or hesitation; simple speech is fluent
   - NO: Frequent breakdowns in speech; unable to maintain flow

2. WILLING_TO_SPEAK (Band 6+): Is the speaker willing to speak at length with some coherence?
   - YES: Provides enough connected development for the question, though coherence may weaken at times and linking may be mechanical
   - NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed

3. SPEAKS_AT_LENGTH (Band 7+): Does the speaker speak at length without noticeable effort?
   - YES: Develops relevant ideas coherently and uses cohesive features with some flexibility across the supplied response
   - NO: Development remains noticeably limited, repetitive, or mechanically linked across the supplied response

4. FLUENT_SPEECH (Band 8+): Across the supplied answers, is the dominant Fluency and Coherence pattern closest to Band 8?
   - YES: Topics are coherent, relevant, and generally well developed; cohesive features are wide and flexible; hesitation is usually content-related
   - NO: Recurring language-search hesitation, repetition, or coherence limitations make Band 7 the closer overall fit
   - CALIBRATION: Occasional repetition, self-correction, or content-related hesitation remains compatible with Band 8

5. FULL_FLUENCY (Band 9): Across the supplied answers, is repetition or self-correction rare, hesitation used only to prepare content, and cohesion fully appropriate?
   - YES: The dominant pattern shows fully appropriate development and cohesion, with only rare non-systematic repair
   - NO: Repetition, self-correction, language-search hesitation, or imprecise cohesion is more than rare

Band Mapping:
- 5 yes = Band 9 (speaks fluently with full coherence)
- 4 yes = Band 8 (speaks fluently with rare hesitation)
- 3 yes = Band 7 (speaks at length without noticeable effort)
- 2 yes = Band 6 (willing to speak at length but some coherence loss)
- 1 yes = Band 5 (maintains flow with effort)
- 0 yes = Band 4 or below

---

**Criterion 2: Lexical Resource**
*Evaluates: Range of vocabulary, precision, and use of idiomatic language*

Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the speaker have sufficient vocabulary for familiar topics?
   - YES: Basic vocabulary allows discussion of the topic; may be repetitive but communicates meaning
   - NO: Vocabulary too limited; cannot express ideas clearly

2. TOPIC_VOCAB (Band 6+): Does the speaker use vocabulary adequate for the topic with some variety?
   - YES: Uses appropriate vocabulary with some less common items; attempts paraphrasing
   - NO: Limited to very basic vocabulary; no variety or failed paraphrasing

3. FLEXIBLE_VOCAB (Band 7+): Does the speaker use vocabulary flexibly with awareness of style and collocation?
   - YES: Uses less common words/idioms appropriately; can paraphrase successfully; shows awareness of collocation
   - NO: Limited flexibility; paraphrasing attempts unsuccessful; inappropriate word combinations

4. WIDE_RANGE (Band 8+): Across the supplied topics, is the dominant Lexical Resource pattern closest to Band 8?
   - YES: Uses a wide resource readily and flexibly to convey precise meaning, with skillful less common or idiomatic use
   - NO: Range, flexibility, or precision remains sufficiently limited that Band 7 is the closer overall fit
   - CALIBRATION: Occasional inaccuracies remain compatible with Band 8; every answer need not contain an idiom

5. FULL_FLEXIBILITY (Band 9): Across all supplied topics, does the candidate use vocabulary with full flexibility and precision?
   - YES: Choice is consistently natural, accurate, and precise, with only rare slips compatible with otherwise full control
   - NO: Recurring imprecision or inappropriacy prevents full flexible control across the topics

Band Mapping:
- 5 yes = Band 9 (full flexibility and accuracy)
- 4 yes = Band 8 (wide range with skillful use)
- 3 yes = Band 7 (flexible use with less common items)
- 2 yes = Band 6 (adequate vocabulary for topic)
- 1 yes = Band 5 (sufficient for familiar topics only)
- 0 yes = Band 4 or below

---

**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: Clarity and intelligibility as inferred from transcript*

⚠️ IMPORTANT - Transcript Limitations:
**CANNOT be assessed from transcript:** Actual pronunciation sounds, intonation patterns, word stress, rhythm, accent
**CAN be assessed from transcript:** Word clarity, consistency of forms, potential transcription errors suggesting pronunciation issues

Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Can the speaker be generally understood despite some unclear words?
   - YES: Most words are recognizable and meaning is generally clear from context
   - NO: Many words unclear or unrecognizable; meaning difficult to follow

2. WORD_CLARITY (Band 7+): Are words transcribed clearly without confusion patterns or ambiguity?
   - YES: Words are clear; no systematic confusion patterns (e.g., live/leave, think/thing); no obvious transcription errors
   - NO: Multiple instances of word confusion, unclear transcription, or ambiguous phrases

3. FULL_CLARITY (Band 8+): Is the transcript fully clear with no ambiguous words and would be easily understood?
   - YES: All words transcribed clearly; no ambiguity; response would be understood without any clarification needed
   - NO: Some words or phrases remain ambiguous or potentially mis-transcribed

**Note for feedback:** "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This Pronunciation score reflects only word clarity, transcription consistency, and intelligibility."

**Scoring Rationale:**
Pronunciation can reach Band 8-9 based on transcript clarity alone, following CELPIP's Listenability approach.
While actual pronunciation sounds cannot be assessed from text, perfect transcript clarity (no ambiguous words,
no transcription errors) indicates highly intelligible speech. A speaker whose words are transcribed with
complete accuracy demonstrates the clarity component of pronunciation worthy of high bands.

Band Mapping:
- 3 yes = Band 8 (default) or Band 9 (if exceptional clarity - see below)
- 2 yes = Band 7
- 1 yes = Band 5 or Band 6 (use 6 if close to WORD_CLARITY threshold)
- 0 yes = Band 4 or below

**Band 8 vs 9 Distinction (when all 3 checks pass):**
- Band 8: Transcript is clear with no ambiguous words
- Band 9: Transcript shows exceptional clarity AND includes sophisticated/technical vocabulary transcribed accurately (e.g., idioms, academic terms, proper nouns all captured correctly)

---

**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: Range of structures and grammatical accuracy*

Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the speaker produce basic sentence forms with reasonable accuracy?
   - YES: Simple sentences are mostly correct; attempts complex structures but with errors
   - NO: Frequent errors even in basic structures; causes comprehension problems

2. MIXED_STRUCTURES (Band 6+): Does the speaker use a mix of simple and complex structures?
   - YES: Attempts complex structures (conditionals, relative clauses); errors rarely cause comprehension problems
   - NO: Only simple structures or complex attempts cause confusion

3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the speaker use a range of complex structures with some flexibility?
   - YES: Frequently produces error-free sentences; uses variety of complex structures; some errors persist but don't impede
   - NO: Complex structures have frequent errors; limited flexibility

4. WIDE_RANGE (Band 8+): Across the supplied answers, is the dominant Grammatical Range and Accuracy pattern closest to Band 8?
   - YES: Uses a wide range flexibly; The majority of sentences are error-free; errors are occasional and non-systematic
   - NO: Range is not wide or errors recur sufficiently that Band 7 is the closer overall fit
   - CALIBRATION: Occasional inappropriate or non-systematic forms remain compatible with Band 8

5. FULL_RANGE (Band 9): Across the supplied answers, is the grammatical range full, natural, flexible, and consistently accurate?
   - YES: Full control is sustained, with only rare slips compatible with otherwise natural, precise use
   - NO: Recurring inappropriacies, basic errors, or range limitations prevent full natural control

Band Mapping:
- 5 yes = Band 9 (full range with consistent accuracy)
- 4 yes = Band 8 (wide range with majority error-free)
- 3 yes = Band 7 (range of complex structures with flexibility)
- 2 yes = Band 6 (mix of simple and complex)
- 1 yes = Band 5 (basic structures with reasonable accuracy)
- 0 yes = Band 4 or below

---

### Band 8 vs Band 9 Distinction

When a criterion scores at the 8+ threshold, apply these criteria to distinguish:

**Award Band 9 when:**
- Fluency: Only rare repetition; any hesitation is content-related; topics developed fully and appropriately
- Lexical: Full flexibility and precision; wide range used accurately and effortlessly throughout
- Grammar: Full range used naturally; only native-like 'slips'; structures always appropriate
- Pronunciation: All words fully clear in transcript; no ambiguity whatsoever

**Award Band 8 when:**
- Fluency: Occasional self-correction; hesitation usually content-related but not always
- Lexical: Wide range but occasional less precise choices; rare minor errors
- Grammar: Majority error-free but occasional minor errors or inappropriacies
- Pronunciation: Generally clear with minor ambiguities in a few words

**Default to Band 8 when:** Criteria met but without the consistent excellence markers of Band 9.

---

### Overall Band Calculation

**Formula:** Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4

**IELTS Speaking Rounding Rule:**
Speaking scores always ROUND DOWN (floor) to the nearest 0.5.

**Examples:**
- 6.125 → 6.0 (floor to .0)
- 6.25 → 6.0 (floor to .0)
- 6.5 → 6.5 (already at .5)
- 6.75 → 6.5 (floor to .5)
- 6.875 → 6.5 (floor to .5)

**Calculation Examples:**
- (7 + 7 + 6 + 6) ÷ 4 = 6.5 → **Band 6.5**
- (8 + 7 + 7 + 8) ÷ 4 = 7.5 → **Band 7.5**
- (9 + 8 + 8 + 8) ÷ 4 = 8.25 → **Band 8.0**
- (8 + 7 + 6 + 6) ÷ 4 = 6.75 → **Band 6.5**

**Important:** Criterion scores (FC, LR, GRA, P) must be whole numbers (5, 6, 7, 8, 9). Only the overall score uses 0.5 increments.

---

### Key Observations
Provide 3-4 specific observations tied DIRECTLY to check results, citing evidence from the transcript.
Examples:
- "Fluency limited by repetition of 'I think' (5 times) - WILLING_TO_SPEAK check: yes but close to threshold"
- "Strong vocabulary with idioms like 'take into account', 'on the other hand' - FLEXIBLE_VOCAB check: yes"
- "6 of 8 sentences use 'I + verb' pattern. Missing: compound openers with 'although/while'. Try: 'Although the weather was hot, I still enjoyed the walk' — MIXED_STRUCTURES check: no"

**Grammar & Structure Observations (MIXED_STRUCTURES, RANGE_WITH_FLEXIBILITY, WIDE_RANGE):**

When a grammar or structure check FAILS or is borderline, the observation MUST:
- NAME the specific sentence structures present AND absent (e.g., "simple conditionals present, no inversions or cleft sentences")
- QUOTE the actual sentence pattern from the transcript (e.g., "'I went to the park. I saw the lake. I liked it.' — all simple Subject-Verb-Object")
- SUGGEST one level-appropriate alternative structure using the learner's own content (e.g., "'Having visited the park, I was struck by the beautiful lake' — starting with a participle phrase")

BAD observation: "Limited sentence structures — MIXED_STRUCTURES check: no"
GOOD observation: "6 of 8 sentences use 'I + verb' pattern ('I went...', 'I saw...', 'I liked...'). Missing: compound sentences with 'although/while'. Try: 'Although the weather was hot, I still enjoyed the walk' — MIXED_STRUCTURES check: no"

LEVEL-APPROPRIATE GUIDANCE:
- Only diagnose structure monotony when there is enough material to show a pattern (multiple clauses or repeated openers across 50+ words — do NOT rely on sentence count, as ASR transcripts are often lightly punctuated or run-on)
- Suggest the NEXT step up from the learner's current level, not advanced rhetorical forms
- For lower levels (Band 5-6 / CELPIP L4-7): suggest basic compound sentences, simple conditionals ("If I had time, I would..."), basic relative clauses ("The place where I live...")
- For mid levels (Band 6-7 / CELPIP L8-9): suggest fronted adverbials ("Having considered this, ..."), concessive clauses ("Although it can be challenging, ..."), cleft sentences ("What I enjoy most is...")
- For higher levels (Band 7+ / CELPIP L10+): suggest inversions ("Not only did I..."), participle phrases ("Fascinated by the idea, I..."), mixed conditionals
- Always prefer spoken-natural alternatives over academic/written forms
- Use plain-English names for structures (e.g., "starting with a time phrase" not just "fronted adverbial") alongside the grammar term
- Do NOT invent structures that aren't supported by the transcript evidence

---

### JSON Output Format
{
  "testBased": {
    "test": "IELTS",
    "overall": <number 5.0-9.0 in 0.5 increments>,
    "criteria": [
      {
        "name": "Fluency and Coherence",
        "checks": [
          {"criterion": "MAINTAINS_FLOW", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "WILLING_TO_SPEAK", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "SPEAKS_AT_LENGTH", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FLUENT_SPEECH", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FULL_FLUENCY", "result": true/false, "evidence": "<specific observation from transcript>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      },
      {
        "name": "Lexical Resource",
        "checks": [
          {"criterion": "ADEQUATE_VOCAB", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "TOPIC_VOCAB", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FLEXIBLE_VOCAB", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "WIDE_RANGE", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FULL_FLEXIBILITY", "result": true/false, "evidence": "<specific observation from transcript>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      },
      {
        "name": "Pronunciation (Transcript-Based)",
        "checks": [
          {"criterion": "INTELLIGIBILITY", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "WORD_CLARITY", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FULL_CLARITY", "result": true/false, "evidence": "<specific observation from transcript>"}
        ],
        "yesCount": <0-3>,
        "band": <5-9>,
        "note": "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."
      },
      {
        "name": "Grammatical Range and Accuracy",
        "checks": [
          {"criterion": "BASIC_STRUCTURES", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "MIXED_STRUCTURES", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "RANGE_WITH_FLEXIBILITY", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "WIDE_RANGE", "result": true/false, "evidence": "<specific observation from transcript>"},
          {"criterion": "FULL_RANGE", "result": true/false, "evidence": "<specific observation from transcript>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      }
    ],
    "keyObservations": [
      "Specific observation with transcript evidence and check reference",
      "Another observation explaining score with examples",
      "Third observation tied to check results"
    ]
  }
}

Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.

Review or customize your scoring instructions

You can review or customize scoring instructions in Settings. Custom prompts were not included in the V33 evaluation above.

Open Settings
05What the score cannot measure

This score cannot confirm your actual pronunciation

Recording-based feedback scores the transcript and question context. Because the model does not receive your recording audio, it cannot confirm how you pronounce words.

The scorer only reads text

It cannot assess sounds, stress, rhythm, intonation, connected speech, or accent.

Two practice modes

Choose the mode that fits your goal

Recording Practice is designed for repeatable answers and detailed text-based feedback. Live Conversation listens and responds to your voice in real time, simulating the back-and-forth of a real speaking test.

Try Live Conversation

Transcription errors can change the score

Transcription errors can distort vocabulary, grammar, and meaning. Listen back and correct the text before requesting feedback.

The development gap is not closed

Three higher-band samples remained more than 0.5 band low. A single score can vary, so use the average across several comparable attempts as a more stable estimate. More attempts usually reduce the influence of one unusual result.

06What we do next

More data and better feedback

We will expand the evaluation set when reliable samples are available and learn from scores users dispute.

01

Expand the dataset

Add reliable samples across IELTS parts, topics, and bands.

02

Learn from disputed scores

Review the transcript, model, prompt version, duration, estimated band, and the band the learner expected.

Your feedback is welcome

Tell us what score you expected and why. Include the IELTS part, transcript, model, prompt version, and duration when relevant.

Share scoring feedback

Joe Speaking is an independent practice product and is not affiliated with or endorsed by IELTS. Scores shown here are AI practice estimates, not official test results.