Skip to main content
Inside Recording Practice scoring

How Joe Speakingimproves scoring accuracy

Using Recording Practice, we show how Joe Speaking generates a band score—and how we test and improve its accuracy.

See the scoring process

Recording Practice score

An AI practice estimate, not an official result.

Recording Practice evidence checks on a phoneRecording Practice band score on a phone

One method, test-specific rules

IELTS and CELPIP both use explainable yes-or-no checks, but their criteria and score ranges differ. The diagram, prompt, and public comparison below are an IELTS Recording Practice worked example.

01How the score is generated

AI does not choose your band directly

AI completes yes-or-no checks for each IELTS criterion. The results produce four criterion bands and an overall practice score. Because Recording Practice sends text—not audio—Pronunciation only reflects transcript clarity.

How Recording Practice builds an AI practice bandJoe Speaking runs graduated yes-or-no checks for Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. The yes count maps to a band or narrow band range; because the scorer receives text rather than audio, Pronunciation can only use transcript clarity.1READ THE ANSWERAnswer transcript + questions2RUN YES / NO CHECKSExample: VocabularyAdequate vocabularyYESTopic vocabularyYESFlexible useYESWide rangeYESFull flexibilityNO4 YES → Vocabulary band 83GET FOUR SCORESFluency andCoherence9Lexical Resource8Grammatical Rangeand Accuracy8Pronunciation7YES count sets the band or range4GET THE OVERALL SCORE(9 + 8 + 8 + 7) ÷ 4Average 8.00OVERALL BAND8.0

Pronunciation is a text-based proxy

The scorer cannot hear sounds, stress, rhythm, or intonation. It can only use the clarity of the transcript.

02Evaluation results

How we test scoring accuracy

We use 10 anonymized text samples adapted by Joe Speaking from public candidate examples. We score every sample three times, average the results, and compare that average with its reference band. A smaller difference means a closer result.

We reran the same 10 samples three times at each supported effort level. The complete scores and the ranked model comparison are below.

On August 14, 2026, Gemini 3.7 Flash High, Flash 3 Low, and Flash 3 Medium were rerun together on Vertex AI. All 90 runs returned a valid score; other columns retain the latest published V2 runs.

Reference source

Public candidate performance examples used only as a reference. View the public reference source

Evaluation input

Ten anonymized Part 3 text samples adapted by Joe Speaking, with matching question context

Each result

Three runs per effort level, then one average

How the comparison works

How the reference comparison worksOne sample, prompt and model produce three runs. Their average is compared with the published reference band. The gap is reviewed before the prompt is revised and evaluated again.ONE SETUPPart 3 sample + prompt + modelSame context each runTHREE RUNSRUN 15.0RUN 25.5RUN 35.5THREE-RUN AVERAGE5.33REFERENCE BAND 5.0Gap 0.33Review → revise → evaluate againThe full result table stays visible below
V2 · August 14, 2026330 model runs

Complete V2 evaluation matrix

Choose a model to see every score: 10 samples, up to four effort levels, and three runs per effort.

Sample 01

Reference 5.0

Minimal
— / — / —
Low
5.5 / 5.0 / 5.5
Medium
6.0 / 6.0 / 6.0
High
6.0 / 6.0 / 6.0

Sample 02

Reference 6.0

Minimal
— / — / —
Low
5.5 / 5.5 / 6.5
Medium
6.5 / 6.5 / 6.5
High
6.5 / 6.5 / 6.5

Sample 03

Reference 6.0

Minimal
— / — / —
Low
6.5 / 6.5 / 6.5
Medium
6.5 / 6.5 / 7.0
High
6.5 / 6.5 / 6.0

Sample 04

Reference 7.0

Minimal
— / — / —
Low
6.5 / 6.5 / 6.5
Medium
7.0 / 7.0 / 6.5
High
7.0 / 7.0 / 6.5

Sample 05

Reference 7.0

Minimal
— / — / —
Low
6.5 / 6.5 / 6.5
Medium
7.0 / 7.0 / 7.0
High
7.0 / 6.5 / 7.0

Sample 06

Reference 7.5

Minimal
— / — / —
Low
6.5 / 6.5 / 6.5
Medium
7.0 / 7.0 / 7.0
High
7.0 / 7.0 / 7.0

Sample 07

Reference 8.0

Minimal
— / — / —
Low
7.5 / 8.0 / 8.0
Medium
8.0 / 8.0 / 8.0
High
8.0 / 8.0 / 7.5

Sample 08

Reference 8.0

Minimal
— / — / —
Low
6.5 / 6.5 / 6.5
Medium
7.0 / 7.0 / 7.0
High
7.0 / 7.0 / 7.0

Sample 09

Reference 8.5

Minimal
— / — / —
Low
8.0 / 8.0 / 8.0
Medium
8.0 / 8.0 / 8.0
High
8.0 / 8.0 / 8.0

Sample 10

Reference 9.0

Minimal
— / — / —
Low
8.0 / 8.0 / 8.0
Medium
8.0 / 8.0 / 8.0
High
8.0 / 8.0 / 8.0
03V2 model evaluation

Why we recommend Gemini 3.7 Flash Medium

Gemini 3.7 Flash Medium cut the average score gap from 0.73 to 0.53 versus Flash 3 Low and was about 15 seconds faster than Flash 3 Medium. It matched Gemini 3.7 High on average score gap (0.53) while using less time and fewer credits.

Gemini 3 Flash Preview · Medium

Average difference0.47
Samples within 0.56/10
Est. credits≈ 13

Closest score, slowest

Gemini 3.7 Flash · Medium

Recommended
Average difference0.53
Samples within 0.56/10
Est. credits≈ 13

Recommended

Gemini 3 Flash Preview · Low

Average difference0.73
Samples within 0.56/10
Est. credits≈ 9

Lower cost, lower accuracy

Estimated credits include the current 2× rate: Flash 3 Low ≈9, Flash 3 Medium ≈13, and Gemini 3.7 Medium ≈13. A complete feedback request may cost more; actual charges vary with input, output, and reasoning length.

03Evaluation prompt

The prompt used for this evaluation

Switch between Parts 1–3. V33 was tested above; it is not the current production prompt.

We publish it so learners and teachers can inspect the rules, question the results, and help us improve.

Evaluated prompt
V33
Updated
July 17, 2026
Template
feedback-combined-plus-ielts-recording-v33-v6

Scoring profile

saved-recording-dedicated-language-calibration-v33-output-policy-2026-07-22

You are scoring one saved IELTS Speaking recording from an ASR transcript.

IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}

Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}

Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}

## 2. IELTS Test-Based Feedback

Topic/Prompt: {{PART_1_TOPIC}}

### Evaluation Method
Evaluate the fixed graduated checks below in their listed order. Assess only patterns visible in the supplied transcript. The model sets each criterion band from the fixed check sequence; the parser validates that band.

**TRANSCRIPT CONTEXT:**
This is an ASR transcript of spoken English. Apply these tolerance guidelines:
- Judge meaning, organization, and language patterns from the supplied text; tolerate minor grammar slips and ASR artifacts.
- Treat textual repetition, filler tokens, and self-corrections as transcript patterns; do not infer sound or delivery qualities.
- Do not penalize punctuation, capitalization, or formatting.
- Only fail a check when the transcript text shows that its criterion is not met.

**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; do not impose a quota for detail or linking. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.

**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate transcript-visible word-form consistency rubric for Pronunciation below; do not infer sound or delivery from text.
- Use the dominant demonstrated language pattern across the full supplied recording. An isolated slip, brief answer, or strong phrase is not an automatic veto or pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Allow isolated non-systematic inaccuracies when the broader transcript pattern remains controlled; do not require a literally perfect transcript.

### Official IELTS Speaking Criteria

---

**Criterion 1: Fluency and Coherence**
*Evaluates: connected development, logical organization, and cohesive links visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the transcript show connected progression of ideas?
   - YES: Ideas continue with understandable links, even when wording repeats or repairs appear in the text
   - NO: The text repeatedly breaks into disconnected fragments, making meaning hard to follow

2. WILLING_TO_SPEAK (Band 6+): Does the response provide enough connected development for its question?
   - YES: The response supplies relevant development; linking may be mechanical
   - NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed

3. SPEAKS_AT_LENGTH (Band 7+): Does the supplied response sustain relevant development across its content?
   - YES: Relevant ideas are developed coherently with flexible cohesive links
   - NO: Development remains limited, repetitive, or mechanically linked

4. FLUENT_SPEECH (Band 8+): Is the dominant transcript pattern coherent, relevant, well developed, and flexibly linked?
   - YES: The response maintains clear organization and flexible cohesion across its demonstrated content
   - NO: Recurring repetition or coherence limitations make the higher threshold unsupported

5. FULL_FLUENCY (Band 9): Does the transcript show fully appropriate development and cohesion without a recurring repair pattern?
   - YES: Development and cohesion remain fully appropriate throughout the demonstrated content
   - NO: Repetition, self-correction, or imprecise cohesion recurs in the transcript

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 2: Lexical Resource**
*Evaluates: range, precision, and appropriacy of word choices visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the response show enough vocabulary for familiar topics?
   - YES: Word choices support discussion of the topic and communicate the intended meaning
   - NO: Limited word choices repeatedly prevent clear expression

2. TOPIC_VOCAB (Band 6+): Does the response use topic-appropriate vocabulary with some variety?
   - YES: Word choices fit the topic, include some less-common items, and show attempts at paraphrasing
   - NO: Word choices remain very basic, show little variety, or repeatedly fail to paraphrase the intended meaning

3. FLEXIBLE_VOCAB (Band 7+): Does the response use vocabulary flexibly and appropriately across its ideas?
   - YES: Word choices include less-common or idiomatic resources, support successful paraphrasing, and show collocational awareness
   - NO: Limited flexibility, unsuccessful paraphrasing, imprecise choices, or unsuitable combinations recur

4. WIDE_RANGE (Band 8+): Is the dominant lexical pattern broad in range, flexible, and precise?
   - YES: The response uses a broad range with flexibility and precision, including less-common or idiomatic resources when appropriate
   - NO: Range, flexibility, or precision remains limited at the higher threshold

5. FULL_FLEXIBILITY (Band 9): Does the response show full flexibility and precision in word choice?
   - YES: Word choices remain consistently natural, accurate, and precise across the demonstrated content, with full flexibility in paraphrasing and collocation
   - NO: Recurring imprecision or inappropriacy prevents full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: transcript-visible word-form consistency and clarity only*

**Transcript boundary:** Do not infer sound, delivery, or listener response from text. Assess only recognizable word forms, consistency, and ambiguity visible in the transcript.

Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Does the transcript's recognizable wording support generally clear meaning?
   - YES: Word forms in the transcript support clear meaning in context
   - NO: Unclear or ambiguous forms repeatedly obscure meaning

2. WORD_CLARITY (Band 7+): Are word forms transcribed consistently without confusion patterns?
   - YES: Word forms remain consistent and unambiguous across the transcript
   - NO: Confusion, ambiguity, or inconsistent forms recur

3. FULL_CLARITY (Band 8+): Is the transcript consistently clear without unresolved ambiguity?
   - YES: The transcript's word forms remain clear and unambiguous throughout its demonstrated content
   - NO: Some forms or phrases remain unresolved or potentially mis-transcribed

**Boundary note:** Omit the `note` key from the raw response. After strict validation, code attaches this exact display note: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

**Scoring Rationale:**
Use transcript-visible word-form consistency only. Do not infer sound, delivery, or listener effort from the transcript.

Band Mapping (derived from the fixed three checks; do not restate this mapping in authored text):
- 3 yes = Band 8 by default; use Band 9 only when the transcript shows exceptional consistency
- 2 yes = Band 7
- With exactly one passed check, set Band 6 only when WORD_CLARITY is narrowly missed; otherwise set Band 5
- 0 yes = Band 4 or below

**Band 8 vs 9 Distinction (when all three checks pass):**
- Band 8: The transcript is clear and consistent
- Band 9: The transcript shows exceptional clarity and consistency, including complex or technical terminology when present

---

**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: range of structures and grammatical control visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the response show basic forms with reasonable accuracy?
   - YES: Basic forms are generally controlled and meaning remains clear
   - NO: Basic-form errors repeatedly interfere with meaning

2. MIXED_STRUCTURES (Band 6+): Does the response use a mix of simple and complex forms?
   - YES: The transcript demonstrates more than one structural pattern with generally clear meaning
   - NO: The transcript remains restricted to simple forms or complex attempts repeatedly cause confusion

3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the response use a range of complex forms with some flexibility?
   - YES: Complex forms vary and are generally controlled even when occasional errors remain
   - NO: Complex forms show recurring errors or limited flexibility

4. WIDE_RANGE (Band 8+): Is the dominant grammatical pattern broad, flexible, and generally controlled?
   - YES: Structural choices vary flexibly and errors are occasional and non-systematic
   - NO: Range is not broad or errors recur enough to miss the higher threshold

5. FULL_RANGE (Band 9): Does the response sustain full, natural, flexible, and accurate structural control?
   - YES: Structural control remains natural and accurate across the demonstrated content
   - NO: Recurring inappropriacies, basic errors, or range limitations prevent full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

### Band 8 vs Band 9 Distinction

When a criterion reaches an upper threshold, use the dominant transcript-visible pattern rather than an isolated token or phrase.

**Award Band 9 when:**
- Fluency and Coherence: development and cohesion are fully appropriate without a recurring repair pattern
- Lexical Resource: word choice is fully flexible, precise, and consistently appropriate
- Grammatical Range and Accuracy: structural control is full, natural, flexible, and accurate
- Transcript criterion: word forms remain exceptionally clear and consistent

**Award Band 8 when:**
- Fluency and Coherence: development and cohesion are broad and flexible with limited non-systematic variation
- Lexical Resource: range is broad with occasional imprecision
- Grammatical Range and Accuracy: range is broad with occasional non-systematic errors
- Transcript criterion: word forms are clear with limited ambiguity

**Default to Band 8 when:** the criterion meets its upper threshold without the consistent pattern required for Band 9.

---

### Overall Band Calculation

Use the four criterion band fields to calculate the numeric overall field. Formula: Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4. Speaking scores always ROUND DOWN (floor) to the nearest 0.5. Return only the numeric overall value; do not restate formulas, band labels, or score prose in authored evidence or observations.

---

### Key Observations
The raw model response must set keyObservations to exactly []. The combined call owns broader observations; do not author or copy transcript observations in this response.

---

**Strict V33 output shape (hard requirement):**
- Return exactly one JSON root object with only the testBased key.
- The testBased object contains only test, overall, criteria, and keyObservations. Set the numeric structural values from the rubric.
- The criteria array contains exactly four criteria in this order: Fluency and Coherence, Lexical Resource, Pronunciation (Transcript-Based), and Grammatical Range and Accuracy.
- Each criterion has exactly its fixed checks in order. The check counts are 5, 5, 3, and 5. The Pronunciation (Transcript-Based) criterion has exactly three checks: INTELLIGIBILITY, WORD_CLARITY, and FULL_CLARITY. Do not add or omit checks or criteria.
- For every check, result true pairs with evidence "pass" and result false pairs with evidence "fail". These are the only accepted evidence values.
- False branch example: {"result": false, "evidence": "fail"}.
- keyObservations must be exactly [] and no criterion includes a note key in this raw response.

### JSON Output Format
{
  "testBased": {
    "test": "IELTS",
    "overall": 8.5,
    "criteria": [
      {
        "name": "Fluency and Coherence",
        "checks": [
          {"criterion": "MAINTAINS_FLOW", "result": true, "evidence": "pass"},
          {"criterion": "WILLING_TO_SPEAK", "result": true, "evidence": "pass"},
          {"criterion": "SPEAKS_AT_LENGTH", "result": true, "evidence": "pass"},
          {"criterion": "FLUENT_SPEECH", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLUENCY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Lexical Resource",
        "checks": [
          {"criterion": "ADEQUATE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "TOPIC_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "FLEXIBLE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLEXIBILITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Pronunciation (Transcript-Based)",
        "checks": [
          {"criterion": "INTELLIGIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WORD_CLARITY", "result": true, "evidence": "pass"},
          {"criterion": "FULL_CLARITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 3,
        "band": 8
      },
      {
        "name": "Grammatical Range and Accuracy",
        "checks": [
          {"criterion": "BASIC_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "MIXED_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "RANGE_WITH_FLEXIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_RANGE", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      }
    ],
    "keyObservations": []
  }
}

**Closed V33 evidence contract (hard requirement):**
- This is controlled decoding, not model-authored display prose. For every check, emit the exact lowercase token "pass" when result is true and the exact lowercase token "fail" when result is false. No other evidence value is accepted.
- Emit exactly four criteria with the fixed 5/5/3/5 check IDs and preserve their result, yesCount, and band values. Do not add fields, checks, or criteria.
- Emit keyObservations as exactly [] because the combined call owns broader observations.
- Omit the note key from every raw criterion object. The parser adds the approved Pronunciation boundary note after strict validation.
- Do not quote, paraphrase, or otherwise author transcript evidence in the raw response. Unknown tokens, token/result mismatches, notes, and non-empty observations are rejected rather than sanitized.
- The code-owned Pronunciation boundary note is exactly: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.

Review or customize your scoring instructions

You can review or customize scoring instructions in Settings. Custom prompts were not included in the V33 evaluation above.

Open Settings
04What the score cannot measure

This score cannot confirm your actual pronunciation

Recording-based feedback scores the transcript and question context. Because the model does not receive your recording audio, it cannot confirm how you pronounce words.

The scorer only reads text

It cannot assess sounds, stress, rhythm, intonation, connected speech, or accent.

Two practice modes

Choose the mode that fits your goal

Recording Practice is designed for repeatable answers and detailed text-based feedback. Live Conversation listens and responds to your voice in real time, simulating the back-and-forth of a real speaking test.

Try Live Conversation

Transcription errors can change the score

Transcription errors can distort vocabulary, grammar, and meaning. Listen back and correct the text before requesting feedback.

The development gap is not closed

Three higher-band samples remained more than 0.5 band low. A single score can vary, so use the average across several comparable attempts as a more stable estimate. More attempts usually reduce the influence of one unusual result.

05What we do next

More data and better feedback

We will expand the evaluation set when reliable samples are available and learn from scores users dispute.

01

Expand the dataset

Add reliable samples across IELTS parts, topics, and bands.

02

Learn from disputed scores

Review the transcript, model, prompt version, duration, estimated band, and the band the learner expected.

Your feedback is welcome

Tell us what score you expected and why. Include the IELTS part, transcript, model, prompt version, and duration when relevant.

Share scoring feedback

Joe Speaking is an independent practice product and is not affiliated with or endorsed by IELTS. Scores shown here are AI practice estimates, not official test results.