跳到主要内容
录音练习评分揭秘

Joe Speaking 如何提高评分准确度

我们以录音练习为例,公开说明 Joe Speaking 如何生成分数,以及如何持续测试和提高准确度。

查看评分过程

录音练习分数

这是 AI 练习估分,不是官方成绩。

手机上的录音练习评分依据手机上的录音练习分数

同一种方法,不同考试规则

IELTS 和 CELPIP 都使用可解释的“是/否”检查,但评分标准和分数范围不同。下方图示、Prompt 与公开参考对比均以 IELTS 录音练习为例。

01分数如何生成

AI 不会直接决定你的分数

AI 会逐项完成“是/否”检查,再根据结果得出四项小分和练习总分。录音练习只发送文字,因此发音项只能参考回答文字是否清晰。

录音练习如何生成 AI 练习分数Joe Speaking 依次进行“是/否”检查,由通过数得出四项小分和练习总分。评分器只收到文字,因此发音项只能参考回答文字是否清晰。1读取回答回答文字 + 题目2逐项 是/否 检查示例:词汇词汇足够话题词汇灵活使用范围广泛完全灵活4 个“是” → 词汇 8 分3得出四项小分流利与连贯9词汇8语法8发音(文字参考)7通过数确定分数或范围4计算总分(9 + 8 + 8 + 7) ÷ 4平均 8.00总分8.0

发音分数只参考文字清晰度

评分器听不到发音、重音、节奏或语调,只能根据回答文字是否清晰来判断。

02评测结果

我们如何测试评分准确度

我们参考公开考生案例,由 Joe Speaking 匿名化整理成 10 份文字样本。每种评分设置对每个样本运行三次,取平均值,再与参考分数比较。差距越小,结果越接近。

我们对同一组 10 个样本、每种支持的推理强度各运行三次。下方是完整评分矩阵与模型评测排名。

2026 年 8 月 14 日,我们在 Vertex AI 上同时重跑了 Gemini 3.7 Flash 高推理强度、Flash 3 低推理强度和 Flash 3 中推理强度;90 次运行均返回有效评分,其余列保留最新发布的 V2 运行结果。

参考来源

公开考生表现案例,仅作为参考。 查看公开参考来源

评测输入

10 份由 Joe Speaking 匿名化整理的 Part 3 文字样本及对应问题

每个结果

每种推理强度运行三次,再取平均

对比测试如何进行

参考分数对比如何进行同一个样本、Prompt 和模型运行三次,平均值与公开参考分数比较,再根据差距复盘和改进。一组设置Part 3 样本 + Prompt + 模型每次上下文相同运行三次 15.0 25.5 35.5三次平均值5.33参考分数 5.0差距 0.33检查 → 修改 → 再次评测下方公开完整结果表
V2 · 2026 年 8 月 14 日330 次模型运行

完整 V2 评测矩阵

选择模型查看每个分数:10 个样本、最多四种推理强度,每种运行三次。

Sample 01

参考分 5.0

最小
— / — / —
5.5 / 5.0 / 5.5
6.0 / 6.0 / 6.0
6.0 / 6.0 / 6.0

Sample 02

参考分 6.0

最小
— / — / —
5.5 / 5.5 / 6.5
6.5 / 6.5 / 6.5
6.5 / 6.5 / 6.5

Sample 03

参考分 6.0

最小
— / — / —
6.5 / 6.5 / 6.5
6.5 / 6.5 / 7.0
6.5 / 6.5 / 6.0

Sample 04

参考分 7.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 6.5
7.0 / 7.0 / 6.5

Sample 05

参考分 7.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 6.5 / 7.0

Sample 06

参考分 7.5

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 7.0 / 7.0

Sample 07

参考分 8.0

最小
— / — / —
7.5 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 7.5

Sample 08

参考分 8.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 7.0 / 7.0

Sample 09

参考分 8.5

最小
— / — / —
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0

Sample 10

参考分 9.0

最小
— / — / —
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
03V2 模型评测

为什么我们推荐 Gemini 3.7 Flash 中推理强度

Gemini 3.7 Flash 中推理强度把平均评分差距从 Flash 3 低推理强度的 0.73 降到 0.53,且比 Flash 3 中推理强度快约 15 秒。它与 Gemini 3.7 高推理强度的平均评分差距相同(0.53),但耗时和积分更少。

Gemini 3 Flash Preview · 中

平均差距0.47
误差 ≤ 0.5 的样本6/10
预估积分≈ 13

差距最小,但最慢

Gemini 3.7 Flash · 中

推荐
平均差距0.53
误差 ≤ 0.5 的样本6/10
预估积分≈ 13

推荐

Gemini 3 Flash Preview · 低

平均差距0.73
误差 ≤ 0.5 的样本6/10
预估积分≈ 9

成本更低,准确度较低

预估积分已包含当前 2 倍费率:Flash 3 低推理强度约 9、Flash 3 中推理强度约 13、Gemini 3.7 中推理强度约 13。完整反馈请求可能更贵;实际消耗会随输入、输出和推理长度变化。

03评测 Prompt

查看本次评测使用的 Prompt

可切换 Part 1、Part 2 和 Part 3。这是上方结果所测试的 V33 评测 Prompt,不是当前生产评分 Prompt。

我们公开 Prompt,让学习者和老师可以检查规则、提出质疑并帮助我们改进。

评测 Prompt
V33
更新日期
2026 年 7 月 17 日
模板
feedback-combined-plus-ielts-recording-v33-v6

评分配置

saved-recording-dedicated-language-calibration-v33-output-policy-2026-07-22

You are scoring one saved IELTS Speaking recording from an ASR transcript.

IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}

Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}

Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}

## 2. IELTS Test-Based Feedback

Topic/Prompt: {{PART_1_TOPIC}}

### Evaluation Method
Evaluate the fixed graduated checks below in their listed order. Assess only patterns visible in the supplied transcript. The model sets each criterion band from the fixed check sequence; the parser validates that band.

**TRANSCRIPT CONTEXT:**
This is an ASR transcript of spoken English. Apply these tolerance guidelines:
- Judge meaning, organization, and language patterns from the supplied text; tolerate minor grammar slips and ASR artifacts.
- Treat textual repetition, filler tokens, and self-corrections as transcript patterns; do not infer sound or delivery qualities.
- Do not penalize punctuation, capitalization, or formatting.
- Only fail a check when the transcript text shows that its criterion is not met.

**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; do not impose a quota for detail or linking. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.

**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate transcript-visible word-form consistency rubric for Pronunciation below; do not infer sound or delivery from text.
- Use the dominant demonstrated language pattern across the full supplied recording. An isolated slip, brief answer, or strong phrase is not an automatic veto or pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Allow isolated non-systematic inaccuracies when the broader transcript pattern remains controlled; do not require a literally perfect transcript.

### Official IELTS Speaking Criteria

---

**Criterion 1: Fluency and Coherence**
*Evaluates: connected development, logical organization, and cohesive links visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the transcript show connected progression of ideas?
   - YES: Ideas continue with understandable links, even when wording repeats or repairs appear in the text
   - NO: The text repeatedly breaks into disconnected fragments, making meaning hard to follow

2. WILLING_TO_SPEAK (Band 6+): Does the response provide enough connected development for its question?
   - YES: The response supplies relevant development; linking may be mechanical
   - NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed

3. SPEAKS_AT_LENGTH (Band 7+): Does the supplied response sustain relevant development across its content?
   - YES: Relevant ideas are developed coherently with flexible cohesive links
   - NO: Development remains limited, repetitive, or mechanically linked

4. FLUENT_SPEECH (Band 8+): Is the dominant transcript pattern coherent, relevant, well developed, and flexibly linked?
   - YES: The response maintains clear organization and flexible cohesion across its demonstrated content
   - NO: Recurring repetition or coherence limitations make the higher threshold unsupported

5. FULL_FLUENCY (Band 9): Does the transcript show fully appropriate development and cohesion without a recurring repair pattern?
   - YES: Development and cohesion remain fully appropriate throughout the demonstrated content
   - NO: Repetition, self-correction, or imprecise cohesion recurs in the transcript

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 2: Lexical Resource**
*Evaluates: range, precision, and appropriacy of word choices visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the response show enough vocabulary for familiar topics?
   - YES: Word choices support discussion of the topic and communicate the intended meaning
   - NO: Limited word choices repeatedly prevent clear expression

2. TOPIC_VOCAB (Band 6+): Does the response use topic-appropriate vocabulary with some variety?
   - YES: Word choices fit the topic, include some less-common items, and show attempts at paraphrasing
   - NO: Word choices remain very basic, show little variety, or repeatedly fail to paraphrase the intended meaning

3. FLEXIBLE_VOCAB (Band 7+): Does the response use vocabulary flexibly and appropriately across its ideas?
   - YES: Word choices include less-common or idiomatic resources, support successful paraphrasing, and show collocational awareness
   - NO: Limited flexibility, unsuccessful paraphrasing, imprecise choices, or unsuitable combinations recur

4. WIDE_RANGE (Band 8+): Is the dominant lexical pattern broad in range, flexible, and precise?
   - YES: The response uses a broad range with flexibility and precision, including less-common or idiomatic resources when appropriate
   - NO: Range, flexibility, or precision remains limited at the higher threshold

5. FULL_FLEXIBILITY (Band 9): Does the response show full flexibility and precision in word choice?
   - YES: Word choices remain consistently natural, accurate, and precise across the demonstrated content, with full flexibility in paraphrasing and collocation
   - NO: Recurring imprecision or inappropriacy prevents full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: transcript-visible word-form consistency and clarity only*

**Transcript boundary:** Do not infer sound, delivery, or listener response from text. Assess only recognizable word forms, consistency, and ambiguity visible in the transcript.

Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Does the transcript's recognizable wording support generally clear meaning?
   - YES: Word forms in the transcript support clear meaning in context
   - NO: Unclear or ambiguous forms repeatedly obscure meaning

2. WORD_CLARITY (Band 7+): Are word forms transcribed consistently without confusion patterns?
   - YES: Word forms remain consistent and unambiguous across the transcript
   - NO: Confusion, ambiguity, or inconsistent forms recur

3. FULL_CLARITY (Band 8+): Is the transcript consistently clear without unresolved ambiguity?
   - YES: The transcript's word forms remain clear and unambiguous throughout its demonstrated content
   - NO: Some forms or phrases remain unresolved or potentially mis-transcribed

**Boundary note:** Omit the `note` key from the raw response. After strict validation, code attaches this exact display note: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

**Scoring Rationale:**
Use transcript-visible word-form consistency only. Do not infer sound, delivery, or listener effort from the transcript.

Band Mapping (derived from the fixed three checks; do not restate this mapping in authored text):
- 3 yes = Band 8 by default; use Band 9 only when the transcript shows exceptional consistency
- 2 yes = Band 7
- With exactly one passed check, set Band 6 only when WORD_CLARITY is narrowly missed; otherwise set Band 5
- 0 yes = Band 4 or below

**Band 8 vs 9 Distinction (when all three checks pass):**
- Band 8: The transcript is clear and consistent
- Band 9: The transcript shows exceptional clarity and consistency, including complex or technical terminology when present

---

**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: range of structures and grammatical control visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the response show basic forms with reasonable accuracy?
   - YES: Basic forms are generally controlled and meaning remains clear
   - NO: Basic-form errors repeatedly interfere with meaning

2. MIXED_STRUCTURES (Band 6+): Does the response use a mix of simple and complex forms?
   - YES: The transcript demonstrates more than one structural pattern with generally clear meaning
   - NO: The transcript remains restricted to simple forms or complex attempts repeatedly cause confusion

3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the response use a range of complex forms with some flexibility?
   - YES: Complex forms vary and are generally controlled even when occasional errors remain
   - NO: Complex forms show recurring errors or limited flexibility

4. WIDE_RANGE (Band 8+): Is the dominant grammatical pattern broad, flexible, and generally controlled?
   - YES: Structural choices vary flexibly and errors are occasional and non-systematic
   - NO: Range is not broad or errors recur enough to miss the higher threshold

5. FULL_RANGE (Band 9): Does the response sustain full, natural, flexible, and accurate structural control?
   - YES: Structural control remains natural and accurate across the demonstrated content
   - NO: Recurring inappropriacies, basic errors, or range limitations prevent full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

### Band 8 vs Band 9 Distinction

When a criterion reaches an upper threshold, use the dominant transcript-visible pattern rather than an isolated token or phrase.

**Award Band 9 when:**
- Fluency and Coherence: development and cohesion are fully appropriate without a recurring repair pattern
- Lexical Resource: word choice is fully flexible, precise, and consistently appropriate
- Grammatical Range and Accuracy: structural control is full, natural, flexible, and accurate
- Transcript criterion: word forms remain exceptionally clear and consistent

**Award Band 8 when:**
- Fluency and Coherence: development and cohesion are broad and flexible with limited non-systematic variation
- Lexical Resource: range is broad with occasional imprecision
- Grammatical Range and Accuracy: range is broad with occasional non-systematic errors
- Transcript criterion: word forms are clear with limited ambiguity

**Default to Band 8 when:** the criterion meets its upper threshold without the consistent pattern required for Band 9.

---

### Overall Band Calculation

Use the four criterion band fields to calculate the numeric overall field. Formula: Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4. Speaking scores always ROUND DOWN (floor) to the nearest 0.5. Return only the numeric overall value; do not restate formulas, band labels, or score prose in authored evidence or observations.

---

### Key Observations
The raw model response must set keyObservations to exactly []. The combined call owns broader observations; do not author or copy transcript observations in this response.

---

**Strict V33 output shape (hard requirement):**
- Return exactly one JSON root object with only the testBased key.
- The testBased object contains only test, overall, criteria, and keyObservations. Set the numeric structural values from the rubric.
- The criteria array contains exactly four criteria in this order: Fluency and Coherence, Lexical Resource, Pronunciation (Transcript-Based), and Grammatical Range and Accuracy.
- Each criterion has exactly its fixed checks in order. The check counts are 5, 5, 3, and 5. The Pronunciation (Transcript-Based) criterion has exactly three checks: INTELLIGIBILITY, WORD_CLARITY, and FULL_CLARITY. Do not add or omit checks or criteria.
- For every check, result true pairs with evidence "pass" and result false pairs with evidence "fail". These are the only accepted evidence values.
- False branch example: {"result": false, "evidence": "fail"}.
- keyObservations must be exactly [] and no criterion includes a note key in this raw response.

### JSON Output Format
{
  "testBased": {
    "test": "IELTS",
    "overall": 8.5,
    "criteria": [
      {
        "name": "Fluency and Coherence",
        "checks": [
          {"criterion": "MAINTAINS_FLOW", "result": true, "evidence": "pass"},
          {"criterion": "WILLING_TO_SPEAK", "result": true, "evidence": "pass"},
          {"criterion": "SPEAKS_AT_LENGTH", "result": true, "evidence": "pass"},
          {"criterion": "FLUENT_SPEECH", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLUENCY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Lexical Resource",
        "checks": [
          {"criterion": "ADEQUATE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "TOPIC_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "FLEXIBLE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLEXIBILITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Pronunciation (Transcript-Based)",
        "checks": [
          {"criterion": "INTELLIGIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WORD_CLARITY", "result": true, "evidence": "pass"},
          {"criterion": "FULL_CLARITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 3,
        "band": 8
      },
      {
        "name": "Grammatical Range and Accuracy",
        "checks": [
          {"criterion": "BASIC_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "MIXED_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "RANGE_WITH_FLEXIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_RANGE", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      }
    ],
    "keyObservations": []
  }
}

**Closed V33 evidence contract (hard requirement):**
- This is controlled decoding, not model-authored display prose. For every check, emit the exact lowercase token "pass" when result is true and the exact lowercase token "fail" when result is false. No other evidence value is accepted.
- Emit exactly four criteria with the fixed 5/5/3/5 check IDs and preserve their result, yesCount, and band values. Do not add fields, checks, or criteria.
- Emit keyObservations as exactly [] because the combined call owns broader observations.
- Omit the note key from every raw criterion object. The parser adds the approved Pronunciation boundary note after strict validation.
- Do not quote, paraphrase, or otherwise author transcript evidence in the raw response. Unknown tokens, token/result mismatches, notes, and non-empty observations are rejected rather than sanitized.
- The code-owned Pronunciation boundary note is exactly: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.

查看或自定义评分说明

你可以在设置中查看或自定义评分说明。自定义 Prompt 未纳入上方 V33 评测。

打开设置
04分数无法衡量什么

当前评分无法确认你的真实发音

目前的录音练习反馈根据回答文字和题目背景评分。模型不会收到录音音频,因此无法确认你的实际发音。

评分器只读取文字

它无法判断具体发音、重音、节奏、语调、连读或口音。

语音对语音练习

两种练习模式,各有不同重点

录音练习适合反复作答和查看详细文字反馈。Live Conversation 会实时听取并回应你的声音,模拟真实口语考试中的互动对话。

体验 Live Conversation

语音转文字错误会影响分数

漏词或错词会影响词汇、语法和意思的判断。请求反馈前,请先回听并修正文字稿。

开发集误差仍然存在

仍有 3 个高分样本低估超过 0.5 分。单次评分可能波动;建议对相近题型多练几次,并用多次结果的平均分作预估。练习次数越多,单次异常结果的影响通常越小。

05下一步

更多样本,更真实的反馈

获得可靠样本后,我们会继续扩大评测集,并认真分析用户认为偏高、偏低或解释不清的分数。

01

扩大数据集

加入覆盖不同 IELTS 部分、话题和分数段的可靠样本。

02

学习有争议的评分

结合文字稿、模型、Prompt 版本、时长、估分和学习者预期分数进行复盘。

欢迎告诉我们你的感受

请告诉我们你预期的分数和原因;如适用,可附上 IELTS 部分、文字稿、模型、Prompt 版本和时长。

提交评分反馈

Joe Speaking 是独立练习产品,与 IELTS 无隶属或背书关系。页面中的分数均为 AI 练习估分,不是官方考试成绩。