跳至主要內容
錄音練習評分揭秘

Joe Speaking 如何提高評分準確度

我們以錄音練習為例,公開說明 Joe Speaking 如何產生分數,以及如何持續測試並提高準確度。

查看評分過程

錄音練習分數

這是 AI 練習估分,不是官方成績。

手機上的錄音練習評分依據手機上的錄音練習分數

同一種方法,不同考試規則

IELTS 和 CELPIP 都使用可解釋的「是/否」檢查,但評分標準和分數範圍不同。下方圖示、Prompt 與公開參考比較均以 IELTS 錄音練習為例。

01分數如何產生

AI 不會直接決定你的分數

AI 會逐項完成「是/否」檢查,再根據結果得出四項分數和練習總分。錄音練習只傳送文字,因此發音項只能參考回答文字是否清晰。

錄音練習如何產生 AI 練習分數Joe Speaking 依序進行「是/否」檢查,由通過數得出四項分數和練習總分。評分器只收到文字,因此發音項只能參考回答文字是否清晰。1讀取回答回答文字 + 題目2逐項 是/否 檢查範例:詞彙詞彙足夠主題詞彙靈活使用範圍廣泛完全靈活4 個「是」→ 詞彙 8 分3得出四項分數流利與連貫9詞彙8文法8發音(文字參考)7通過數確定分數或範圍4計算總分(9 + 8 + 8 + 7) ÷ 4平均 8.00總分8.0

發音分數只參考文字清晰度

評分器聽不到發音、重音、節奏或語調,只能根據回答文字是否清晰來判斷。

02評測結果

我們如何測試評分準確度

我們參考公開考生案例,由 Joe Speaking 匿名化整理成 10 份文字樣本。每種評分設定對每個樣本執行三次,取平均值,再與參考分數比較。差距越小,結果越接近。

我們對同一組 10 個樣本、每種支援的推理強度各執行三次。下方是完整評分矩陣與模型評測排名。

2026 年 8 月 14 日,我們在 Vertex AI 上同時重跑了 Gemini 3.7 Flash 高推理強度、Flash 3 低推理強度和 Flash 3 中推理強度;90 次執行均傳回有效評分,其餘欄位保留最新發布的 V2 執行結果。

參考來源

公開考生表現案例,僅作為參考。 查看公開參考來源

評測輸入

10 份由 Joe Speaking 匿名化整理的 Part 3 文字樣本及對應題目

每個結果

每種推理強度執行三次,再取平均

對比測試如何進行

參考分數比較如何進行同一個樣本、Prompt 和模型執行三次,平均值與公開參考分數比較,再根據差距檢視和改進。一組設定Part 3 樣本 + Prompt + 模型每次上下文相同執行三次 15.0 25.5 35.5三次平均值5.33參考分數 5.0差距 0.33檢查 → 修改 → 再次評測下方公開完整結果表
V2 · 2026 年 8 月 14 日330 次模型執行

完整 V2 評測矩陣

選擇模型查看每個分數:10 個樣本、最多四種推理強度,每種執行三次。

Sample 01

參考分 5.0

最小
— / — / —
5.5 / 5.0 / 5.5
6.0 / 6.0 / 6.0
6.0 / 6.0 / 6.0

Sample 02

參考分 6.0

最小
— / — / —
5.5 / 5.5 / 6.5
6.5 / 6.5 / 6.5
6.5 / 6.5 / 6.5

Sample 03

參考分 6.0

最小
— / — / —
6.5 / 6.5 / 6.5
6.5 / 6.5 / 7.0
6.5 / 6.5 / 6.0

Sample 04

參考分 7.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 6.5
7.0 / 7.0 / 6.5

Sample 05

參考分 7.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 6.5 / 7.0

Sample 06

參考分 7.5

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 7.0 / 7.0

Sample 07

參考分 8.0

最小
— / — / —
7.5 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 7.5

Sample 08

參考分 8.0

最小
— / — / —
6.5 / 6.5 / 6.5
7.0 / 7.0 / 7.0
7.0 / 7.0 / 7.0

Sample 09

參考分 8.5

最小
— / — / —
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0

Sample 10

參考分 9.0

最小
— / — / —
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
8.0 / 8.0 / 8.0
03V2 模型評測

為什麼我們推薦 Gemini 3.7 Flash 中推理強度

Gemini 3.7 Flash 中推理強度把平均評分差距從 Flash 3 低推理強度的 0.73 降到 0.53,而且比 Flash 3 中推理強度快約 15 秒。它與 Gemini 3.7 高推理強度的平均評分差距相同(0.53),但耗時和積分更少。

Gemini 3 Flash Preview · 中

平均差距0.47
誤差 ≤ 0.5 的樣本6/10
預估積分≈ 13

差距最小,但最慢

Gemini 3.7 Flash · 中

推薦
平均差距0.53
誤差 ≤ 0.5 的樣本6/10
預估積分≈ 13

推薦

Gemini 3 Flash Preview · 低

平均差距0.73
誤差 ≤ 0.5 的樣本6/10
預估積分≈ 9

成本較低,準確度較低

預估積分已包含目前 2 倍費率:Flash 3 低推理強度約 9、Flash 3 中推理強度約 13、Gemini 3.7 中推理強度約 13。完整回饋請求可能更貴;實際消耗會隨輸入、輸出和推理長度變化。

03評測 Prompt

查看本次評測使用的 Prompt

可切換 Part 1、Part 2 和 Part 3。這是上方結果所測試的 V33 評測 Prompt,不是目前正式環境的評分 Prompt。

我們公開 Prompt,讓學習者和老師可以檢查規則、提出質疑並幫助我們改進。

評測 Prompt
V33
更新日期
2026 年 7 月 17 日
範本
feedback-combined-plus-ielts-recording-v33-v6

評分設定

saved-recording-dedicated-language-calibration-v33-output-policy-2026-07-22

You are scoring one saved IELTS Speaking recording from an ASR transcript.

IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}

Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}

Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}

## 2. IELTS Test-Based Feedback

Topic/Prompt: {{PART_1_TOPIC}}

### Evaluation Method
Evaluate the fixed graduated checks below in their listed order. Assess only patterns visible in the supplied transcript. The model sets each criterion band from the fixed check sequence; the parser validates that band.

**TRANSCRIPT CONTEXT:**
This is an ASR transcript of spoken English. Apply these tolerance guidelines:
- Judge meaning, organization, and language patterns from the supplied text; tolerate minor grammar slips and ASR artifacts.
- Treat textual repetition, filler tokens, and self-corrections as transcript patterns; do not infer sound or delivery qualities.
- Do not penalize punctuation, capitalization, or formatting.
- Only fail a check when the transcript text shows that its criterion is not met.

**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; do not impose a quota for detail or linking. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.

**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate transcript-visible word-form consistency rubric for Pronunciation below; do not infer sound or delivery from text.
- Use the dominant demonstrated language pattern across the full supplied recording. An isolated slip, brief answer, or strong phrase is not an automatic veto or pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Allow isolated non-systematic inaccuracies when the broader transcript pattern remains controlled; do not require a literally perfect transcript.

### Official IELTS Speaking Criteria

---

**Criterion 1: Fluency and Coherence**
*Evaluates: connected development, logical organization, and cohesive links visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the transcript show connected progression of ideas?
   - YES: Ideas continue with understandable links, even when wording repeats or repairs appear in the text
   - NO: The text repeatedly breaks into disconnected fragments, making meaning hard to follow

2. WILLING_TO_SPEAK (Band 6+): Does the response provide enough connected development for its question?
   - YES: The response supplies relevant development; linking may be mechanical
   - NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed

3. SPEAKS_AT_LENGTH (Band 7+): Does the supplied response sustain relevant development across its content?
   - YES: Relevant ideas are developed coherently with flexible cohesive links
   - NO: Development remains limited, repetitive, or mechanically linked

4. FLUENT_SPEECH (Band 8+): Is the dominant transcript pattern coherent, relevant, well developed, and flexibly linked?
   - YES: The response maintains clear organization and flexible cohesion across its demonstrated content
   - NO: Recurring repetition or coherence limitations make the higher threshold unsupported

5. FULL_FLUENCY (Band 9): Does the transcript show fully appropriate development and cohesion without a recurring repair pattern?
   - YES: Development and cohesion remain fully appropriate throughout the demonstrated content
   - NO: Repetition, self-correction, or imprecise cohesion recurs in the transcript

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 2: Lexical Resource**
*Evaluates: range, precision, and appropriacy of word choices visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the response show enough vocabulary for familiar topics?
   - YES: Word choices support discussion of the topic and communicate the intended meaning
   - NO: Limited word choices repeatedly prevent clear expression

2. TOPIC_VOCAB (Band 6+): Does the response use topic-appropriate vocabulary with some variety?
   - YES: Word choices fit the topic, include some less-common items, and show attempts at paraphrasing
   - NO: Word choices remain very basic, show little variety, or repeatedly fail to paraphrase the intended meaning

3. FLEXIBLE_VOCAB (Band 7+): Does the response use vocabulary flexibly and appropriately across its ideas?
   - YES: Word choices include less-common or idiomatic resources, support successful paraphrasing, and show collocational awareness
   - NO: Limited flexibility, unsuccessful paraphrasing, imprecise choices, or unsuitable combinations recur

4. WIDE_RANGE (Band 8+): Is the dominant lexical pattern broad in range, flexible, and precise?
   - YES: The response uses a broad range with flexibility and precision, including less-common or idiomatic resources when appropriate
   - NO: Range, flexibility, or precision remains limited at the higher threshold

5. FULL_FLEXIBILITY (Band 9): Does the response show full flexibility and precision in word choice?
   - YES: Word choices remain consistently natural, accurate, and precise across the demonstrated content, with full flexibility in paraphrasing and collocation
   - NO: Recurring imprecision or inappropriacy prevents full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: transcript-visible word-form consistency and clarity only*

**Transcript boundary:** Do not infer sound, delivery, or listener response from text. Assess only recognizable word forms, consistency, and ambiguity visible in the transcript.

Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Does the transcript's recognizable wording support generally clear meaning?
   - YES: Word forms in the transcript support clear meaning in context
   - NO: Unclear or ambiguous forms repeatedly obscure meaning

2. WORD_CLARITY (Band 7+): Are word forms transcribed consistently without confusion patterns?
   - YES: Word forms remain consistent and unambiguous across the transcript
   - NO: Confusion, ambiguity, or inconsistent forms recur

3. FULL_CLARITY (Band 8+): Is the transcript consistently clear without unresolved ambiguity?
   - YES: The transcript's word forms remain clear and unambiguous throughout its demonstrated content
   - NO: Some forms or phrases remain unresolved or potentially mis-transcribed

**Boundary note:** Omit the `note` key from the raw response. After strict validation, code attaches this exact display note: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

**Scoring Rationale:**
Use transcript-visible word-form consistency only. Do not infer sound, delivery, or listener effort from the transcript.

Band Mapping (derived from the fixed three checks; do not restate this mapping in authored text):
- 3 yes = Band 8 by default; use Band 9 only when the transcript shows exceptional consistency
- 2 yes = Band 7
- With exactly one passed check, set Band 6 only when WORD_CLARITY is narrowly missed; otherwise set Band 5
- 0 yes = Band 4 or below

**Band 8 vs 9 Distinction (when all three checks pass):**
- Band 8: The transcript is clear and consistent
- Band 9: The transcript shows exceptional clarity and consistency, including complex or technical terminology when present

---

**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: range of structures and grammatical control visible in the transcript*

Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the response show basic forms with reasonable accuracy?
   - YES: Basic forms are generally controlled and meaning remains clear
   - NO: Basic-form errors repeatedly interfere with meaning

2. MIXED_STRUCTURES (Band 6+): Does the response use a mix of simple and complex forms?
   - YES: The transcript demonstrates more than one structural pattern with generally clear meaning
   - NO: The transcript remains restricted to simple forms or complex attempts repeatedly cause confusion

3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the response use a range of complex forms with some flexibility?
   - YES: Complex forms vary and are generally controlled even when occasional errors remain
   - NO: Complex forms show recurring errors or limited flexibility

4. WIDE_RANGE (Band 8+): Is the dominant grammatical pattern broad, flexible, and generally controlled?
   - YES: Structural choices vary flexibly and errors are occasional and non-systematic
   - NO: Range is not broad or errors recur enough to miss the higher threshold

5. FULL_RANGE (Band 9): Does the response sustain full, natural, flexible, and accurate structural control?
   - YES: Structural control remains natural and accurate across the demonstrated content
   - NO: Recurring inappropriacies, basic errors, or range limitations prevent full control

Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below

---

### Band 8 vs Band 9 Distinction

When a criterion reaches an upper threshold, use the dominant transcript-visible pattern rather than an isolated token or phrase.

**Award Band 9 when:**
- Fluency and Coherence: development and cohesion are fully appropriate without a recurring repair pattern
- Lexical Resource: word choice is fully flexible, precise, and consistently appropriate
- Grammatical Range and Accuracy: structural control is full, natural, flexible, and accurate
- Transcript criterion: word forms remain exceptionally clear and consistent

**Award Band 8 when:**
- Fluency and Coherence: development and cohesion are broad and flexible with limited non-systematic variation
- Lexical Resource: range is broad with occasional imprecision
- Grammatical Range and Accuracy: range is broad with occasional non-systematic errors
- Transcript criterion: word forms are clear with limited ambiguity

**Default to Band 8 when:** the criterion meets its upper threshold without the consistent pattern required for Band 9.

---

### Overall Band Calculation

Use the four criterion band fields to calculate the numeric overall field. Formula: Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4. Speaking scores always ROUND DOWN (floor) to the nearest 0.5. Return only the numeric overall value; do not restate formulas, band labels, or score prose in authored evidence or observations.

---

### Key Observations
The raw model response must set keyObservations to exactly []. The combined call owns broader observations; do not author or copy transcript observations in this response.

---

**Strict V33 output shape (hard requirement):**
- Return exactly one JSON root object with only the testBased key.
- The testBased object contains only test, overall, criteria, and keyObservations. Set the numeric structural values from the rubric.
- The criteria array contains exactly four criteria in this order: Fluency and Coherence, Lexical Resource, Pronunciation (Transcript-Based), and Grammatical Range and Accuracy.
- Each criterion has exactly its fixed checks in order. The check counts are 5, 5, 3, and 5. The Pronunciation (Transcript-Based) criterion has exactly three checks: INTELLIGIBILITY, WORD_CLARITY, and FULL_CLARITY. Do not add or omit checks or criteria.
- For every check, result true pairs with evidence "pass" and result false pairs with evidence "fail". These are the only accepted evidence values.
- False branch example: {"result": false, "evidence": "fail"}.
- keyObservations must be exactly [] and no criterion includes a note key in this raw response.

### JSON Output Format
{
  "testBased": {
    "test": "IELTS",
    "overall": 8.5,
    "criteria": [
      {
        "name": "Fluency and Coherence",
        "checks": [
          {"criterion": "MAINTAINS_FLOW", "result": true, "evidence": "pass"},
          {"criterion": "WILLING_TO_SPEAK", "result": true, "evidence": "pass"},
          {"criterion": "SPEAKS_AT_LENGTH", "result": true, "evidence": "pass"},
          {"criterion": "FLUENT_SPEECH", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLUENCY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Lexical Resource",
        "checks": [
          {"criterion": "ADEQUATE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "TOPIC_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "FLEXIBLE_VOCAB", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_FLEXIBILITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      },
      {
        "name": "Pronunciation (Transcript-Based)",
        "checks": [
          {"criterion": "INTELLIGIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WORD_CLARITY", "result": true, "evidence": "pass"},
          {"criterion": "FULL_CLARITY", "result": true, "evidence": "pass"}
        ],
        "yesCount": 3,
        "band": 8
      },
      {
        "name": "Grammatical Range and Accuracy",
        "checks": [
          {"criterion": "BASIC_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "MIXED_STRUCTURES", "result": true, "evidence": "pass"},
          {"criterion": "RANGE_WITH_FLEXIBILITY", "result": true, "evidence": "pass"},
          {"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
          {"criterion": "FULL_RANGE", "result": true, "evidence": "pass"}
        ],
        "yesCount": 5,
        "band": 9
      }
    ],
    "keyObservations": []
  }
}

**Closed V33 evidence contract (hard requirement):**
- This is controlled decoding, not model-authored display prose. For every check, emit the exact lowercase token "pass" when result is true and the exact lowercase token "fail" when result is false. No other evidence value is accepted.
- Emit exactly four criteria with the fixed 5/5/3/5 check IDs and preserve their result, yesCount, and band values. Do not add fields, checks, or criteria.
- Emit keyObservations as exactly [] because the combined call owns broader observations.
- Omit the note key from every raw criterion object. The parser adds the approved Pronunciation boundary note after strict validation.
- Do not quote, paraphrase, or otherwise author transcript evidence in the raw response. Unknown tokens, token/result mismatches, notes, and non-empty observations are rejected rather than sanitized.
- The code-owned Pronunciation boundary note is exactly: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."

Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.

查看或自訂評分說明

你可以在設定中查看或自訂評分說明。自訂 Prompt 未納入上方 V33 評測。

開啟設定
04分數無法衡量什麼

目前評分無法確認你的真實發音

目前的錄音練習回饋根據回答文字和題目背景評分。模型不會收到錄音音訊,因此無法確認你的實際發音。

評分器只讀取文字

它無法判斷具體發音、重音、節奏、語調、連音或口音。

語音對語音練習

兩種練習模式,各有不同重點

錄音練習適合反覆作答和查看詳細文字回饋。Live Conversation 會即時聽取並回應你的聲音,模擬真實口說考試中的互動對話。

體驗 Live Conversation

語音轉文字錯誤會影響分數

漏字或錯字會影響詞彙、文法和意思的判斷。請在要求回饋前先回聽並修正文字稿。

開發集誤差仍然存在

仍有 3 個高分樣本低估超過 0.5 分。單次評分可能波動;建議針對相近題型多練幾次,並以多次結果的平均分作為預估。練習次數越多,單次異常結果的影響通常越小。

05下一步

更多樣本,更真實的回饋

取得可靠樣本後,我們會繼續擴大評測集,並認真分析使用者認為偏高、偏低或解釋不清的分數。

01

擴大資料集

加入涵蓋不同 IELTS 部分、主題和分數帶的可靠樣本。

02

從有爭議的評分學習

結合逐字稿、模型、Prompt 版本、時長、估分和學習者預期分數進行檢視。

歡迎告訴我們你的感受

請告訴我們你預期的分數和原因;如適用,可附上 IELTS 部分、逐字稿、模型、Prompt 版本和時長。

提交評分回饋

Joe Speaking 是獨立練習產品,與 IELTS 無隸屬或背書關係。頁面中的分數均為 AI 練習估分,不是官方考試成績。