Sample 01
참고 5.0
- 최소
- — / — / —
- 낮음
- 5.5 / 5.0 / 5.5
- 중간
- 6.0 / 6.0 / 6.0
- 높음
- 6.0 / 6.0 / 6.0
녹음 연습을 예로 들어 Joe Speaking이 점수를 만드는 방법과 정확도를 테스트하고 개선하는 과정을 공개합니다.
채점 과정 보기녹음 연습 점수
AI 연습용 추정치이며 공식 결과가 아닙니다.


IELTS와 CELPIP 모두 설명 가능한 예/아니요 검사를 사용하지만 기준과 점수 범위는 다릅니다. 아래 도표, 프롬프트, 공개 비교는 IELTS 녹음 연습을 예로 듭니다.
AI가 각 IELTS 평가 항목을 예/아니요로 확인하고, 그 결과로 네 영역 점수와 전체 연습 점수를 계산합니다. 녹음 연습은 텍스트만 보내므로 발음 항목은 답변 텍스트의 명확성만 참고합니다.
소리, 강세, 리듬, 억양을 들을 수 없으며 답변 텍스트의 명확성만 사용합니다.
공개 응시자 사례를 참고해 Joe Speaking에서 익명화하고 정리한 텍스트 샘플 10개를 사용합니다. 각 설정으로 같은 샘플을 세 번 채점하고 평균을 낸 뒤 참고 밴드와 비교합니다. 차이가 작을수록 더 가깝습니다.
같은 샘플 10개를 지원되는 각 사고 수준으로 세 번씩 채점했습니다. 아래에 전체 점수와 모델 평가 순위가 있습니다.
2026년 8월 14일 Vertex AI에서 Gemini 3.7 Flash 높음, Flash 3 낮음, Flash 3 중간을 함께 다시 실행했고 90회 모두 유효한 점수를 반환했습니다. 나머지 열은 최신 공개 V2 실행 결과를 유지합니다.
참고 출처
평가 입력
각 결과
모델을 선택하면 샘플 10개, 최대 사고 수준 4개, 수준별 3회의 모든 점수를 볼 수 있습니다.
| 샘플 | 참고 밴드 | 최소1회 · 2회 · 3회 | 낮음1회 · 2회 · 3회 | 중간1회 · 2회 · 3회 | 높음1회 · 2회 · 3회 |
|---|---|---|---|---|---|
| Sample 01 | 5.0 | — / — / — | 5.5 / 5.0 / 5.5 | 6.0 / 6.0 / 6.0 | 6.0 / 6.0 / 6.0 |
| Sample 02 | 6.0 | — / — / — | 5.5 / 5.5 / 6.5 | 6.5 / 6.5 / 6.5 | 6.5 / 6.5 / 6.5 |
| Sample 03 | 6.0 | — / — / — | 6.5 / 6.5 / 6.5 | 6.5 / 6.5 / 7.0 | 6.5 / 6.5 / 6.0 |
| Sample 04 | 7.0 | — / — / — | 6.5 / 6.5 / 6.5 | 7.0 / 7.0 / 6.5 | 7.0 / 7.0 / 6.5 |
| Sample 05 | 7.0 | — / — / — | 6.5 / 6.5 / 6.5 | 7.0 / 7.0 / 7.0 | 7.0 / 6.5 / 7.0 |
| Sample 06 | 7.5 | — / — / — | 6.5 / 6.5 / 6.5 | 7.0 / 7.0 / 7.0 | 7.0 / 7.0 / 7.0 |
| Sample 07 | 8.0 | — / — / — | 7.5 / 8.0 / 8.0 | 8.0 / 8.0 / 8.0 | 8.0 / 8.0 / 7.5 |
| Sample 08 | 8.0 | — / — / — | 6.5 / 6.5 / 6.5 | 7.0 / 7.0 / 7.0 | 7.0 / 7.0 / 7.0 |
| Sample 09 | 8.5 | — / — / — | 8.0 / 8.0 / 8.0 | 8.0 / 8.0 / 8.0 | 8.0 / 8.0 / 8.0 |
| Sample 10 | 9.0 | — / — / — | 8.0 / 8.0 / 8.0 | 8.0 / 8.0 / 8.0 | 8.0 / 8.0 / 8.0 |
| 평균 차이 | — | 0.617 | 0.533 | 0.533 | |
참고 5.0
참고 6.0
참고 6.0
참고 7.0
참고 7.0
참고 7.5
참고 8.0
참고 8.0
참고 8.5
참고 9.0
Gemini 3.7 Flash 중간은 Flash 3 낮음보다 평균 점수 차이를 0.73에서 0.53으로 줄였고, Flash 3 중간보다 약 15초 빨랐습니다. Gemini 3.7 높음과 평균 점수 차이(0.53)가 같았고, 더 짧은 시간과 적은 크레딧으로 제공했습니다.
Gemini 3 Flash Preview · 중간
가장 가깝지만 가장 느림
Gemini 3.7 Flash · 중간
추천추천
Gemini 3 Flash Preview · 낮음
비용은 낮고 정확도도 낮음
예상 크레딧에는 현재 2배 요금이 포함됩니다. Flash 3 낮음은 약 9, Flash 3 중간은 약 13, Gemini 3.7 중간은 약 13입니다. 전체 피드백 요청은 더 비쌀 수 있으며 실제 사용량은 입력, 출력, 추론 길이에 따라 달라집니다.
Part 1, Part 2, Part 3를 전환할 수 있습니다. 위 결과에서 테스트한 V33이며 현재 실제 채점에 쓰이는 프롬프트는 아닙니다.
학습자와 교사가 규칙을 확인하고 결과에 의문을 제기하며 개선에 참여할 수 있도록 공개합니다.
채점 프로필
saved-recording-dedicated-language-calibration-v33-output-policy-2026-07-22
You are scoring one saved IELTS Speaking recording from an ASR transcript.
IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}
Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}
Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}
## 2. IELTS Test-Based Feedback
Topic/Prompt: {{PART_1_TOPIC}}
### Evaluation Method
Evaluate the fixed graduated checks below in their listed order. Assess only patterns visible in the supplied transcript. The model sets each criterion band from the fixed check sequence; the parser validates that band.
**TRANSCRIPT CONTEXT:**
This is an ASR transcript of spoken English. Apply these tolerance guidelines:
- Judge meaning, organization, and language patterns from the supplied text; tolerate minor grammar slips and ASR artifacts.
- Treat textual repetition, filler tokens, and self-corrections as transcript patterns; do not infer sound or delivery qualities.
- Do not penalize punctuation, capitalization, or formatting.
- Only fail a check when the transcript text shows that its criterion is not met.
**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; do not impose a quota for detail or linking. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.
**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate transcript-visible word-form consistency rubric for Pronunciation below; do not infer sound or delivery from text.
- Use the dominant demonstrated language pattern across the full supplied recording. An isolated slip, brief answer, or strong phrase is not an automatic veto or pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Allow isolated non-systematic inaccuracies when the broader transcript pattern remains controlled; do not require a literally perfect transcript.
### Official IELTS Speaking Criteria
---
**Criterion 1: Fluency and Coherence**
*Evaluates: connected development, logical organization, and cohesive links visible in the transcript*
Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the transcript show connected progression of ideas?
- YES: Ideas continue with understandable links, even when wording repeats or repairs appear in the text
- NO: The text repeatedly breaks into disconnected fragments, making meaning hard to follow
2. WILLING_TO_SPEAK (Band 6+): Does the response provide enough connected development for its question?
- YES: The response supplies relevant development; linking may be mechanical
- NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed
3. SPEAKS_AT_LENGTH (Band 7+): Does the supplied response sustain relevant development across its content?
- YES: Relevant ideas are developed coherently with flexible cohesive links
- NO: Development remains limited, repetitive, or mechanically linked
4. FLUENT_SPEECH (Band 8+): Is the dominant transcript pattern coherent, relevant, well developed, and flexibly linked?
- YES: The response maintains clear organization and flexible cohesion across its demonstrated content
- NO: Recurring repetition or coherence limitations make the higher threshold unsupported
5. FULL_FLUENCY (Band 9): Does the transcript show fully appropriate development and cohesion without a recurring repair pattern?
- YES: Development and cohesion remain fully appropriate throughout the demonstrated content
- NO: Repetition, self-correction, or imprecise cohesion recurs in the transcript
Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below
---
**Criterion 2: Lexical Resource**
*Evaluates: range, precision, and appropriacy of word choices visible in the transcript*
Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the response show enough vocabulary for familiar topics?
- YES: Word choices support discussion of the topic and communicate the intended meaning
- NO: Limited word choices repeatedly prevent clear expression
2. TOPIC_VOCAB (Band 6+): Does the response use topic-appropriate vocabulary with some variety?
- YES: Word choices fit the topic, include some less-common items, and show attempts at paraphrasing
- NO: Word choices remain very basic, show little variety, or repeatedly fail to paraphrase the intended meaning
3. FLEXIBLE_VOCAB (Band 7+): Does the response use vocabulary flexibly and appropriately across its ideas?
- YES: Word choices include less-common or idiomatic resources, support successful paraphrasing, and show collocational awareness
- NO: Limited flexibility, unsuccessful paraphrasing, imprecise choices, or unsuitable combinations recur
4. WIDE_RANGE (Band 8+): Is the dominant lexical pattern broad in range, flexible, and precise?
- YES: The response uses a broad range with flexibility and precision, including less-common or idiomatic resources when appropriate
- NO: Range, flexibility, or precision remains limited at the higher threshold
5. FULL_FLEXIBILITY (Band 9): Does the response show full flexibility and precision in word choice?
- YES: Word choices remain consistently natural, accurate, and precise across the demonstrated content, with full flexibility in paraphrasing and collocation
- NO: Recurring imprecision or inappropriacy prevents full control
Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below
---
**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: transcript-visible word-form consistency and clarity only*
**Transcript boundary:** Do not infer sound, delivery, or listener response from text. Assess only recognizable word forms, consistency, and ambiguity visible in the transcript.
Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Does the transcript's recognizable wording support generally clear meaning?
- YES: Word forms in the transcript support clear meaning in context
- NO: Unclear or ambiguous forms repeatedly obscure meaning
2. WORD_CLARITY (Band 7+): Are word forms transcribed consistently without confusion patterns?
- YES: Word forms remain consistent and unambiguous across the transcript
- NO: Confusion, ambiguity, or inconsistent forms recur
3. FULL_CLARITY (Band 8+): Is the transcript consistently clear without unresolved ambiguity?
- YES: The transcript's word forms remain clear and unambiguous throughout its demonstrated content
- NO: Some forms or phrases remain unresolved or potentially mis-transcribed
**Boundary note:** Omit the `note` key from the raw response. After strict validation, code attaches this exact display note: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."
**Scoring Rationale:**
Use transcript-visible word-form consistency only. Do not infer sound, delivery, or listener effort from the transcript.
Band Mapping (derived from the fixed three checks; do not restate this mapping in authored text):
- 3 yes = Band 8 by default; use Band 9 only when the transcript shows exceptional consistency
- 2 yes = Band 7
- With exactly one passed check, set Band 6 only when WORD_CLARITY is narrowly missed; otherwise set Band 5
- 0 yes = Band 4 or below
**Band 8 vs 9 Distinction (when all three checks pass):**
- Band 8: The transcript is clear and consistent
- Band 9: The transcript shows exceptional clarity and consistency, including complex or technical terminology when present
---
**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: range of structures and grammatical control visible in the transcript*
Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the response show basic forms with reasonable accuracy?
- YES: Basic forms are generally controlled and meaning remains clear
- NO: Basic-form errors repeatedly interfere with meaning
2. MIXED_STRUCTURES (Band 6+): Does the response use a mix of simple and complex forms?
- YES: The transcript demonstrates more than one structural pattern with generally clear meaning
- NO: The transcript remains restricted to simple forms or complex attempts repeatedly cause confusion
3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the response use a range of complex forms with some flexibility?
- YES: Complex forms vary and are generally controlled even when occasional errors remain
- NO: Complex forms show recurring errors or limited flexibility
4. WIDE_RANGE (Band 8+): Is the dominant grammatical pattern broad, flexible, and generally controlled?
- YES: Structural choices vary flexibly and errors are occasional and non-systematic
- NO: Range is not broad or errors recur enough to miss the higher threshold
5. FULL_RANGE (Band 9): Does the response sustain full, natural, flexible, and accurate structural control?
- YES: Structural control remains natural and accurate across the demonstrated content
- NO: Recurring inappropriacies, basic errors, or range limitations prevent full control
Band Mapping (derived from the fixed five checks; do not restate this mapping in authored text):
- 5 yes = Band 9
- 4 yes = Band 8
- 3 yes = Band 7
- 2 yes = Band 6
- 1 yes = Band 5
- 0 yes = Band 4 or below
---
### Band 8 vs Band 9 Distinction
When a criterion reaches an upper threshold, use the dominant transcript-visible pattern rather than an isolated token or phrase.
**Award Band 9 when:**
- Fluency and Coherence: development and cohesion are fully appropriate without a recurring repair pattern
- Lexical Resource: word choice is fully flexible, precise, and consistently appropriate
- Grammatical Range and Accuracy: structural control is full, natural, flexible, and accurate
- Transcript criterion: word forms remain exceptionally clear and consistent
**Award Band 8 when:**
- Fluency and Coherence: development and cohesion are broad and flexible with limited non-systematic variation
- Lexical Resource: range is broad with occasional imprecision
- Grammatical Range and Accuracy: range is broad with occasional non-systematic errors
- Transcript criterion: word forms are clear with limited ambiguity
**Default to Band 8 when:** the criterion meets its upper threshold without the consistent pattern required for Band 9.
---
### Overall Band Calculation
Use the four criterion band fields to calculate the numeric overall field. Formula: Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4. Speaking scores always ROUND DOWN (floor) to the nearest 0.5. Return only the numeric overall value; do not restate formulas, band labels, or score prose in authored evidence or observations.
---
### Key Observations
The raw model response must set keyObservations to exactly []. The combined call owns broader observations; do not author or copy transcript observations in this response.
---
**Strict V33 output shape (hard requirement):**
- Return exactly one JSON root object with only the testBased key.
- The testBased object contains only test, overall, criteria, and keyObservations. Set the numeric structural values from the rubric.
- The criteria array contains exactly four criteria in this order: Fluency and Coherence, Lexical Resource, Pronunciation (Transcript-Based), and Grammatical Range and Accuracy.
- Each criterion has exactly its fixed checks in order. The check counts are 5, 5, 3, and 5. The Pronunciation (Transcript-Based) criterion has exactly three checks: INTELLIGIBILITY, WORD_CLARITY, and FULL_CLARITY. Do not add or omit checks or criteria.
- For every check, result true pairs with evidence "pass" and result false pairs with evidence "fail". These are the only accepted evidence values.
- False branch example: {"result": false, "evidence": "fail"}.
- keyObservations must be exactly [] and no criterion includes a note key in this raw response.
### JSON Output Format
{
"testBased": {
"test": "IELTS",
"overall": 8.5,
"criteria": [
{
"name": "Fluency and Coherence",
"checks": [
{"criterion": "MAINTAINS_FLOW", "result": true, "evidence": "pass"},
{"criterion": "WILLING_TO_SPEAK", "result": true, "evidence": "pass"},
{"criterion": "SPEAKS_AT_LENGTH", "result": true, "evidence": "pass"},
{"criterion": "FLUENT_SPEECH", "result": true, "evidence": "pass"},
{"criterion": "FULL_FLUENCY", "result": true, "evidence": "pass"}
],
"yesCount": 5,
"band": 9
},
{
"name": "Lexical Resource",
"checks": [
{"criterion": "ADEQUATE_VOCAB", "result": true, "evidence": "pass"},
{"criterion": "TOPIC_VOCAB", "result": true, "evidence": "pass"},
{"criterion": "FLEXIBLE_VOCAB", "result": true, "evidence": "pass"},
{"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
{"criterion": "FULL_FLEXIBILITY", "result": true, "evidence": "pass"}
],
"yesCount": 5,
"band": 9
},
{
"name": "Pronunciation (Transcript-Based)",
"checks": [
{"criterion": "INTELLIGIBILITY", "result": true, "evidence": "pass"},
{"criterion": "WORD_CLARITY", "result": true, "evidence": "pass"},
{"criterion": "FULL_CLARITY", "result": true, "evidence": "pass"}
],
"yesCount": 3,
"band": 8
},
{
"name": "Grammatical Range and Accuracy",
"checks": [
{"criterion": "BASIC_STRUCTURES", "result": true, "evidence": "pass"},
{"criterion": "MIXED_STRUCTURES", "result": true, "evidence": "pass"},
{"criterion": "RANGE_WITH_FLEXIBILITY", "result": true, "evidence": "pass"},
{"criterion": "WIDE_RANGE", "result": true, "evidence": "pass"},
{"criterion": "FULL_RANGE", "result": true, "evidence": "pass"}
],
"yesCount": 5,
"band": 9
}
],
"keyObservations": []
}
}
**Closed V33 evidence contract (hard requirement):**
- This is controlled decoding, not model-authored display prose. For every check, emit the exact lowercase token "pass" when result is true and the exact lowercase token "fail" when result is false. No other evidence value is accepted.
- Emit exactly four criteria with the fixed 5/5/3/5 check IDs and preserve their result, yesCount, and band values. Do not add fields, checks, or criteria.
- Emit keyObservations as exactly [] because the combined call owns broader observations.
- Omit the note key from every raw criterion object. The parser adds the approved Pronunciation boundary note after strict validation.
- Do not quote, paraphrase, or otherwise author transcript evidence in the raw response. Unknown tokens, token/result mismatches, notes, and non-empty observations are rejected rather than sanitized.
- The code-owned Pronunciation boundary note is exactly: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."
Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.설정에서 채점 지침을 확인하거나 변경할 수 있습니다. 사용자 프롬프트는 위 V33 평가에 포함되지 않았습니다.
녹음 연습은 답변 텍스트와 질문 문맥을 채점합니다. 녹음 오디오가 모델에 전달되지 않으므로 실제 발음을 확인할 수 없습니다.
소리, 강세, 리듬, 억양, 연음, 악센트를 평가할 수 없습니다.
두 가지 연습 모드
Recording Practice는 답변을 반복하고 상세한 텍스트 피드백을 확인하는 데 적합합니다. Live Conversation은 음성을 실시간으로 듣고 응답하며 실제 말하기 시험의 대화 흐름을 재현합니다.
빠진 단어나 잘못 인식된 단어는 어휘, 문법, 의미 판단을 바꿀 수 있습니다. 피드백을 요청하기 전에 다시 듣고 수정하세요.
고밴드 샘플 3개는 0.5보다 더 낮게 평가되었습니다. 한 번의 점수는 달라질 수 있으므로 비슷한 문제를 여러 번 연습하고 그 평균을 더 안정적인 추정치로 활용하세요. 연습 횟수가 늘수록 한 번의 이례적인 결과가 미치는 영향은 보통 줄어듭니다.
신뢰할 수 있는 샘플이 생기면 평가 세트를 넓히고, 너무 높거나 낮거나 설명이 부족하다고 느낀 점수를 조사하겠습니다.
IELTS 파트, 주제, 밴드를 폭넓게 포함하는 신뢰할 수 있는 샘플을 추가합니다.
전사문, 모델, 프롬프트 버전, 시간, 추정 밴드와 기대 밴드를 함께 검토합니다.
예상한 밴드와 그 이유를 알려 주세요. 필요한 경우 IELTS 파트, 전사문, 모델, 프롬프트 버전과 시간도 포함해 주세요.
Joe Speaking은 독립적인 연습 제품이며 IELTS와 제휴하거나 보증받지 않았습니다. 표시된 점수는 AI 연습용 추정치이며 공식 시험 결과가 아닙니다.