Writing practice research · October 2026
IELTS Academic Writing criterion profiles: grammar, ties and score gaps
Grammatical Range and Accuracy had the lowest mean stored estimate in both studied cohorts. Shared-lowest and uniquely-lowest counts show how often that pattern applied to individual practice responses.
Published · Research publisher: Lucas Weaver
- Retained responses
- 1,600
- Publication
- 5 October 2026
- Evidence
- AI practice estimates
Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.
01
What does a lowest criterion actually mean?
The earlier IELTS benchmark reported that grammar had the lowest mean criterion estimate in its Academic cohorts. A mean alone cannot tell whether grammar is the sole low criterion in most responses, or whether several criteria frequently tie. This study adds response-level lowest-criterion counts and a separate first-response-per-account comparison.
We distinguish a unique lowest criterion from a shared lowest criterion. If two criteria have the same minimum value, both count as shared lowest and neither counts as uniquely lowest. If all four estimates are equal, all four count as shared lowest. The shared-lowest percentages can therefore sum to more than 100%.
The database uses one task-coverage field for both tasks. In the public tables we label it Task Achievement for Academic Task 1 and Task Response for Academic Task 2. This is a display mapping, not an assertion that the two tasks assess exactly the same communicative behavior.
Task and rubric reference: IELTS Writing band descriptors.
02
How often was language control the lowest estimate?
In Academic Writing Task 1, Grammatical Range and Accuracy shared the lowest estimate in 619 of 677 responses (91.4%). It was the sole lowest criterion in 186 (27.5%).
In Academic Writing Task 2, Grammatical Range and Accuracy shared the lowest estimate in 872 of 923 responses (94.5%). It was the sole lowest criterion in 307 (33.3%).
| Task cohort | Criterion | Mean | Median | Shared lowest: count / share | Unique lowest: count / share | Mean gap to highest |
|---|---|---|---|---|---|---|
| Academic · Writing Task 1 | Coherence and Cohesion | 5.71 | 6 | 280 / 41.4% | 1 / 0.1% | 0.1 |
| Academic · Writing Task 1 | Grammatical Range and Accuracy | 5.25 | 5 | 619 / 91.4% | 186 / 27.5% | 0.56 |
| Academic · Writing Task 1 | Lexical Resource | 5.56 | 5.5 | 390 / 57.6% | 8 / 1.2% | 0.25 |
| Academic · Writing Task 1 | Task Achievement | 5.57 | 5 | 396 / 58.5% | 25 / 3.7% | 0.24 |
| Academic · Writing Task 2 | Coherence and Cohesion | 5.83 | 6 | 333 / 36.1% | 1 / 0.1% | 0.11 |
| Academic · Writing Task 2 | Grammatical Range and Accuracy | 5.32 | 5 | 872 / 94.5% | 307 / 33.3% | 0.63 |
| Academic · Writing Task 2 | Lexical Resource | 5.62 | 5 | 563 / 61.0% | 6 / 0.7% | 0.33 |
| Academic · Writing Task 2 | Task Response | 5.81 | 6 | 370 / 40.1% | 13 / 1.4% | 0.14 |
Native practice-estimate scales: IELTS 0–9; Cambridge 0–5. Gap means compare each criterion with the highest criterion in the same response.
03
How much of the pattern comes from tied estimates?
All four criterion estimates were equal in 210 of 677 Academic Writing Task 1 responses (31.0%). Those responses contribute to every shared-lowest count.
All four criterion estimates were equal in 227 of 923 Academic Writing Task 2 responses (24.6%). Those responses contribute to every shared-lowest count.
Shared minima overlap. The unique-lowest shares are mutually exclusive, but they do not sum to 100% because responses with tied minima do not have a unique lowest criterion. This prevents a tied profile from being reported as four separate problems.
04
What happens when repeat responses are removed?
Grammatical Range and Accuracy remains the lowest mean criterion under the first-response rule in both cohorts. The full criterion table is provided so that the ordering can be checked directly.
A systematic ordering can also be an evaluator behavior. No blinded human-rating comparison was performed, and the historical evaluation records do not identify every model and prompt version.
| Task cohort | Criterion | Responses | Mean estimate | Shared lowest | Unique lowest |
|---|---|---|---|---|---|
| Academic · Writing Task 1 | Coherence and Cohesion | 372 | 5.54 | 43.0% | 0.0% |
| Academic · Writing Task 1 | Grammatical Range and Accuracy | 372 | 5.09 | 91.1% | 28.2% |
| Academic · Writing Task 1 | Lexical Resource | 372 | 5.43 | 54.3% | 1.3% |
| Academic · Writing Task 1 | Task Achievement | 372 | 5.42 | 57.8% | 4.0% |
| Academic · Writing Task 2 | Coherence and Cohesion | 546 | 5.66 | 35.2% | 0.2% |
| Academic · Writing Task 2 | Grammatical Range and Accuracy | 546 | 5.13 | 94.9% | 34.6% |
| Academic · Writing Task 2 | Lexical Resource | 546 | 5.46 | 58.6% | 0.5% |
| Academic · Writing Task 2 | Task Response | 546 | 5.65 | 36.8% | 1.1% |
05
What should change in teaching or revision?
The distinction between a shared and a unique minimum changes the practical interpretation. A criterion that frequently shares the lowest value may be part of a broad profile rather than the single obstacle holding a response back. We report both counts so readers can see that distinction instead of treating every shared minimum as a separate diagnosis.
Use a lower criterion estimate as a prompt for inspection. Review the actual feedback and the response before choosing an exercise. The tables cannot establish which grammatical structures, vocabulary choices, task omissions or organizational problems are present in an individual answer.
Means summarize the checker’s estimate pattern in this practice cohort. The native scales are ordered assessment estimates, and differences between decimal means should not be interpreted as precise amounts of language ability. We retain medians and response counts alongside the means for that reason.
The first-response-per-account check asks whether repeat contributors change the ordering. It does not remove external assistance, prompt differences or potential scorer bias. Agreement between the two summaries supports the stability of this observed pattern under one sampling change, rather than validating the scoring system.
Methods
How this report was prepared
The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.
Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.
The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.
Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.
Read the full dataset definition, formulas, exclusions and reporting thresholds.
Questions about the findings
Are shared-lowest percentages supposed to total 100%?
No. Every criterion tied at the minimum is counted. The same response can contribute to several shared-lowest counts.
Does this prove the lowest criterion is the hardest exam skill?
No. The observed ordering belongs to this checker’s practice estimates. Writing differences and scorer behavior can both produce the pattern.
Evidence
Aggregate data, sources and citation
The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.
Dataset version and integrity
Version 2026-10-05.1. SHA-256 of this report’s JSON file:
e2da62d69ac962809577e5d535bf7db3d02d6946c20b52fe4aeb84a2f9270f15Official and contextual sources
Suggested citation
Weaver, Lucas. “IELTS Academic Writing criterion profiles: grammar, ties and score gaps.” IELTS Writing Checker, 5 October 2026. Version 2026-10-05.1. https://ieltswritingchecker.com/research/ielts-writing-criterion-bottlenecks-october-2026
Related writing research
IELTS Writing Checker is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.
