Writing practice research · October 2026
IELTS General Training Task 2 practice: a 123-response pilot study
The General Training Task 2 pilot contains 123 responses from 51 accounts. The median length was 286 words, but the length–score association weakened when repeat responses were removed.
Published · Research publisher: Lucas Weaver
- Retained responses
- 123
- Publication
- 5 October 2026
- Evidence
- AI practice estimates
Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.
01
Why does General Training need its own analysis?
The earlier IELTS benchmark was dominated by Academic tasks. General Training writers were represented by much smaller cohorts, which limited the claims the benchmark could support. This recent-window study keeps General Training Task 2 separate rather than borrowing Academic findings as evidence about its users.
The cohort passes our descriptive reporting threshold of 100 responses and 30 contributing accounts, but it remains a pilot. Several accounts contribute multiple answers. The number of responses should not be confused with the number of independent writers or with representative coverage of General Training candidates.
General Training Task 2 involves writing about a point of view, argument or problem, with a 250-word minimum. We use the selected task label to define the sample; we do not infer the writer’s profession, migration plans, first language, country or real examination result.
Task and rubric reference: IELTS General Training Writing: official format.
02
What did the retained practice responses look like?
The middle half of responses contained 257–310.5 words. The mean overall estimate was 5.92, and the median was 5.5. These values describe the checker’s stored practice outputs on a 0–9 scale.
The cohort contains no responses below 250 words under our counting method. Product submission checks restrict the lower tail, so the study does not estimate the effect of writing an under-length answer.
The score bands below use the full 123-response cohort as their denominator. The small lowest-score tail is suppressed. Suppression leaves a gap in the displayed percentages; it does not mean that the missing band had no responses.
| Estimate band | Responses | Share of cohort | Contributing accounts |
|---|---|---|---|
| 5 to below 6 | 59 | 48.0% | 24 |
| 6 to below 7 | 36 | 29.3% | 23 |
| 7–9 | 23 | 18.7% | 11 |
03
Which criterion most often received the lowest estimate?
Grammatical Range and Accuracy shared the lowest estimate in 108 responses (87.8%) and was uniquely lowest in 37 (30.1%).
Shared minima overlap when criteria tie. The pattern is evidence about this scoring system’s outputs within the pilot cohort; it does not establish the most common real-world writing error. No error-frequency classification or independent human score was collected for this study.
| Criterion | Mean estimate | Shared lowest | Unique lowest |
|---|---|---|---|
| Coherence and Cohesion | 5.92 | 39.0% | 0.0% |
| Grammatical Range and Accuracy | 5.5 | 87.8% | 30.1% |
| Lexical Resource | 5.75 | 59.3% | 1.6% |
| Task Response | 5.84 | 51.2% | 4.9% |
04
Why does repeated practice change the interpretation?
The length–score rank correlation changes from 0.279 across all retained responses to 0.098 in the 51-response first-per-account sample. The median length changes from 286 to 295 words.
That sensitivity makes a simple recommendation to write longer especially difficult to defend. The all-response result gives more weight to frequent contributors, while the first-response result gives each account one observation for this task. Neither design follows a controlled revision or establishes a causal length effect.
The first-per-account sample is smaller than the 100-response primary reporting threshold. It is displayed as a sensitivity check on the qualifying primary cohort, not as a separate full benchmark.
| Task cohort | Sampling rule | Responses | Median words | Mean estimate | Median estimate | Length–score correlation |
|---|---|---|---|---|---|---|
| IELTS · General · Writing Task 2 | all responses | 123 | 286 | 5.92 | 5.5 | 0.279 |
| IELTS · General · Writing Task 2 | first per account and task | 51 | 295 | 5.93 | 6 | 0.098 |
05
What is useful now, and what needs more evidence?
The pilot supports transparent descriptions of response length and stored criterion profiles. It also shows why contributing-account counts belong beside headline essay totals. For a teacher, these distributions can identify questions to investigate in a class; they cannot diagnose a particular learner.
A stronger General Training study would recruit more distinct writers, identify task prompts consistently, evaluate responses with a fixed scoring setup and obtain independent ratings. General Training Task 1 remains below this series’ primary response threshold and receives no detailed score analysis.
Readers should resist comparing the pilot’s average with the Academic average as a ranking of exam difficulty. The groups contain different writers and practice choices. This release has no design capable of separating those differences.
Methods
How this report was prepared
The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.
Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.
The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.
Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.
Read the full dataset definition, formulas, exclusions and reporting thresholds.
Questions about the findings
Does the pilot show that longer General Training essays are better?
No. The correlation becomes much smaller with one response per account, and neither analysis tests an editing intervention.
Why is there no detailed General Training Task 1 report?
Its recent cleaned cohort falls below the series’ 100-response threshold. A dedicated score report would give that sample more authority than the reporting design supports.
Evidence
Aggregate data, sources and citation
The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.
Dataset version and integrity
Version 2026-10-05.1. SHA-256 of this report’s JSON file:
90f35dab1dfbba98e47778c74f9befeda3cb9be428641598f8558d9269c8cf64Official and contextual sources
Suggested citation
Weaver, Lucas. “IELTS General Training Task 2 practice: a 123-response pilot study.” IELTS Writing Checker, 5 October 2026. Version 2026-10-05.1. https://ieltswritingchecker.com/research/ielts-general-training-task-2-practice-study-october-2026
Related writing research
IELTS Writing Checker is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.
