Language Proficiency Classification Tools vs. Generative Artificial Intelligence: Scoring Behavior on English Language Learner Materials
DOI:
https://doi.org/10.13021/jssr2026.5627Abstract
There is an increasing variety of automatic text analyzers that can assess both student-generated text and learning materials, often according to metrics such as the Common European Framework of Reference for Languages (CEFR). This can help English second language learners find educational resources that suit their proficiency level; however, the various algorithms used by these tools can produce inconsistent or inaccurate results. This study used 3 tools designed for CEFR analysis and 3 GenAI tools to score 30 writing samples from the Education First Cambridge Open Language Database (EFCAMDAT) and 27 reading passages from Lingua, an online language-learning platform. Scoring behavior was analyzed for accuracy, agreement, and differences across tools and underlying levels. On average, each tool estimated proficiency within one level of the reference classification. For writing samples, analyses using Krippendorff’s alpha indicated moderate agreement among the 3 GenAI tools (α = 0.748, 95% BCa CI = [0.691, 0.796]). For reading passages, agreement was weaker when considering all 6 tools together (α = 0.61, [0.4668, 0.7194]), but stronger when comparing only the 3 GenAI tools (α = 0.863, [0.7942, 0.9244]) and moderate when comparing only the 3 CEFR-specific tools (α = .705, [0.6619, 0.8312]). Mixed-effects models indicated that the underlying level and tool significantly influenced estimations, and for reading passages, their interaction was also significant. While analyzers may be able to detect overall proficiency differences, scoring behavior varies across tools and text levels, emphasizing the need for careful interpretation of automated CEFR classifications.


