Classifying German Language Proficiency Levels Using Large Language Models

2512.06483v1 cs.CL, cs.AI 2025-12-09

Авторы:

Elias-Leander Ahlers, Witold Brunsmann, Malte Schilling

Abstract

Assessing language proficiency is essential for education, as it enables instruction tailored to learners needs. This paper investigates the use of Large Language Models (LLMs) for automatically classifying German texts according to the Common European Framework of Reference for Languages (CEFR) into different proficiency levels. To support robust training and evaluation, we construct a diverse dataset by combining multiple existing CEFR-annotated corpora with synthetic data. We then evaluate prompt-engineering strategies, fine-tuning of a LLaMA-3-8B-Instruct model and a probing-based approach that utilizes the internal neural state of the LLM for classification. Our results show a consistent performance improvement over prior methods, highlighting the potential of LLMs for reliable and scalable CEFR classification.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

Classifying German Language Proficiency Levels Using Large Language Models

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personali...

Leveraging KV Similarity for Online Structured Pruning in LLMs

Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum L...

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

SPAD: Seven-Source Token Probability Attribution with Syntactic Aggregation for ...

Навигация