▲ AI Grading for CSAT Essay and Short-Answer Questions
As the National Education Commission reviews plans to introduce essay and short-answer evaluations to the College Scholastic Ability Test (CSAT), attention is turning to artificial intelligence (AI) automated grading technology capable of handling large-scale answer sheets.
While expectations are rising that utilizing AI could quickly grade the answers of hundreds of thousands of test-takers and reduce discrepancies among human graders, critics point out that current technology cannot sufficiently guarantee fairness and reliability in grading.
According to the education community on the 28th, grading has emerged as the biggest challenge in introducing essay and short-answer evaluations for the CSAT and expanding similar assessments for school records, both of which are currently being reviewed by the commission.
If exams continue to consist mostly of multiple-choice questions like the current CSAT, correctness can be uniformly determined. However, essay and short-answer evaluations require a comprehensive assessment of students' logical structures, supporting evidence, and conceptual understanding in their written answers.
Cha Jeong-in, chairperson of the National Education Commission who directly leads the Special Committee on College Admission Systems, argues that AI grading can solve such problems.
However, as a result of reviewing domestic and international AI automated grading research trends in an issue paper recently published by researchers at the Korea Institute for Curriculum and Evaluation (KICE), it was found that there are numerous hurdles to overcome before entrusting grading to AI.
A representative issue is the "black box" problem, where it is difficult to know what grounds and processes AI used to calculate a specific score.
According to the researchers, Educational Testing Service (ETS) in the United States did not adopt automated scoring models in actual grading despite developing models with high consistency with human graders, citing the difficulty of explaining the basis for score calculation.
This is because it would be difficult to explain the grounds if test-takers or parents file objections regarding their scores.
The researchers also pointed out that even if generative AI provides detailed explanations for score calculations, the problem is not completely resolved.
This is because there is no guarantee that the explanations provided by AI match the actual grounds used to calculate the scores.
Doubts are also raised about whether AI properly evaluates the competencies intended to be measured in essay and short-answer evaluations.
There is a risk that AI might calculate scores based on superficial characteristics such as word count rather than the quality of logic or evidence in the answers.
The researchers pointed out that in this case, it could lead to results that fail to properly fulfill the purpose of introducing essay and short-answer evaluations, which is to assess critical thinking and argumentation skills.
Consistency and reproducibility in grading are also cited as tasks to be resolved.
Generative AI probabilistically generates highly probable results based on input contexts rather than determining correct answers according to fixed rules.
Because of this, results can vary depending on prompts, answer input sequences, model versions, and detailed settings. The researchers pointed out that it is difficult to guarantee that the exact same scores and evaluations will be produced every time, even if the same answers and grading criteria are entered.
Due to these limitations, the researchers suggested that stricter standards are required when utilizing AI for "high-stakes assessments" like the CSAT, which determine students' college admissions and selection.
They also added that the role of AI should be minimized while increasing the proportion of human review.
The researchers recommended that "AI should be positioned as a supportive system to assist teachers' judgments rather than the entity that determines scores," advising that grading methods with high interpretability should be prioritized in high-stakes assessments.