▲ Deputy Prime Minister and Minister of Science and ICT Bae Kyung-hoon delivers a welcoming speech at the first presentation event for the "Indigenous AI Foundation Model" project held at COEX in Gangnam-gu, Seoul, on December 30 of last year.
In the second-round evaluation of the Indigenous AI Foundation Model (Dokpamo) project to select South Korea's representative artificial intelligence, startups outperformed major conglomerates in global performance metrics.
Motif Technologies and Upstage claimed the first and second places, respectively, in the global AI performance evaluation index, while LG AI Research, which ranked first in the first-round evaluation, dropped to the bottom.
However, industry observers point out that it is difficult to conclude these results as the actual competitiveness or final evaluation outcome of the Dokpamo models, as this index represents only a portion of the overall evaluation, does not reflect Korean-language performance, and faces controversy over so-called "benchmaxing," which targets benchmark scores themselves.
According to the AI industry on August 13, the latest scores on the Artificial Intelligence Index (AAII) calculated by the global AI performance evaluation organization Artificial Analysis showed that Motif Technologies' "Motif 3" recorded the highest score among the four participating models at 47.
It was followed by Upstage's "Solar Open 2" at 37, SK Telecom's "A.X-K2" at 35, and LG AI Research's "K-EXAONE 2.0" at 31.
This result draws particular attention because LG AI Research previously ranked first in the first-round Dokpamo evaluation in December of last year with a comprehensive benchmark score of 33.6.
However, the latest AAII scores do not directly dictate the final results of the second-round Dokpamo evaluation.
In the second-round Dokpamo evaluation, benchmark testing accounts for 40 out of a total 100 points, with the AAII making up 25 points and the National Information Society Agency (NIA) benchmark evaluation accounting for 15 points.
The remaining weights consist of expert evaluations (35 points) and user evaluations (25 points).
The AAII is a comprehensive indexing method that combines multiple individual benchmarks across areas such as knowledge, reasoning, and coding into a single score.
While useful for comparing the overall intellectual capacity of models with a single figure, scores fluctuate significantly depending on which benchmarks are grouped together.
With each version update, specific fields such as mathematics are entirely omitted, and changes to major benchmarks can cause fluctuations of 10 to 20 points.
Notably, its limitation lies in the fact that Korean-language benchmarks are not reflected.
Experts note that it is difficult to gauge multilingual comprehension including Korean, performance in specific industrial tasks, or operational efficiency based solely on the AAII.
Meanwhile, controversy surrounding the reliability of benchmark scores themselves persists.
In the AI industry, the practice of aggressively targeting weak benchmark items to inflate scores regardless of actual performance is referred to as "benchmaxing."
Andrej Karpathy, an AI researcher at Anthropic and former co-founder of OpenAI, stated in an annual report last year, "We have lost trust in benchmarks."
An official from the domestic AI industry also remarked, "A model with raised benchmark scores does not necessarily guarantee actual performance."
The government maintains that because expert and user evaluations are conducted in parallel, any benchmark overfitting issues can be sufficiently compensated for.
Following the public evaluation that closed on August 12, the government plans to announce the results of the second-round Dokpamo evaluation soon.
In the final evaluation, three out of the four participating teams will survive and advance to the next stage.
(Photo: Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Video News
Video News