POLITICS

Audit finds public agencies wasting budget on duplicate AI training data

by
Kim Hae-sol
Published : Aug. 12, 2026 - 12:00:00
    • Copy Completed!

View Korean Original

[Board of Audit and Inspection]
[Board of Audit and Inspection]

Public agencies have been building duplicate AI training datasets and wasting government funds by failing to check whether similar data already existed or to share project plans with one another in advance, a Board of Audit and Inspection review has found.

According to the audit body's report on "AI Industry Promotion Status III (AI Training Data)," released Wednesday, the Ministry of Science and ICT had invested a total of 1.63 trillion won ($1.15 billion) since 2017 to build 908 types of training datasets, which it has made publicly available through its AI Hub platform. Separately, 26 other public institutions — including the Ministry of Food and Drug Safety, the Seoul Metropolitan Government and Korea Expressway Corporation — have built an additional 313 types of training datasets and operate them through their own portals. The figures are as of last November.

Despite requirements to review whether existing data could be reused and to coordinate with other agencies before launching new projects, the audit found numerous cases in which agencies proceeded independently without such checks, resulting in near-identical datasets.

The most common pattern involved agencies building new datasets without first checking what had already been completed. After the Ministry of Science and ICT spent 7.38 billion won to build four types of image data — covering wildlife, lifestyle waste, concrete cracks and eggs — five agencies including the Seoul Metropolitan Government and Korea Expressway Corporation spent an additional 610 million won to build similar datasets from scratch. Of the wildlife training data Korea Expressway Corporation built in 2022 for roadkill prevention — covering 12 species across 60,000 images — 42 percent, or 25,000 images, were similar to data the Ministry of Science and ICT had already built in 2021 covering 11 species and roughly 330,000 images. Lifestyle waste detection and classification datasets built separately by the Seoul Metropolitan Government in 2020 (2,303 images) and by Seo-gu in Daejeon in 2025 (9,000 images) also overlapped significantly with a dataset the Ministry of Science and ICT built in 2020 covering 128 categories and 150,000 images — by 63 percent (1,454 images) and 58 percent (5,200 images), respectively.

The audit also uncovered duplication that arose even when agencies ran projects in the same year, simply because they did not share or coordinate their plans in advance. Between 2021 and 2023, while the Ministry of Science and ICT invested 23.3 billion won to build datasets related to pills, oral care and autonomous driving, three other agencies — including the Ministry of Food and Drug Safety — independently spent 840 million won to build similar datasets without prior coordination. In 2021, the ministry built pill auto-identification image data covering 4,999 types while the Ministry of Food and Drug Safety built a separate set covering 5,098 types; 1,411 types of images were found to be mutually similar. In 2023, oral care image data built by the Personal Information Protection Commission — 1,000 images in total — was found to be entirely similar to a dataset the Ministry of Science and ICT built the same year.

The Board of Audit and Inspection said the practice of agencies operating independently not only undermines the efficient use of public funds but also risks causing confusion for the companies and users who rely on the data.

The audit body has accordingly notified the minister of Science and ICT to consult with the minister of Interior and Safety and establish a pre-review system requiring individual public agencies to assess whether existing data could serve as a substitute before launching new AI training data projects, and to share and coordinate their construction plans in advance.


sunpine@heraldcorp.com
This content was produced with the assistance of AI translation services.

MOST READ