IT·SCIENCE

UNIST develops 3D spatial AI that locates objects in 5 seconds, 400 times faster than existing tech

by
Koo Bon-hyuk
Published : June 8, 2026 - 09:33:46
    • Copy Completed!

View Korean Original

Memory usage cut by factor of 64

UNIST professor Joo Kyung-don (left) and researcher Bang Jae-hun. [UNIST]
UNIST professor Joo Kyung-don (left) and researcher Bang Jae-hun. [UNIST]

Researchers have developed an AI technology that can locate a user-specified object inside a robot's three-dimensional field of view in just five seconds using a simple text query.

A research team led by professor Joo Kyung-don of the AI Graduate School at UNIST (Ulsan National Institute of Science and Technology) announced Monday the development of "LightSplat," an open-vocabulary 3D spatial recognition system that identifies targets inside a reconstructed 3D space based on natural-language input.

Open-vocabulary 3D spatial recognition — technology that links three-dimensional space to human language — has drawn growing attention in recent years. Unlike conventional systems that recognize only a fixed list of predefined objects, it allows users to search a 3D environment using any word or phrase they choose.

Existing approaches, however, have faced limitations in processing speed, memory consumption and the precision with which object boundaries are distinguished.

LightSplat addresses those shortcomings by enabling searches through free-form natural language rather than preset categories such as chairs, desks or doors. Users can query for specific descriptions — "white sofa" or "egg on top of ramyun" — and the system will locate the target. Compared with previous open-vocabulary 3D recognition technology, LightSplat reduces memory usage to one sixty-fourth of the original level. The time needed to link semantic information to 3D Gaussians and make the space searchable by natural language has been cut to roughly five seconds — 50 to 400 times faster than the previous state of the art.

Results of a 3D semantic segmentation experiment using the ScanNet dataset. The system identified objects including an egg on top of ramyun, a teacup and a spatula more accurately than existing methods. [UNIST]
Results of a 3D semantic segmentation experiment using the ScanNet dataset. The system identified objects including an egg on top of ramyun, a teacup and a spatula more accurately than existing methods. [UNIST]

Despite the reductions in memory use and preparation time, recognition performance exceeded that of existing technology. In experiments using the LERF-OVS and DL3DV-OVS datasets, the system clearly distinguished objects ranging from small items — such as an egg on ramyun or tea in a glass — to distant cars and office furniture of varying sizes and arrangements. In a 3D semantic segmentation test on the ScanNet dataset, LightSplat recorded a mean intersection over union (mIoU) score of 37.11 across 19 categories. The mIoU metric measures how closely the area an AI identifies as an object overlaps with the actual ground-truth region.

"This technology could be applied to robot development with enhanced human-machine interaction that executes instructions given in natural speech, to AR and VR content creation where users designate targets by text for editing assistance, and to digital twin technology," professor Joo said.

The research was supported by the Ministry of Science and ICT and the Institute of Information and Communications Technology Planning and Evaluation through their AI Graduate School support project. It has been accepted for presentation at CVPR 2026, an international conference on computer vision.


nbgkoo@heraldcorp.com
This content was produced with the assistance of AI translation services.

MOST READ