SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2609.19916 · September 2026

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

Hyunji Lee, Jin Park, Jun Lee, Jeongwan Shin, Hyeyoung Park, Jin Hyun Park, Soojin Lee, Soha Lee, Heesung Yang, Hyunju Song, Jinsan An, Kilim Nam

arXiv · PDF · Open in the Atlas

Code that ran

We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. The repositories linked to it are listed below.

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage of such recent lexical change, and their English-oriented design makes it difficult to assess the typological properties of Korean, in which content words combine productively with functional morphemes. In this paper, we introduce KoNeoBench, a benchmark for evaluating LLMs' understanding of Korean neologisms. KoNeoBench is built on 1,785 Korean neologisms attested in online news since 2020 and curated through expert lexicographic review. Each entry provides usage examples, word-formation analyses, and dictionary-style definitions. Based on this resource, we define four tasks and report results on recent models, together with a human baseline. Our experiments show that current LLMs exhibit clear limitations in recovering source components, distinguishing semantic categories, and generating accurate definitions. These results reveal specific aspects of recent Korean lexical change that remain challenging for current LLMs. KoNeoBench is available at https://github.com/bcmilab/ko-neobench/ .

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2609.19916")
get_code_for_paper("2609.19916")
have("2609.19916")

The run record, dated, one paper per request, free:

curl https://syntology.ai/api/ran/2609.19916.json

A badge for a README (the split and the date, never a ratio):

[![Syntology run record](https://syntology.ai/api/ran/2609.19916.svg)](https://syntology.ai/paper/2609.19916)

Connect an agent — have() is free.