CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms
LLMs can often explain what Chinese internet slang like 886, yyds, or 彳亍 means, but many fail to reconstruct the original form behind it
Researchers built CNeo-Bench, a benchmark of 4,759 Chinese neologisms such as 886 (a numeric stand-in for 'bye-bye') and yyds (letters from a Pinyin phrase), sorted into 9 subcategories by how they're formed. Testing 18 LLMs including GPT-5.1, DeepSeek, and Kimi in a two-tier setup, most models scored below 40% on simply defining these terms, and on tasks requiring models to restore the original source form, models that correctly described a term often substituted a paraphrase instead of the actual source form. Giving models 1-3 examples recovered many hard cases, but 33-47% of errors persisted even at 3 examples.
METAL LAB explanatory visual
CNeo-Bench's two-tier evaluation pipeline
Evidence statusMeasured results reported
- Collect neologisms4,759 Chinese neologisms with reference definitions gathered from Moegirlpedia, Wikiversity, and other sources, sorted into 5 categories and 9 subcategories by formation mechanism
- Tier 1: Define itModel sees only the neologism and must generate a definition, scored by an LLM judge against the reference definition
- Tier 2: Manipulate itModel must either restore the exact source form (for phonetic substitution and character decomposition types) or pick the right answer from four multiple-choice options (for other types)
- Recognition-manipulation gap analysisCases where Tier 1 was correct but Tier 2 failed are isolated to reveal the pattern of substituting a paraphrase for the true source form
- Few-shot recovery test1,058 items that three frontier models all failed at zero-shot are retested with 1-3 example sentences to see how much difficulty is due to lack of context
What they did
- The team collected 4,759 Chinese internet neologisms with reference definitions from sources like Moegirlpedia and Wikiversity, organizing them into 5 top-level categories and 9 subcategories based on the linguistic mechanism behind each (phonetic substitution, character decomposition, foreign borrowing, etc.).
- Evaluation uses two tiers: Tier 1 asks a model to define a neologism from scratch, while Tier 2 asks the model to either restore the exact original form (for phonetic/decomposition types) or pick the correct answer from four choices (for other types).
- 18 models were tested at zero temperature (deterministic, most-confident output), spanning GPT-5.1, DeepSeek-V3/V3.2, Kimi-K2-Instruct/K2.5, the Qwen3 family, GLM-4 family, InternLM3, Gemma3 family, and Llama3.2 family.
- The best-performing model, Kimi-K2.5, reached only 67.74% on the definition task, well below the human evaluators' 85.63%. GPT-5.1, despite strong general benchmark performance, scored only 50.89%, lower than Chinese-focused models.
- A systematic 'recognition-manipulation gap' appeared: for neologisms formed by phonetic substitution or character decomposition, models that correctly explained the meaning in Tier 1 frequently failed to reproduce the exact source form in Tier 2, instead outputting a paraphrase or synonym.

| Subcategory | Full | Hard | Hard Rate |
|---|---|---|---|
| Pinyin Abbr. | 192 | 35 | 18.23% |
| English Abbr. | 58 | 7 | 12.07% |
| Chinese Homo. | 377 | 65 | 17.24% |
| Number Homo. | 29 | 7 | 24.14% |
| Lexical Neo. | 1161 | 298 | 25.67% |
| Semantic Neo. | 850 | 185 | 21.76% |
| Unconv. Char. | 23 | 3 | 13.04% |
| Formulaic. | 1667 | 347 | 20.82% |
| Foreign | 402 | 111 | 27.61% |
| Total | 4759 | 1058 | 22.23% |
| Model | Alphabetic | Homophonic | Lexical | Stylistics | Foreign | Avg.∗ | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pinyin | English | Chinese | Number | Lexical | Semantic | Unconv. Char. | Formulaic | |||
| GPT-5.1 | 47.40 | 74.14 | 48.81 | 44.83 | 49.10 | 54.94 | 21.74 | 52.67 | 42.54 | 50.89 |
| DeepSeek-V3 | 64.06 | 74.14 | 55.44 | 72.41 | 51.51 | 54.59 | 65.22 | 54.35 | 45.27 | 53.81 |
| DeepSeek-V3.2 | 58.85 | 79.31 | 60.21 | 72.41 | 54.61 | 55.65 | 69.57 | 54.17 | 49.00 | 55.26 |
| Kimi-K2-Instruct | 66.15 | 70.69 | 61.27 | 72.41 | 58.14 | 64.47 | 65.22 | 62.57 | 53.48 | 61.27 |
| Kimi-K2.5 | 76.56 | 79.31 | 70.82 | 72.41 | 66.06 | 67.06 | 82.61 | 67.55 | 61.69 | 67.74 |
| Qwen3-4B | 9.38 | 41.38 | 15.53 | 17.24 | 24.89 | 28.12 | 8.70 | 26.03 | 13.93 | 23.49 |
| Qwen3-4B-Instruct-2507 | 7.81 | 53.45 | 16.71 | 20.69 | 27.82 | 32.94 | 13.04 | 28.91 | 16.67 | 26.69 |
| Qwen3-8B | 17.71 | 53.45 | 18.04 | 27.59 | 31.35 | 33.65 | 4.55 | 32.03 | 18.16 | 29.40 |
| Qwen3-32B | 21.88 | 63.79 | 25.99 | 34.48 | 36.69 | 39.18 | 17.39 | 37.79 | 25.37 | 35.34 |
| GLM-4-9B-0414 | 18.75 | 48.28 | 20.16 | 20.69 | 29.46 | 34.00 | 8.70 | 29.03 | 21.39 | 28.35 |
| GLM-4-32B-0414 | 30.21 | 63.79 | 35.28 | 65.52 | 41.26 | 45.88 | 34.78 | 41.09 | 30.60 | 40.60 |
| InternLM3-8B-Instruct | 17.19 | 58.62 | 18.30 | 34.48 | 25.84 | 31.06 | 4.55 | 20.94 | 15.17 | 23.56 |
| Gemma3-4B-it | 5.21 | 37.93 | 7.69 | 3.45 | 17.23 | 19.29 | 0.00 | 16.68 | 10.45 | 15.68 |
| Gemma3-12B-it | 19.27 | 48.28 | 14.32 | 17.24 | 25.58 | 27.06 | 4.55 | 23.34 | 18.16 | 23.41 |
| Gemma3-27B-it | 26.04 | 63.79 | 21.22 | 13.79 | 31.52 | 30.00 | 4.55 | 29.57 | 19.40 | 28.66 |
| Llama3.2-1B-Instruct | 1.56 | 3.45 | 3.18 | 10.34 | 5.77 | 4.47 | 0.00 | 6.24 | 2.24 | 5.00 |
| Llama3.2-3B-Instruct | 5.21 | 24.14 | 3.71 | 3.45 | 5.77 | 8.71 | 0.00 | 8.64 | 6.22 | 7.33 |
| Llama3.2-11B-Instruct | 4.69 | 34.48 | 7.16 | 3.45 | 14.13 | 17.29 | 0.00 | 13.44 | 10.70 | 13.34 |
| Model | Alphabetic | Homophonic | Lexical | Stylistics | Foreign | Avg.∗ | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pinyin | English | Chinese† | Number† | Lexical | Semantic | Unconv. Char.† | Formulaic | |||
| GPT-5.1 | 72.40 | 98.28 | 56.76 | 72.41 | 73.47 | 96.46 | 22.73 | 68.77 | 86.36 | 75.70 |
| DeepSeek-V3 | 77.60 | 98.28 | 65.25 | 82.76 | 77.43 | 96.58 | 45.45 | 70.10 | 81.06 | 77.76 |
| DeepSeek-V3.2 | 76.04 | 94.83 | 64.46 | 82.76 | 78.12 | 95.75 | 54.55 | 64.20 | 81.31 | 75.61 |
| Kimi-K2-Instruct | 78.65 | 96.55 | 68.97 | 93.10 | 83.03 | 96.58 | 59.09 | 68.17 | 84.60 | 79.20 |
| Kimi-K2.5 | 80.21 | 96.55 | 68.17 | 79.31 | 82.00 | 97.76 | 59.09 | 75.39 | 88.64 | 81.94 |
| Qwen3-4B | 23.44 | 86.21 | 9.28 | 0.00 | 49.53 | 86.79 | 13.64 | 44.65 | 51.77 | 50.40 |
| Qwen3-4B-Instruct-2507 | 22.40 | 84.48 | 16.18 | 6.90 | 51.34 | 92.10 | 13.64 | 46.03 | 54.04 | 52.99 |
| Qwen3-8B | 30.21 | 86.21 | 19.10 | 3.45 | 52.63 | 90.21 | 4.55 | 49.22 | 58.59 | 54.97 |
| Qwen3-32B | 56.25 | 94.83 | 33.69 | 13.79 | 69.42 | 96.11 | 13.64 | 58.90 | 66.92 | 66.63 |
| GLM-4-9B-0414 | 38.54 | 89.66 | 26.26 | 10.34 | 61.41 | 91.16 | 4.55 | 37.67 | 49.75 | 53.47 |
| GLM-4-32B-0414 | 65.62 | 93.10 | 37.93 | 51.72 | 70.03 | 93.63 | 18.18 | 48.98 | 66.67 | 63.79 |
| InternLM3-8B-Instruct | 16.15 | 79.31 | 15.12 | 0.00 | 53.32 | 92.69 | 0.00 | 59.33 | 46.97 | 58.60 |
| Gemma3-4B-it | 30.73 | 86.21 | 9.55 | 3.45 | 39.79 | 91.04 | 4.55 | 29.60 | 44.95 | 43.22 |
| Gemma3-12B-it | 39.58 | 89.66 | 21.22 | 10.34 | 58.66 | 92.69 | 4.55 | 44.65 | 57.07 | 55.78 |
| Gemma3-27B-it | 41.67 | 94.83 | 29.44 | 20.69 | 66.32 | 95.64 | 4.55 | 48.13 | 62.37 | 60.71 |
| Llama3.2-1B-Instruct | 24.48 | 48.28 | 0.27 | 0.00 | 25.40 | 31.13 | 0.00 | 4.03 | 24.24 | 16.81 |
| Llama3.2-3B-Instruct | 26.04 | 70.69 | 0.27 | 0.00 | 36.61 | 54.72 | 0.00 | 5.60 | 28.79 | 25.03 |
| Llama3.2-11B-Instruct | 24.48 | 74.14 | 6.90 | 6.90 | 39.28 | 83.84 | 4.55 | 36.04 | 45.96 | 43.57 |
| Subcategory | Human | Kimi-K2.5 |
|---|---|---|
| Pinyin Abbr. | 91.00% | 76.56% |
| English Abbr. | 84.48% | 79.31% |
| Chinese Homo. | 94.00% | 70.82% |
| Number Homo. | 82.76% | 72.41% |
| Lexical Neo. | 87.00% | 66.06% |
| Semantic Neo. | 83.00% | 67.06% |
| Unconv. Chars. | 95.65% | 82.61% |
| Formulaic | 78.00% | 67.55% |
| Foreign | 80.00% | 61.69% |
| Overall | 85.63% | 67.74% |

| Categories | Subcategories | Explanations | Examples |
|---|---|---|---|
| 1. Alphabetic Abbreviations | Pinyin Abbr. | Expressions formed by taking the initial letters of the Pinyin transliteration of Chinese phrases to create abbreviated forms. | yyds from “永远的神” (yong yuan de shen), which means forever a legend. |
| English Abbr. | Expressions formed by taking the initial letters of the English phrases to create abbreviated forms. | ACG from “Anime, Comics, Games”. | |
| 2. Homophonic Expressions | Number Homo. | Expressions that use numbers to convey meaning based on phonetic similarity. | “拜拜咯” becomes “886” as they have similar pronunciations in Chinese, which means bye bye. |
| Chinese Homo. | Expressions that replace standard Chinese characters with characters of similar pronunciation. | “与你无关” (yu ni wu guan) becomes “雨女无瓜” (yu nv wu gua), which means none of your business. | |
| 3. Lexical Innovations | Lexical Neo. | Newly coined expressions that introduce novel lexical items. | “键盘侠”, which means keyboard warrior, a person who aggressively criticizes or attacks others online without engaging in real-world action. |
| Semantic Neo. | Existing words that acquire new meanings in emerging usage. | “养鱼”, literal meaning: fish farming; new meaning: to maintain multiple romantic prospects simultaneously, often by keeping others as backups. | |
| 4. Stylistic and Pragmatic Expressions | Unconv. Chars. | Expressions that employ rare, archaic, or visually stylized Chinese characters in non-standard ways to convey meaning or achieve stylistic and expressive effects. | “彳亍”, an expression composed of two archaic characters “彳” and “亍” to form “行”, which means okay. |
| Formulaic | Fixed or semi-fixed expressions derived from specific memes or events, which are reused across contexts to convey particular pragmatic meanings. | “没活了可以咬打火机”, a meme-based expression originating from a viral livestream clip, used to humorously suggest that someone has run out of ideas or content. | |
| 5. Foreign Borrowings | Foreign∗ | Expressions derived from foreign-language sources through phonetic reinterpretation or direct lexical adoptiont. | “逮虾户” from “Déjà Vu” (phonetic, via Japanese anime Initial D), used for high-speed driving scenes; “文化祭”, which means school cultural festival is a direct loanword from Japanese. |
| Subcategory | Task | Example | Answer |
|---|---|---|---|
| Pinyin Abbr. | Cloze | Context: 这家楼下新开的烧烤摊味道真的太绝了,朋友吃完直接说它是夜宵界的(____)。 English: The new BBQ stall that just opened downstairs is absolutely amazing, my friend straight-up called it the (____) of the late-night food scene. Options: (A) cxk (B) xswl (C) yysy (D) yyds | (D) yyds (“forever legendary”) |
| English Abbr. | Cloze | Context: 他平时看番、追漫画还爱玩手游,朋友圈里一看就是个(____)重度爱好者。 English: He watches anime, follows manga, and loves playing mobile games, and you can tell he’s a hardcore (____) enthusiast from his social media. Options: (A) KPI (B) ACG (C) PDF (D) CPU | (B) ACG (Anime, Comics, Games) |
| Chinese Homo.† | Restoration | Neologism: 雨女无瓜 Example context: 这件事是他们部门内部的安排,雨女无瓜,别再到处打听了。 English: This is an internal arrangement within their department, none of your business, so stop asking around about it. | 与你无关 |
| Number Homo.† | Restoration | Neologism: 886 Example context: "时间不早了, 我先撤, 886, 明天见。 English: It’s getting late, I’m gonna head out. Bye-bye, see you tomorrow! | 拜拜咯 |
| Lexical Neo. | Cloze | Context: 事情刚上热搜、真相还没出来,他就在评论区里对当事人冷嘲热讽、指点江山,活脱脱一个(____)。 English: The story had just started trending and the truth hadn’t even come out yet, but there he was in the comments section mocking the people involved and pontificating about everything, a textbook (____). Options: (A) 程序员 (B) 黑客 (C) 杠精 (D) 键盘侠 | (D) 键盘侠 (keyboard warrior) |
| Semantic Neo. | Semantic Discrimination | Target word: 养鱼 Sentences: (A) 他在院子里挖了个小池子,专门养鱼和种睡莲。(fish farming) (B) 爷爷退休后最大的乐趣就是养鱼,每天一早都去喂食。(fish farming) (C) 她总是同时和好几个人保持暧昧,朋友都说她是在养鱼。 (D) 这家农场除了养鸡养鸭,还养鱼,逢年过节卖得特别好。(fish farming) | (C) maintaining multiple romantic relationships |
| Unconv. Chars.† | Restoration | Neologism: 彳亍 | 行 (“okay”) |
| Formulaic | Scenario Matching | Scenario: 家庭聚会上,62岁的老周本来在陪孙子搭积木,结果看到院里积雪够厚,突然兴奋地拉着几个晚辈去堆“巨型雪怪”,还非要拿胡萝卜做獠牙。被老伴笑着数落一把年纪还这么能闹,他却越玩越起劲。 English: At the family gathering, 62-year-old Zhou had been building blocks with his grandson, but when he saw how thick the snow had piled up in the yard, he suddenly got excited and dragged a few of the younger family members out to build a “giant snow monster”, and he insisted on using carrots for the fangs. His wife laughingly scolded him for being so rowdy at his age, but the more he played, the more into it he got. Options: (A) 老夫聊发少年狂 (B) 人间清醒 (C) 差生文具多 (D) 男人至死是少年 | (D) 男人至死是少年 (Men never really grow up) |
| Foreign† | Source Identification | Neologism: 文化祭 Options: (A) 它是从日语“文化祭(ぶんかさい)”转译扩展的词,“祭”在日语里有集中展示、主题活动之意,“文化祭”本质上是对日本校园文化展示周这一活动与说法的概括吸收。(Expanded translation) (B) 它是从日语“文化祭(ぶんかさい)”直接借入的词,“祭”在日语里有节庆、庆典之意,“文化祭”本质上是对日本校园“学园文化祭”这一制度与说法的直接沿用。(Direct loanword) (C) 它是从日语“文化祭(ぶんかさい)”类推形成的词,“祭”在日语里有公开活动、校内集会之意,“文化祭”本质上是对日本校园“社团发表会”这一形式与说法的重新命名。(Analogical formation) (D) 它是从日语“文化祭(ぶんかさい)”借形改义的词,“祭”在日语里有热闹活动、集体庆祝之意,“文化祭”本质上是对日本校园“校园开放日”这一安排与说法的本地化套用。(Borrowed form, new meaning) | (B) Direct loanword |
| Model | Subcategory | Cell B Failure Mode (%) | |||
|---|---|---|---|---|---|
| Copy Verbatim | Explanation/Paraphrase | Wrong Restoration | Empty/Other | ||
| Kimi-K2.5 | Chinese Homo. | 26.6% | 38.0% | 32.9% | 2.5% |
| Number Homo. | 0.0% | 60.0% | 40.0% | 0.0% | |
| Unconv. Chars. | 0.0% | 71.4% | 28.6% | 0.0% | |
| DeepSeek-V3.2 | Chinese Homo. | 25.5% | 35.3% | 35.3% | 3.9% |
| Number Homo. | 0.0% | 100% | 0.0% | 0.0% | |
| Unconv. Chars. | 0.0% | 62.5% | 37.5% | 0.0% | |
| Qwen3-32B | Chinese Homo. | 13.6% | 22.7% | 54.5% | 9.1% |
| Number Homo. | 0.0% | 42.9% | 57.1% | 0.0% | |
| Unconv. Chars. | 50.0% | 50.0% | 0.0% | 0.0% | |
| GLM-4-32B | Chinese Homo. | 27.3% | 38.2% | 32.7% | 1.8% |
| Number Homo. | 0.0% | 66.7% | 33.3% | 0.0% | |
| Unconv. Chars. | 33.3% | 33.3% | 33.3% | 0.0% | |
| Gemma3-27B | Chinese Homo. | 22.2% | 44.4% | 26.7% | 6.7% |
| Number Homo. | 0.0% | 100% | 0.0% | 0.0% | |
| Unconv. Chars. | 0.0% | 0.0% | 100% | 0.0% |
| Subcategory | Model | Neologism | Gold | Tier 1 Model Output (correct) | Tier 2 Output (wrong) |
|---|---|---|---|---|---|
| Chinese Homo. | Kimi-K2.5 | v我50 | 微我50 | 网络流行语,指通过微信转账50元人民币。源自肯德基“疯狂星期四”促销活动… Referring to transferring 50 yuan RMB via WeChat. It originated from KFC’s ‘Crazy Thursday’ promotional campaign. | 微信我50 (additional characters) |
| Chinese Homo. | Kimi-K2.5 | 石乐志 | 失了智 | 网络流行语,“失了智”的谐音,指失去理智、做出愚蠢或不理智的行为… A homophonic pun on “失了智” (lost one’s wits), meaning to lose one’s reason and act in a stupid or irrational way. | 失智 (dementia) |
| Chinese Homo. | DeepSeek-V3.2 | 河蟹 | 和谐 | 网络用语中常被用作“删除”或“屏蔽”的委婉说法,因“和谐”与“河蟹”谐音,后者常被用来讽刺内容审查。 Often used as a euphemism for ‘deleting’ or ‘blocking’ content. Because 和谐 (harmony) is a homophone of 河蟹 (river crab). | 删除 (delete) |
| Chinese Homo. | DeepSeek-V3.2 | 大清药丸 | 大清要完 | 大清药丸是网络流行语,源自“大清要完”的谐音,实际用来调侃某事物或局势即将崩溃或失败,带有讽刺或幽默意味。 A homophonic pun on ‘大清要完’ (the Qing dynasty is finished). It’s used to jokingly suggest that something or some situation is on the verge of collapse or failure, with a satirical or humorous undertone. | 大清朝要完 (additional characters) |
| Number Homo. | Kimi-K2.5 | 3q | thank you | 英文“thank you”的谐音数字缩写,“3”对应“th”音、“q”对应“you”音。 ‘3Q’ is a homophonic abbreviation of the English ‘thank you’, using numbers and letters: ‘3’ represents the ‘th’ sound and ‘q’ represents the ‘you’ sound. | 谢谢 (Chinese expression for thank you) |
| Number Homo. | DeepSeek-V3.2 | 666 | 溜溜溜 | 表示非常厉害、令人惊叹的意思,源自数字6的谐音“溜”。 ‘666’ means something is awesome or impressive. It derives from the number 6, a homophone of ‘溜’ (‘slick/skilled’) | 牛牛牛 (semantically equivalent) |
| Unconv. Chars. | Kimi-K2.5 | 占戈哥欠走已 | 战歌起 | 将“战歌起”三字拆解为偏旁部首(战=占+戈,歌=哥+欠,起=走+己/已)的隐晦写法… A cryptic way of writing ‘战歌起’ (cue the battle anthem) by splitting each of the three characters into their radical components. | 战歌响起 (additional characters) |
| Unconv. Chars. | Kimi-K2.5 | 口区 | 呕 | 将汉字“呕”拆分为“口”和“区”二字输入的网络黑话… The Chinese character ‘呕’ (to vomit/retch) is broken apart and typed as its two components, ‘口’ and ‘区’. | 呕吐 (additional characters) |
| Unconv. Chars. | DeepSeek-V3.2 | 彳亍口巴 | 行吧 | 网络用语,由“彳亍”和“口巴”拼接而成。其中“彳亍”是“行”的拆分形式,“口巴”是“吧”的拆分形式。 ‘彳亍’ is the disassembled form of ‘行’ and ‘口巴’ is the disassembled form of ‘吧’. | 好吧 (another expression for okay) |
| Unconv. Chars. | DeepSeek-V3.2 | 占戈土也 | 战地 | “占戈土也”是“战地”的拆分写法… ‘占戈土也’ is the disassembled-character way of ‘战地’(‘battlefield’) | 战争 (war) |
Findings
- On the Tier 1 definition task, the top model Kimi-K2.5 reached 67.74% overall accuracy, while most open-source models scored below 40%; human evaluators averaged 85.63%.
- GPT-5.1, despite frontier status on general benchmarks, scored only 50.89% overall on CNeo-Bench, below Chinese-focused models, with especially large gaps on Pinyin abbreviations (47.40% vs. Kimi-K2.5's 76.56%) and character decomposition (21.74% vs. 82.61%).
- Every model scored higher on Tier 2 than its own Tier 1 score, but on the three open-ended restoration subcategories (two homophonic types plus character decomposition), 24.2% to 57.1% of items correctly described in Tier 1 failed restoration in Tier 2.
- Analyzing cases where Tier 1 was correct but Tier 2 was wrong, the dominant failure mode across models was substituting a paraphrase or near-synonym instead of the exact source form.
- On 1,058 hard items that GPT-5.1, DeepSeek-V3.2, and Kimi-K2.5 all failed at zero-shot, giving just 1 example recovered 37-50% of cases across models, and 3 examples recovered 53-67%, but 33-47% remained unrecovered even at 3 examples.
Where it can be used
- Can serve as a diagnostic tool for checking how well chatbots, translation systems, or content moderation tools handle Chinese internet slang and memes.
- Can help pinpoint where a model knows the meaning of a neologism but can't manipulate its underlying form, informing prompt design or example-based strategies for neologism-related tasks.
- Provides a reference for which linguistic mechanisms (phonetic substitution, character decomposition) are weakest, useful when designing targeted training data or fine-tuning approaches.
Limits and open work
- Tier 1 scoring relies on an LLM judge checking semantic adequacy, so a definition can be marked correct without reproducing the actual source form.
- Two subcategories, Number Homophone (29 items) and Unconventional Character decomposition (23 items), have very few items, so per-model comparisons on these should be read as indicative rather than definitive.
- The study focuses on diagnosing the gap rather than fixing it; mitigation strategies are left to future work.
- Data was collected only through January 2026, so newer neologisms emerging afterward are not covered.
- The 33-47% of hard samples that remain unrecovered even with 3-shot prompting are not further explained or resolved in this work.
Why it matters
Because neologisms sit at the edge of what a model likely saw during training, how well a model handles them is a useful probe of genuine linguistic flexibility rather than memorized surface performance. This shows that scoring well on 'what does this mean' tasks doesn't guarantee a model actually understands the underlying mechanism that produced the expression.
Terms in this paper
- Neologism · A newly coined expression or an existing word/phrase used with a new meaning
- Phonetic substitution / Homophonic expression · Replacing a phrase with numbers or characters that sound similar but are unrelated in meaning, e.g., 886 sounds like '拜拜咯' (bye-bye)
- Unconventional character decomposition · Splitting a single Chinese character into its visual components and writing them separately, e.g., 行 (okay) split into 彳 and 亍
- Recognition-manipulation gap · The pattern where a model correctly describes a neologism's meaning but fails to reproduce the exact original source form it came from
- LLM-as-judge · Using another language model to automatically score whether a model's generated answer matches the correct meaning
Original abstract (English)
Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench, a benchmark of 4,759 such neologisms with reference definitions, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression. CNeo-Bench is paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whether it can operate on its underlying mechanism. Evaluating 18 LLMs, we find that Chinese neologisms remain an open challenge; most models fall below 40\% on definition generation, and on several subcategories a systematic recognition-manipulation gap emerges: models describe neologisms correctly but, in source-form restoration tasks, substitute a semantic equivalent (paraphrase) for the source form rather than producing the source form itself. A few-shot analysis on 1,058 hard items shows that in-context examples can solve many difficult cases, but leave a noticeable portion of errors remaining, indicating challenges beyond prompting alone can address.
Read on arXivLatest papers
- FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial OutcomesA dataset that finally teaches AI what biology, chemistry, and physics peer reviewers actually argue about, not just CS reviewers
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness EvolutionAn AI system that writes a custom 'operating scaffold' for other AI agents on the spot, for every new task
- The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling PipelineAI language models still charge a hidden 'dialect tax' on AAVE and other non-standard English at every stage, not just tokenization
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal BayesiansA math model shows that even a perfectly rational person can be talked into delusion by a chatbot that keeps agreeing with them
- Autonomous Mathematical Discovery in an Open-World Multi-Agent EnvironmentAI agents from different companies self-organized in an open-world simulation and produced new results on five math problems, with no one directing them
- Automata from Agent Traces: Failure and Next-Step PredictionCompressing thousands of LLM agent execution logs into one tiny 7-to-43-state machine that predicts both the next action and eventual failure
- MARS: Multi-Specialist LLM Relay System for Competitive ProgrammingLetting topic-specialist AIs take turns fixing code beats one generalist coder on programming contest problems
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared WorkspaceLetting multiple AI coding agents share one workspace and coordinate in real time beats running them one-by-one or in uncoordinated parallel
Latest from METAL LAB
- 1,050 Japanese Local Governments Adopt OpenAI-Based QommonsAI
- Runway unveils Solaris, which redraws the screen with every click
- Google Research unveils TimesFM-3, a multivariate time-series forecasting model
- OpenClaw 2.0 automates setup, adds shared cloud sessions
- Pentagon adds ChatGPT and Grok to AI portal, leaves Claude out
Figures: Kaiyan Zhao et al., arXiv:2608.28053, arxiv-nonexclusive
