METAL LAB

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

arXiv:2608.280532026-08-31

886、yyds、彳亍这类中文新造词,大模型往往能解释是什么意思,却还原不出它原本的样子

研究团队构建了CNeo-Bench,收录4,759个中文网络新词(如用886谐音代表拜拜、用yyds取自拼音首字母、用彳亍拆解行字),按构词机制分为5大类9小类。团队用两层评测测试了包括GPT-5.1、DeepSeek、Kimi在内的18个大模型,发现在仅要求解释含义的任务上大多数模型准确率不到40%,而在要求还原新词原始形式的任务上,不少模型即便正确解释了含义,也会用近义改写代替原始形式。给模型看1到3个示例后,大约一半难题能被解决,但即便给3个示例,仍有33%到47%的错误无法解决。

METAL LAB 解读图

CNeo-Bench两层评测流程

证据状态已报告实测结果

  1. 收集新词从萌娘百科、维基学院等来源收集4,759个中文新词及参考释义,按构词机制分为5大类9小类
  2. 第一层:解释含义模型仅看到新词本身,需要生成释义,由大模型评判员与参考释义比对打分
  3. 第二层:操作形式模型需要还原出新词的原始形式(谐音、拆字类)或从四个选项中选出正确答案(缩写、语义新词等类),按精确匹配或选项匹配打分
  4. 识别-操作差距分析挑出第一层答对但第二层答错的案例,发现模型倾向于用改写代替原始形式的失败模式
  5. 少样本恢复实验对三个模型共同答错的1,058个难题,给出1到3个示例句子重新测试,观察恢复率和仍未解决的残留错误
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 团队从萌娘百科、维基学院等来源收集了4,759个中文新词及其参考释义,按构词机制分为5个大类和9个小类,涵盖谐音替代、汉字拆解、拼音缩写、外来借词等类型。
  2. 评测分两层:第一层要求模型仅凭新词本身给出释义;第二层要求模型直接还原新词背后的原始形式,或从四个选项中选出正确答案。
  3. 研究以温度0(只输出最确定的答案)测试了18个模型,包括GPT-5.1、DeepSeek-V3和V3.2、Kimi-K2-Instruct和K2.5、Qwen3系列、GLM-4系列、InternLM3、Gemma3系列、Llama3.2系列。
  4. 表现最好的Kimi-K2.5在释义任务上仅达到67.74%的准确率,低于人类评测者85.63%的平均水平;而在通用基准上表现强劲的GPT-5.1在这里只有50.89%,低于多个中文优化模型。
  5. 在谐音和拆字这类需要还原原始形式的任务上,出现了系统性的'识别-操作差距':模型在第一层正确解释了新词含义,却在第二层未能还原原始形式,而是给出一个意思相近的改写。
Figure 1: Categories and tasks in CNeo-Bench.
Figure 1: Categories and tasks in CNeo-Bench.
Table 1: Statistics for CNeo-Bench. “Full” is the total number of items in each subcategory; “Hard” is the number of items on which GPT-5.1, DeepSeek-V3.2, and Kimi-K2.5 jointly fail under definition generation.
SubcategoryFullHardHard Rate
Pinyin Abbr.1923518.23%
English Abbr.58712.07%
Chinese Homo.3776517.24%
Number Homo.29724.14%
Lexical Neo.116129825.67%
Semantic Neo.85018521.76%
Unconv. Char.23313.04%
Formulaic.166734720.82%
Foreign40211127.61%
Total4759105822.23%
Figure 2: Contingency analysis of Tier 1 (definition generation) vs Tier 2 (category-specific) outcomes for five representative models on three open-ended subcategories. Cells A/B/C/D denote (T1✓, T2✓), (T1✓, T2×), (T1×, T2✓), (T1×, T2×) respectively. The highlighted B/(A+B) column reports the operation failure rate (the fraction of correctly described neologisms on which restoration fails) directly quantifying the recognition-manipulation gap.
Figure 2: Contingency analysis of Tier 1 (definition generation) vs Tier 2 (category-specific) outcomes for five representative models on three open-ended subcategories. Cells A/B/C/D denote (T1✓, T2✓), (T1✓, T2×), (T1×, T2✓), (T1×, T2×) respectively. The highlighted B/(A+B) column reports the operation failure rate (the fraction of correctly described neologisms on which restoration fails) directly quantifying the recognition-manipulation gap.
Table 2: Full definition generation task results. Scores are reported as accuracy (%). Bold indicates the best score in each column; underline indicates the second best. Avg.∗ is computed as micro-average across all 4,759 instances.
ModelAlphabeticHomophonicLexicalStylisticsForeignAvg.∗
PinyinEnglishChineseNumberLexicalSemanticUnconv. Char.Formulaic
GPT-5.147.4074.1448.8144.8349.1054.9421.7452.6742.5450.89
DeepSeek-V364.0674.1455.4472.4151.5154.5965.2254.3545.2753.81
DeepSeek-V3.258.8579.3160.2172.4154.6155.6569.5754.1749.0055.26
Kimi-K2-Instruct66.1570.6961.2772.4158.1464.4765.2262.5753.4861.27
Kimi-K2.576.5679.3170.8272.4166.0667.0682.6167.5561.6967.74
Qwen3-4B9.3841.3815.5317.2424.8928.128.7026.0313.9323.49
Qwen3-4B-Instruct-25077.8153.4516.7120.6927.8232.9413.0428.9116.6726.69
Qwen3-8B17.7153.4518.0427.5931.3533.654.5532.0318.1629.40
Qwen3-32B21.8863.7925.9934.4836.6939.1817.3937.7925.3735.34
GLM-4-9B-041418.7548.2820.1620.6929.4634.008.7029.0321.3928.35
GLM-4-32B-041430.2163.7935.2865.5241.2645.8834.7841.0930.6040.60
InternLM3-8B-Instruct17.1958.6218.3034.4825.8431.064.5520.9415.1723.56
Gemma3-4B-it5.2137.937.693.4517.2319.290.0016.6810.4515.68
Gemma3-12B-it19.2748.2814.3217.2425.5827.064.5523.3418.1623.41
Gemma3-27B-it26.0463.7921.2213.7931.5230.004.5529.5719.4028.66
Llama3.2-1B-Instruct1.563.453.1810.345.774.470.006.242.245.00
Llama3.2-3B-Instruct5.2124.143.713.455.778.710.008.646.227.33
Llama3.2-11B-Instruct4.6934.487.163.4514.1317.290.0013.4410.7013.34
Figure 3: Few-shot recovery on 1,058 hard samples. Each bar shows the number of hard samples recovered (i.e., judged correct under the few-shot setting) under 1-shot, 2-shot, and 3-shot prompting. Bars are stacked by subcategory to show the distribution of recoveries across the nine subcategories.
Figure 3: Few-shot recovery on 1,058 hard samples. Each bar shows the number of hard samples recovered (i.e., judged correct under the few-shot setting) under 1-shot, 2-shot, and 3-shot prompting. Bars are stacked by subcategory to show the distribution of recoveries across the nine subcategories.
Table 3: Category-specific task results. Scores are reported as accuracy (%). Bold indicates the best score in each column; underline indicates the second best. Tasks with† use open-ended restoration; all other tasks use multiple-choice format. Avg.∗ is computed as micro-average across all 4,759 instances.
ModelAlphabeticHomophonicLexicalStylisticsForeignAvg.∗
PinyinEnglishChinese†Number†LexicalSemanticUnconv. Char.†Formulaic
GPT-5.172.4098.2856.7672.4173.4796.4622.7368.7786.3675.70
DeepSeek-V377.6098.2865.2582.7677.4396.5845.4570.1081.0677.76
DeepSeek-V3.276.0494.8364.4682.7678.1295.7554.5564.2081.3175.61
Kimi-K2-Instruct78.6596.5568.9793.1083.0396.5859.0968.1784.6079.20
Kimi-K2.580.2196.5568.1779.3182.0097.7659.0975.3988.6481.94
Qwen3-4B23.4486.219.280.0049.5386.7913.6444.6551.7750.40
Qwen3-4B-Instruct-250722.4084.4816.186.9051.3492.1013.6446.0354.0452.99
Qwen3-8B30.2186.2119.103.4552.6390.214.5549.2258.5954.97
Qwen3-32B56.2594.8333.6913.7969.4296.1113.6458.9066.9266.63
GLM-4-9B-041438.5489.6626.2610.3461.4191.164.5537.6749.7553.47
GLM-4-32B-041465.6293.1037.9351.7270.0393.6318.1848.9866.6763.79
InternLM3-8B-Instruct16.1579.3115.120.0053.3292.690.0059.3346.9758.60
Gemma3-4B-it30.7386.219.553.4539.7991.044.5529.6044.9543.22
Gemma3-12B-it39.5889.6621.2210.3458.6692.694.5544.6557.0755.78
Gemma3-27B-it41.6794.8329.4420.6966.3295.644.5548.1362.3760.71
Llama3.2-1B-Instruct24.4848.280.270.0025.4031.130.004.0324.2416.81
Llama3.2-3B-Instruct26.0470.690.270.0036.6154.720.005.6028.7925.03
Llama3.2-11B-Instruct24.4874.146.906.9039.2883.844.5536.0445.9643.57
Figure 4: Accuracy on English-based vs Pinyin-based alphabetic abbreviations in the definition generation task. English-based abbreviations are substantially easier than Pinyin-based ones for all evaluated models.
Figure 4: Accuracy on English-based vs Pinyin-based alphabetic abbreviations in the definition generation task. English-based abbreviations are substantially easier than Pinyin-based ones for all evaluated models.
Table 4: Human performance on a stratified sample of CNeo-Bench Tier 1 definition generation. We leave Kimi-K2.5, the strongest model we evaluate, for comparison.
SubcategoryHumanKimi-K2.5
Pinyin Abbr.91.00%76.56%
English Abbr.84.48%79.31%
Chinese Homo.94.00%70.82%
Number Homo.82.76%72.41%
Lexical Neo.87.00%66.06%
Semantic Neo.83.00%67.06%
Unconv. Chars.95.65%82.61%
Formulaic78.00%67.55%
Foreign80.00%61.69%
Overall85.63%67.74%
Figure 5: Tier 2 minus Tier 1 accuracy (percentage points) across 18 models and 9 subcategories. Positive gaps (red) indicate Tier 2 exceeds Tier 1; negative gaps (blue) indicate Tier 2 underperforms Tier 1. Subcategories marked with † use open-ended restoration; others use multiple-choice.
Figure 5: Tier 2 minus Tier 1 accuracy (percentage points) across 18 models and 9 subcategories. Positive gaps (red) indicate Tier 2 exceeds Tier 1; negative gaps (blue) indicate Tier 2 underperforms Tier 1. Subcategories marked with † use open-ended restoration; others use multiple-choice.
Table 5: Mechanism-driven taxonomy of Chinese neologisms in CNeo-Bench, with five top-level categories and nine subcategories. ∗Foreign Borrowings has no further subcategorization and is treated as a single class.
CategoriesSubcategoriesExplanationsExamples
1. Alphabetic AbbreviationsPinyin Abbr.Expressions formed by taking the initial letters of the Pinyin transliteration of Chinese phrases to create abbreviated forms.yyds from “永远的神” (yong yuan de shen), which means forever a legend.
English Abbr.Expressions formed by taking the initial letters of the English phrases to create abbreviated forms.ACG from “Anime, Comics, Games”.
2. Homophonic ExpressionsNumber Homo.Expressions that use numbers to convey meaning based on phonetic similarity.“拜拜咯” becomes “886” as they have similar pronunciations in Chinese, which means bye bye.
Chinese Homo.Expressions that replace standard Chinese characters with characters of similar pronunciation.“与你无关” (yu ni wu guan) becomes “雨女无瓜” (yu nv wu gua), which means none of your business.
3. Lexical InnovationsLexical Neo.Newly coined expressions that introduce novel lexical items.“键盘侠”, which means keyboard warrior, a person who aggressively criticizes or attacks others online without engaging in real-world action.
Semantic Neo.Existing words that acquire new meanings in emerging usage.“养鱼”, literal meaning: fish farming; new meaning: to maintain multiple romantic prospects simultaneously, often by keeping others as backups.
4. Stylistic and Pragmatic ExpressionsUnconv. Chars.Expressions that employ rare, archaic, or visually stylized Chinese characters in non-standard ways to convey meaning or achieve stylistic and expressive effects.“彳亍”, an expression composed of two archaic characters “彳” and “亍” to form “行”, which means okay.
FormulaicFixed or semi-fixed expressions derived from specific memes or events, which are reused across contexts to convey particular pragmatic meanings.“没活了可以咬打火机”, a meme-based expression originating from a viral livestream clip, used to humorously suggest that someone has run out of ideas or content.
5. Foreign BorrowingsForeign∗Expressions derived from foreign-language sources through phonetic reinterpretation or direct lexical adoptiont.“逮虾户” from “Déjà Vu” (phonetic, via Japanese anime Initial D), used for high-speed driving scenes; “文化祭”, which means school cultural festival is a direct loanword from Japanese.
Table 6: Representative examples for each Tier 2 tasks in CNeo-Bench. Subcategories marked with † require the model to complete open-ended restoration. English references are not included in our dataset.
SubcategoryTaskExampleAnswer
Pinyin Abbr.ClozeContext: 这家楼下新开的烧烤摊味道真的太绝了,朋友吃完直接说它是夜宵界的(____)。 English: The new BBQ stall that just opened downstairs is absolutely amazing, my friend straight-up called it the (____) of the late-night food scene. Options: (A) cxk (B) xswl (C) yysy (D) yyds(D) yyds (“forever legendary”)
English Abbr.ClozeContext: 他平时看番、追漫画还爱玩手游,朋友圈里一看就是个(____)重度爱好者。 English: He watches anime, follows manga, and loves playing mobile games, and you can tell he’s a hardcore (____) enthusiast from his social media. Options: (A) KPI (B) ACG (C) PDF (D) CPU(B) ACG (Anime, Comics, Games)
Chinese Homo.†RestorationNeologism: 雨女无瓜 Example context: 这件事是他们部门内部的安排,雨女无瓜,别再到处打听了。 English: This is an internal arrangement within their department, none of your business, so stop asking around about it.与你无关
Number Homo.†RestorationNeologism: 886 Example context: "时间不早了, 我先撤, 886, 明天见。 English: It’s getting late, I’m gonna head out. Bye-bye, see you tomorrow!拜拜咯
Lexical Neo.ClozeContext: 事情刚上热搜、真相还没出来,他就在评论区里对当事人冷嘲热讽、指点江山,活脱脱一个(____)。 English: The story had just started trending and the truth hadn’t even come out yet, but there he was in the comments section mocking the people involved and pontificating about everything, a textbook (____). Options: (A) 程序员 (B) 黑客 (C) 杠精 (D) 键盘侠(D) 键盘侠 (keyboard warrior)
Semantic Neo.Semantic DiscriminationTarget word: 养鱼 Sentences: (A) 他在院子里挖了个小池子,专门养鱼和种睡莲。(fish farming) (B) 爷爷退休后最大的乐趣就是养鱼,每天一早都去喂食。(fish farming) (C) 她总是同时和好几个人保持暧昧,朋友都说她是在养鱼。 (D) 这家农场除了养鸡养鸭,还养鱼,逢年过节卖得特别好。(fish farming)(C) maintaining multiple romantic relationships
Unconv. Chars.†RestorationNeologism: 彳亍行 (“okay”)
FormulaicScenario MatchingScenario: 家庭聚会上,62岁的老周本来在陪孙子搭积木,结果看到院里积雪够厚,突然兴奋地拉着几个晚辈去堆“巨型雪怪”,还非要拿胡萝卜做獠牙。被老伴笑着数落一把年纪还这么能闹,他却越玩越起劲。 English: At the family gathering, 62-year-old Zhou had been building blocks with his grandson, but when he saw how thick the snow had piled up in the yard, he suddenly got excited and dragged a few of the younger family members out to build a “giant snow monster”, and he insisted on using carrots for the fangs. His wife laughingly scolded him for being so rowdy at his age, but the more he played, the more into it he got. Options: (A) 老夫聊发少年狂 (B) 人间清醒 (C) 差生文具多 (D) 男人至死是少年(D) 男人至死是少年 (Men never really grow up)
Foreign†Source IdentificationNeologism: 文化祭 Options: (A) 它是从日语“文化祭(ぶんかさい)”转译扩展的词,“祭”在日语里有集中展示、主题活动之意,“文化祭”本质上是对日本校园文化展示周这一活动与说法的概括吸收。(Expanded translation) (B) 它是从日语“文化祭(ぶんかさい)”直接借入的词,“祭”在日语里有节庆、庆典之意,“文化祭”本质上是对日本校园“学园文化祭”这一制度与说法的直接沿用。(Direct loanword) (C) 它是从日语“文化祭(ぶんかさい)”类推形成的词,“祭”在日语里有公开活动、校内集会之意,“文化祭”本质上是对日本校园“社团发表会”这一形式与说法的重新命名。(Analogical formation) (D) 它是从日语“文化祭(ぶんかさい)”借形改义的词,“祭”在日语里有热闹活动、集体庆祝之意,“文化祭”本质上是对日本校园“校园开放日”这一安排与说法的本地化套用。(Borrowed form, new meaning)(B) Direct loanword
Table 7: Breakdown of Cell B failure modes (Tier 1 correct, Tier 2 wrong) across five models and three open-ended subcategories. Each row sums to 100%. Copy Verbatim: model outputs the neologism itself. Explanation/Paraphrase: model outputs a semantic gloss or near-synonym of the source form. Wrong Restoration: model outputs a different Chinese expression that is neither the neologism nor a paraphrase. Empty/Other: empty output or unparseable.
ModelSubcategoryCell B Failure Mode (%)
Copy VerbatimExplanation/ParaphraseWrong RestorationEmpty/Other
Kimi-K2.5Chinese Homo.26.6%38.0%32.9%2.5%
Number Homo.0.0%60.0%40.0%0.0%
Unconv. Chars.0.0%71.4%28.6%0.0%
DeepSeek-V3.2Chinese Homo.25.5%35.3%35.3%3.9%
Number Homo.0.0%100%0.0%0.0%
Unconv. Chars.0.0%62.5%37.5%0.0%
Qwen3-32BChinese Homo.13.6%22.7%54.5%9.1%
Number Homo.0.0%42.9%57.1%0.0%
Unconv. Chars.50.0%50.0%0.0%0.0%
GLM-4-32BChinese Homo.27.3%38.2%32.7%1.8%
Number Homo.0.0%66.7%33.3%0.0%
Unconv. Chars.33.3%33.3%33.3%0.0%
Gemma3-27BChinese Homo.22.2%44.4%26.7%6.7%
Number Homo.0.0%100%0.0%0.0%
Unconv. Chars.0.0%0.0%100%0.0%
Table 8: Representative Cell B cases (Tier 1 correct, Tier 2 wrong) across the three open-ended subcategories. In each case, the model correctly describes the neologism in Tier 1 but, in Tier 2, substitutes a semantic equivalent for the source form rather than producing the source form itself (Explanation/Paraphrase), which is the dominant failure mode across all three subcategories (Table 7). English references are provided below the generated definition.
SubcategoryModelNeologismGoldTier 1 Model Output (correct)Tier 2 Output (wrong)
Chinese Homo.Kimi-K2.5v我50微我50网络流行语,指通过微信转账50元人民币。源自肯德基“疯狂星期四”促销活动… Referring to transferring 50 yuan RMB via WeChat. It originated from KFC’s ‘Crazy Thursday’ promotional campaign.微信我50 (additional characters)
Chinese Homo.Kimi-K2.5石乐志失了智网络流行语,“失了智”的谐音,指失去理智、做出愚蠢或不理智的行为… A homophonic pun on “失了智” (lost one’s wits), meaning to lose one’s reason and act in a stupid or irrational way.失智 (dementia)
Chinese Homo.DeepSeek-V3.2河蟹和谐网络用语中常被用作“删除”或“屏蔽”的委婉说法,因“和谐”与“河蟹”谐音,后者常被用来讽刺内容审查。 Often used as a euphemism for ‘deleting’ or ‘blocking’ content. Because 和谐 (harmony) is a homophone of 河蟹 (river crab).删除 (delete)
Chinese Homo.DeepSeek-V3.2大清药丸大清要完大清药丸是网络流行语,源自“大清要完”的谐音,实际用来调侃某事物或局势即将崩溃或失败,带有讽刺或幽默意味。 A homophonic pun on ‘大清要完’ (the Qing dynasty is finished). It’s used to jokingly suggest that something or some situation is on the verge of collapse or failure, with a satirical or humorous undertone.大清朝要完 (additional characters)
Number Homo.Kimi-K2.53qthank you英文“thank you”的谐音数字缩写,“3”对应“th”音、“q”对应“you”音。 ‘3Q’ is a homophonic abbreviation of the English ‘thank you’, using numbers and letters: ‘3’ represents the ‘th’ sound and ‘q’ represents the ‘you’ sound.谢谢 (Chinese expression for thank you)
Number Homo.DeepSeek-V3.2666溜溜溜表示非常厉害、令人惊叹的意思,源自数字6的谐音“溜”。 ‘666’ means something is awesome or impressive. It derives from the number 6, a homophone of ‘溜’ (‘slick/skilled’)牛牛牛 (semantically equivalent)
Unconv. Chars.Kimi-K2.5占戈哥欠走已战歌起将“战歌起”三字拆解为偏旁部首(战=占+戈,歌=哥+欠,起=走+己/已)的隐晦写法… A cryptic way of writing ‘战歌起’ (cue the battle anthem) by splitting each of the three characters into their radical components.战歌响起 (additional characters)
Unconv. Chars.Kimi-K2.5口区将汉字“呕”拆分为“口”和“区”二字输入的网络黑话… The Chinese character ‘呕’ (to vomit/retch) is broken apart and typed as its two components, ‘口’ and ‘区’.呕吐 (additional characters)
Unconv. Chars.DeepSeek-V3.2彳亍口巴行吧网络用语,由“彳亍”和“口巴”拼接而成。其中“彳亍”是“行”的拆分形式,“口巴”是“吧”的拆分形式。 ‘彳亍’ is the disassembled form of ‘行’ and ‘口巴’ is the disassembled form of ‘吧’.好吧 (another expression for okay)
Unconv. Chars.DeepSeek-V3.2占戈土也战地“占戈土也”是“战地”的拆分写法… ‘占戈土也’ is the disassembled-character way of ‘战地’(‘battlefield’)战争 (war)

研究结果

  • 在释义任务上,表现最好的Kimi-K2.5整体准确率为67.74%,多数开源模型不到40%,而人类评测者平均达到85.63%。
  • 在通用基准上处于前沿水平的GPT-5.1在CNeo-Bench整体上仅得50.89%,低于中文优化模型,在拼音缩写(47.40%对Kimi-K2.5的76.56%)和拆字(21.74%对82.61%)上差距尤为明显。
  • 所有模型在第二层任务上的得分都高于自身在第一层的得分,但在三个需要开放式还原的小类(两种谐音类型加拆字)中,第一层答对的题目里有24.2%到57.1%在第二层还原失败。
  • 分析第一层答对但第二层答错的案例,发现最常见的失败模式是模型给出意思相近的改写或近义词,而不是给出原始形式本身。
  • 在GPT-5.1、DeepSeek-V3.2、Kimi-K2.5三个模型共同答错的1,058个难题上,给1个示例就能让37%到50%的题目被答对,给3个示例能提升到53%到67%,但即便给3个示例,仍有33%到47%的题目答不对。

可应用场景

  • 可作为诊断工具,检验聊天机器人、翻译系统或内容审核工具处理中文网络流行语和梗的能力水平。
  • 可用来定位模型知道新词含义却无法操作其底层形式的薄弱环节,为设计涉及新词的提示词或示例策略提供参考。
  • 可作为参考,判断哪种构词机制(谐音替代、拆字等)对模型最难,从而指导针对性训练数据或微调方案的设计。

局限与待验证事项

  • 第一层的打分依赖大模型评判员判断语义是否恰当,因此即便没有还原出原始形式,释义也可能被判为正确。
  • 数字谐音(29条)和非常规拆字(23条)两个小类样本量很少,针对这两类的模型间细致比较只能作为参考,不能视为确定结论。
  • 本研究专注于诊断这一差距的存在,并未提出解决识别-操作差距的训练方法,留待未来工作。
  • 数据收集截止到2026年1月,之后出现的新词未被纳入。
  • 即便给出3个示例,仍有33%到47%的难题未被解决,这部分残留错误的成因尚未在本文中进一步分析或解决。

为什么重要

新词处于模型训练数据的边缘甚至之外,因此模型处理新词的能力是检验其语言灵活性的一个有效探针。这一结果说明,仅凭释义任务的表现来判断模型是否真正理解某种语言现象是不够的。

本文术语

  • 新词(Neologism) · 近期新造出来的词,或原有词被赋予了新的含义
  • 谐音替代(Homophonic substitution) · 用发音相近但意思无关的数字或汉字替代原本的表达,例如886因发音接近拜拜咯而被使用
  • 非常规拆字(Unconventional Character Decomposition) · 把一个汉字拆成视觉上的组成部分分开书写,例如把行拆成彳和亍
  • 识别-操作差距(Recognition-Manipulation Gap) · 模型能正确说出新词的含义,却无法还原出该新词最初是从哪个原始形式变化而来的现象
  • 大模型评判员(LLM-as-judge) · 用另一个大模型来自动判断模型生成的答案是否与参考答案意思相符

论文原文摘要(英文)

Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench, a benchmark of 4,759 such neologisms with reference definitions, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression. CNeo-Bench is paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whether it can operate on its underlying mechanism. Evaluating 18 LLMs, we find that Chinese neologisms remain an open challenge; most models fall below 40\% on definition generation, and on several subcategories a systematic recognition-manipulation gap emerges: models describe neologisms correctly but, in source-form restoration tasks, substitute a semantic equivalent (paraphrase) for the source form rather than producing the source form itself. A few-shot analysis on 1,058 hard items shows that in-context examples can solve many difficult cases, but leave a noticeable portion of errors remaining, indicating challenges beyond prompting alone can address.

作者 · Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Kaiyan Zhao et al., arXiv:2608.28053, arxiv-nonexclusive