METAL LAB

886, yyds, 彳亍 같은 중국 신조어를 LLM이 정말 이해하는지 테스트했더니, 뜻은 설명해도 원래 형태는 못 만드는 모델이 많았다

arXiv:2608.280532026-08-31

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

886, yyds, 彳亍 같은 중국 신조어를 LLM이 정말 이해하는지 테스트했더니, 뜻은 설명해도 원래 형태는 못 만드는 모델이 많았다

연구팀은 886(바이바이), yyds(영원한 신), 彳亍(行을 쪼갠 글자) 같은 중국 인터넷 신조어 4,759개를 모아 언어적 생성 방식에 따라 9개 하위 유형으로 분류한 CNeo-Bench를 만들었다. GPT-5.1, DeepSeek, Kimi 등 18개 LLM을 두 단계로 평가했는데, 신조어의 뜻을 설명하는 과제(Tier 1)에서는 대부분 모델이 40% 미만 정확도에 머물렀고, 원래 형태를 복원하는 과제(Tier 2)에서는 뜻은 맞게 설명한 신조어조차 원형 대신 비슷한 뜻의 다른 표현을 내놓는 경우가 많았다. 어려운 문제 1,058개에 예시를 1~3개 보여주면 절반가량은 풀렸지만, 3개를 보여줘도 33~47%는 여전히 못 풀었다.

METAL LAB 해설 도표

신조어가 뜻풀이 단계로 들어가면 절반 넘는 모델이 의미는 맞게 설명한다. 하지만 그 다음 원형복원 단계로 가는 길목에는 관문이 있다: 유의어대체라는 관문으로, 원래 형태 대신 비슷한 뜻의 다른 표현으로 바꿔치기하는 현상이다. 그 결과 뜻은 맞혀놓고도 정작 원형은 복원하지 못하는 모델이 많다.
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 중국 인터넷에서 쓰이는 신조어(줄임말, 발음이 비슷한 글자로 바꾸기, 한자를 쪼개 쓰기, 외래어 차용 등) 4,759개를 정의문과 함께 모아 5개 대분류, 9개 소분류로 나눈 CNeo-Bench 벤치마크를 만들었다.
  2. 평가는 두 단계로 나뉜다. Tier 1은 신조어를 보고 뜻을 설명하는 과제이고, Tier 2는 신조어가 만들어진 원리(예: 원래 한자, 원래 문구)를 직접 복원하거나 객관식으로 맞추는 과제다.
  3. GPT-5.1, DeepSeek-V3/V3.2, Kimi-K2-Instruct/K2.5, Qwen3 계열, GLM-4 계열, InternLM3, Gemma3 계열, Llama3.2 계열 등 18개 모델을 온도 0(가장 확실한 답만 내놓는 설정)으로 평가했다.
  4. 가장 성적이 좋았던 Kimi-K2.5도 뜻 설명 과제에서 67.74%에 그쳤고, 사람 평가자의 정확도 85.63%보다 낮았다. 일반 벤치마크에서 앞서는 GPT-5.1은 오히려 50.89%로 중국어 특화 모델들보다 낮았다.
  5. 발음이나 한자 분해로 만들어진 신조어의 경우, 모델이 Tier 1에서는 뜻을 맞게 설명하면서도 Tier 2에서 원래 형태를 그대로 복원하지 않고 비슷한 뜻의 다른 표현으로 대신 답하는 패턴, 즉 '인식-조작 격차'가 여러 모델에서 공통으로 나타났다.
Figure 1: Categories and tasks in CNeo-Bench.
Figure 1: Categories and tasks in CNeo-Bench.
Table 1: Statistics for CNeo-Bench. “Full” is the total number of items in each subcategory; “Hard” is the number of items on which GPT-5.1, DeepSeek-V3.2, and Kimi-K2.5 jointly fail under definition generation.
SubcategoryFullHardHard Rate
Pinyin Abbr.1923518.23%
English Abbr.58712.07%
Chinese Homo.3776517.24%
Number Homo.29724.14%
Lexical Neo.116129825.67%
Semantic Neo.85018521.76%
Unconv. Char.23313.04%
Formulaic.166734720.82%
Foreign40211127.61%
Total4759105822.23%
Figure 2: Contingency analysis of Tier 1 (definition generation) vs Tier 2 (category-specific) outcomes for five representative models on three open-ended subcategories. Cells A/B/C/D denote (T1✓, T2✓), (T1✓, T2×), (T1×, T2✓), (T1×, T2×) respectively. The highlighted B/(A+B) column reports the operation failure rate (the fraction of correctly described neologisms on which restoration fails) directly quantifying the recognition-manipulation gap.
Figure 2: Contingency analysis of Tier 1 (definition generation) vs Tier 2 (category-specific) outcomes for five representative models on three open-ended subcategories. Cells A/B/C/D denote (T1✓, T2✓), (T1✓, T2×), (T1×, T2✓), (T1×, T2×) respectively. The highlighted B/(A+B) column reports the operation failure rate (the fraction of correctly described neologisms on which restoration fails) directly quantifying the recognition-manipulation gap.
Table 2: Full definition generation task results. Scores are reported as accuracy (%). Bold indicates the best score in each column; underline indicates the second best. Avg.∗ is computed as micro-average across all 4,759 instances.
ModelAlphabeticHomophonicLexicalStylisticsForeignAvg.∗
PinyinEnglishChineseNumberLexicalSemanticUnconv. Char.Formulaic
GPT-5.147.4074.1448.8144.8349.1054.9421.7452.6742.5450.89
DeepSeek-V364.0674.1455.4472.4151.5154.5965.2254.3545.2753.81
DeepSeek-V3.258.8579.3160.2172.4154.6155.6569.5754.1749.0055.26
Kimi-K2-Instruct66.1570.6961.2772.4158.1464.4765.2262.5753.4861.27
Kimi-K2.576.5679.3170.8272.4166.0667.0682.6167.5561.6967.74
Qwen3-4B9.3841.3815.5317.2424.8928.128.7026.0313.9323.49
Qwen3-4B-Instruct-25077.8153.4516.7120.6927.8232.9413.0428.9116.6726.69
Qwen3-8B17.7153.4518.0427.5931.3533.654.5532.0318.1629.40
Qwen3-32B21.8863.7925.9934.4836.6939.1817.3937.7925.3735.34
GLM-4-9B-041418.7548.2820.1620.6929.4634.008.7029.0321.3928.35
GLM-4-32B-041430.2163.7935.2865.5241.2645.8834.7841.0930.6040.60
InternLM3-8B-Instruct17.1958.6218.3034.4825.8431.064.5520.9415.1723.56
Gemma3-4B-it5.2137.937.693.4517.2319.290.0016.6810.4515.68
Gemma3-12B-it19.2748.2814.3217.2425.5827.064.5523.3418.1623.41
Gemma3-27B-it26.0463.7921.2213.7931.5230.004.5529.5719.4028.66
Llama3.2-1B-Instruct1.563.453.1810.345.774.470.006.242.245.00
Llama3.2-3B-Instruct5.2124.143.713.455.778.710.008.646.227.33
Llama3.2-11B-Instruct4.6934.487.163.4514.1317.290.0013.4410.7013.34
Figure 3: Few-shot recovery on 1,058 hard samples. Each bar shows the number of hard samples recovered (i.e., judged correct under the few-shot setting) under 1-shot, 2-shot, and 3-shot prompting. Bars are stacked by subcategory to show the distribution of recoveries across the nine subcategories.
Figure 3: Few-shot recovery on 1,058 hard samples. Each bar shows the number of hard samples recovered (i.e., judged correct under the few-shot setting) under 1-shot, 2-shot, and 3-shot prompting. Bars are stacked by subcategory to show the distribution of recoveries across the nine subcategories.
Table 3: Category-specific task results. Scores are reported as accuracy (%). Bold indicates the best score in each column; underline indicates the second best. Tasks with† use open-ended restoration; all other tasks use multiple-choice format. Avg.∗ is computed as micro-average across all 4,759 instances.
ModelAlphabeticHomophonicLexicalStylisticsForeignAvg.∗
PinyinEnglishChinese†Number†LexicalSemanticUnconv. Char.†Formulaic
GPT-5.172.4098.2856.7672.4173.4796.4622.7368.7786.3675.70
DeepSeek-V377.6098.2865.2582.7677.4396.5845.4570.1081.0677.76
DeepSeek-V3.276.0494.8364.4682.7678.1295.7554.5564.2081.3175.61
Kimi-K2-Instruct78.6596.5568.9793.1083.0396.5859.0968.1784.6079.20
Kimi-K2.580.2196.5568.1779.3182.0097.7659.0975.3988.6481.94
Qwen3-4B23.4486.219.280.0049.5386.7913.6444.6551.7750.40
Qwen3-4B-Instruct-250722.4084.4816.186.9051.3492.1013.6446.0354.0452.99
Qwen3-8B30.2186.2119.103.4552.6390.214.5549.2258.5954.97
Qwen3-32B56.2594.8333.6913.7969.4296.1113.6458.9066.9266.63
GLM-4-9B-041438.5489.6626.2610.3461.4191.164.5537.6749.7553.47
GLM-4-32B-041465.6293.1037.9351.7270.0393.6318.1848.9866.6763.79
InternLM3-8B-Instruct16.1579.3115.120.0053.3292.690.0059.3346.9758.60
Gemma3-4B-it30.7386.219.553.4539.7991.044.5529.6044.9543.22
Gemma3-12B-it39.5889.6621.2210.3458.6692.694.5544.6557.0755.78
Gemma3-27B-it41.6794.8329.4420.6966.3295.644.5548.1362.3760.71
Llama3.2-1B-Instruct24.4848.280.270.0025.4031.130.004.0324.2416.81
Llama3.2-3B-Instruct26.0470.690.270.0036.6154.720.005.6028.7925.03
Llama3.2-11B-Instruct24.4874.146.906.9039.2883.844.5536.0445.9643.57
Figure 4: Accuracy on English-based vs Pinyin-based alphabetic abbreviations in the definition generation task. English-based abbreviations are substantially easier than Pinyin-based ones for all evaluated models.
Figure 4: Accuracy on English-based vs Pinyin-based alphabetic abbreviations in the definition generation task. English-based abbreviations are substantially easier than Pinyin-based ones for all evaluated models.
Table 4: Human performance on a stratified sample of CNeo-Bench Tier 1 definition generation. We leave Kimi-K2.5, the strongest model we evaluate, for comparison.
SubcategoryHumanKimi-K2.5
Pinyin Abbr.91.00%76.56%
English Abbr.84.48%79.31%
Chinese Homo.94.00%70.82%
Number Homo.82.76%72.41%
Lexical Neo.87.00%66.06%
Semantic Neo.83.00%67.06%
Unconv. Chars.95.65%82.61%
Formulaic78.00%67.55%
Foreign80.00%61.69%
Overall85.63%67.74%
Figure 5: Tier 2 minus Tier 1 accuracy (percentage points) across 18 models and 9 subcategories. Positive gaps (red) indicate Tier 2 exceeds Tier 1; negative gaps (blue) indicate Tier 2 underperforms Tier 1. Subcategories marked with † use open-ended restoration; others use multiple-choice.
Figure 5: Tier 2 minus Tier 1 accuracy (percentage points) across 18 models and 9 subcategories. Positive gaps (red) indicate Tier 2 exceeds Tier 1; negative gaps (blue) indicate Tier 2 underperforms Tier 1. Subcategories marked with † use open-ended restoration; others use multiple-choice.
Table 5: Mechanism-driven taxonomy of Chinese neologisms in CNeo-Bench, with five top-level categories and nine subcategories. ∗Foreign Borrowings has no further subcategorization and is treated as a single class.
CategoriesSubcategoriesExplanationsExamples
1. Alphabetic AbbreviationsPinyin Abbr.Expressions formed by taking the initial letters of the Pinyin transliteration of Chinese phrases to create abbreviated forms.yyds from “永远的神” (yong yuan de shen), which means forever a legend.
English Abbr.Expressions formed by taking the initial letters of the English phrases to create abbreviated forms.ACG from “Anime, Comics, Games”.
2. Homophonic ExpressionsNumber Homo.Expressions that use numbers to convey meaning based on phonetic similarity.“拜拜咯” becomes “886” as they have similar pronunciations in Chinese, which means bye bye.
Chinese Homo.Expressions that replace standard Chinese characters with characters of similar pronunciation.“与你无关” (yu ni wu guan) becomes “雨女无瓜” (yu nv wu gua), which means none of your business.
3. Lexical InnovationsLexical Neo.Newly coined expressions that introduce novel lexical items.“键盘侠”, which means keyboard warrior, a person who aggressively criticizes or attacks others online without engaging in real-world action.
Semantic Neo.Existing words that acquire new meanings in emerging usage.“养鱼”, literal meaning: fish farming; new meaning: to maintain multiple romantic prospects simultaneously, often by keeping others as backups.
4. Stylistic and Pragmatic ExpressionsUnconv. Chars.Expressions that employ rare, archaic, or visually stylized Chinese characters in non-standard ways to convey meaning or achieve stylistic and expressive effects.“彳亍”, an expression composed of two archaic characters “彳” and “亍” to form “行”, which means okay.
FormulaicFixed or semi-fixed expressions derived from specific memes or events, which are reused across contexts to convey particular pragmatic meanings.“没活了可以咬打火机”, a meme-based expression originating from a viral livestream clip, used to humorously suggest that someone has run out of ideas or content.
5. Foreign BorrowingsForeign∗Expressions derived from foreign-language sources through phonetic reinterpretation or direct lexical adoptiont.“逮虾户” from “Déjà Vu” (phonetic, via Japanese anime Initial D), used for high-speed driving scenes; “文化祭”, which means school cultural festival is a direct loanword from Japanese.
Table 6: Representative examples for each Tier 2 tasks in CNeo-Bench. Subcategories marked with † require the model to complete open-ended restoration. English references are not included in our dataset.
SubcategoryTaskExampleAnswer
Pinyin Abbr.ClozeContext: 这家楼下新开的烧烤摊味道真的太绝了,朋友吃完直接说它是夜宵界的(____)。 English: The new BBQ stall that just opened downstairs is absolutely amazing, my friend straight-up called it the (____) of the late-night food scene. Options: (A) cxk (B) xswl (C) yysy (D) yyds(D) yyds (“forever legendary”)
English Abbr.ClozeContext: 他平时看番、追漫画还爱玩手游,朋友圈里一看就是个(____)重度爱好者。 English: He watches anime, follows manga, and loves playing mobile games, and you can tell he’s a hardcore (____) enthusiast from his social media. Options: (A) KPI (B) ACG (C) PDF (D) CPU(B) ACG (Anime, Comics, Games)
Chinese Homo.†RestorationNeologism: 雨女无瓜 Example context: 这件事是他们部门内部的安排,雨女无瓜,别再到处打听了。 English: This is an internal arrangement within their department, none of your business, so stop asking around about it.与你无关
Number Homo.†RestorationNeologism: 886 Example context: "时间不早了, 我先撤, 886, 明天见。 English: It’s getting late, I’m gonna head out. Bye-bye, see you tomorrow!拜拜咯
Lexical Neo.ClozeContext: 事情刚上热搜、真相还没出来,他就在评论区里对当事人冷嘲热讽、指点江山,活脱脱一个(____)。 English: The story had just started trending and the truth hadn’t even come out yet, but there he was in the comments section mocking the people involved and pontificating about everything, a textbook (____). Options: (A) 程序员 (B) 黑客 (C) 杠精 (D) 键盘侠(D) 键盘侠 (keyboard warrior)
Semantic Neo.Semantic DiscriminationTarget word: 养鱼 Sentences: (A) 他在院子里挖了个小池子,专门养鱼和种睡莲。(fish farming) (B) 爷爷退休后最大的乐趣就是养鱼,每天一早都去喂食。(fish farming) (C) 她总是同时和好几个人保持暧昧,朋友都说她是在养鱼。 (D) 这家农场除了养鸡养鸭,还养鱼,逢年过节卖得特别好。(fish farming)(C) maintaining multiple romantic relationships
Unconv. Chars.†RestorationNeologism: 彳亍行 (“okay”)
FormulaicScenario MatchingScenario: 家庭聚会上,62岁的老周本来在陪孙子搭积木,结果看到院里积雪够厚,突然兴奋地拉着几个晚辈去堆“巨型雪怪”,还非要拿胡萝卜做獠牙。被老伴笑着数落一把年纪还这么能闹,他却越玩越起劲。 English: At the family gathering, 62-year-old Zhou had been building blocks with his grandson, but when he saw how thick the snow had piled up in the yard, he suddenly got excited and dragged a few of the younger family members out to build a “giant snow monster”, and he insisted on using carrots for the fangs. His wife laughingly scolded him for being so rowdy at his age, but the more he played, the more into it he got. Options: (A) 老夫聊发少年狂 (B) 人间清醒 (C) 差生文具多 (D) 男人至死是少年(D) 男人至死是少年 (Men never really grow up)
Foreign†Source IdentificationNeologism: 文化祭 Options: (A) 它是从日语“文化祭(ぶんかさい)”转译扩展的词,“祭”在日语里有集中展示、主题活动之意,“文化祭”本质上是对日本校园文化展示周这一活动与说法的概括吸收。(Expanded translation) (B) 它是从日语“文化祭(ぶんかさい)”直接借入的词,“祭”在日语里有节庆、庆典之意,“文化祭”本质上是对日本校园“学园文化祭”这一制度与说法的直接沿用。(Direct loanword) (C) 它是从日语“文化祭(ぶんかさい)”类推形成的词,“祭”在日语里有公开活动、校内集会之意,“文化祭”本质上是对日本校园“社团发表会”这一形式与说法的重新命名。(Analogical formation) (D) 它是从日语“文化祭(ぶんかさい)”借形改义的词,“祭”在日语里有热闹活动、集体庆祝之意,“文化祭”本质上是对日本校园“校园开放日”这一安排与说法的本地化套用。(Borrowed form, new meaning)(B) Direct loanword
Table 7: Breakdown of Cell B failure modes (Tier 1 correct, Tier 2 wrong) across five models and three open-ended subcategories. Each row sums to 100%. Copy Verbatim: model outputs the neologism itself. Explanation/Paraphrase: model outputs a semantic gloss or near-synonym of the source form. Wrong Restoration: model outputs a different Chinese expression that is neither the neologism nor a paraphrase. Empty/Other: empty output or unparseable.
ModelSubcategoryCell B Failure Mode (%)
Copy VerbatimExplanation/ParaphraseWrong RestorationEmpty/Other
Kimi-K2.5Chinese Homo.26.6%38.0%32.9%2.5%
Number Homo.0.0%60.0%40.0%0.0%
Unconv. Chars.0.0%71.4%28.6%0.0%
DeepSeek-V3.2Chinese Homo.25.5%35.3%35.3%3.9%
Number Homo.0.0%100%0.0%0.0%
Unconv. Chars.0.0%62.5%37.5%0.0%
Qwen3-32BChinese Homo.13.6%22.7%54.5%9.1%
Number Homo.0.0%42.9%57.1%0.0%
Unconv. Chars.50.0%50.0%0.0%0.0%
GLM-4-32BChinese Homo.27.3%38.2%32.7%1.8%
Number Homo.0.0%66.7%33.3%0.0%
Unconv. Chars.33.3%33.3%33.3%0.0%
Gemma3-27BChinese Homo.22.2%44.4%26.7%6.7%
Number Homo.0.0%100%0.0%0.0%
Unconv. Chars.0.0%0.0%100%0.0%
Table 8: Representative Cell B cases (Tier 1 correct, Tier 2 wrong) across the three open-ended subcategories. In each case, the model correctly describes the neologism in Tier 1 but, in Tier 2, substitutes a semantic equivalent for the source form rather than producing the source form itself (Explanation/Paraphrase), which is the dominant failure mode across all three subcategories (Table 7). English references are provided below the generated definition.
SubcategoryModelNeologismGoldTier 1 Model Output (correct)Tier 2 Output (wrong)
Chinese Homo.Kimi-K2.5v我50微我50网络流行语,指通过微信转账50元人民币。源自肯德基“疯狂星期四”促销活动… Referring to transferring 50 yuan RMB via WeChat. It originated from KFC’s ‘Crazy Thursday’ promotional campaign.微信我50 (additional characters)
Chinese Homo.Kimi-K2.5石乐志失了智网络流行语,“失了智”的谐音,指失去理智、做出愚蠢或不理智的行为… A homophonic pun on “失了智” (lost one’s wits), meaning to lose one’s reason and act in a stupid or irrational way.失智 (dementia)
Chinese Homo.DeepSeek-V3.2河蟹和谐网络用语中常被用作“删除”或“屏蔽”的委婉说法,因“和谐”与“河蟹”谐音,后者常被用来讽刺内容审查。 Often used as a euphemism for ‘deleting’ or ‘blocking’ content. Because 和谐 (harmony) is a homophone of 河蟹 (river crab).删除 (delete)
Chinese Homo.DeepSeek-V3.2大清药丸大清要完大清药丸是网络流行语,源自“大清要完”的谐音,实际用来调侃某事物或局势即将崩溃或失败,带有讽刺或幽默意味。 A homophonic pun on ‘大清要完’ (the Qing dynasty is finished). It’s used to jokingly suggest that something or some situation is on the verge of collapse or failure, with a satirical or humorous undertone.大清朝要完 (additional characters)
Number Homo.Kimi-K2.53qthank you英文“thank you”的谐音数字缩写,“3”对应“th”音、“q”对应“you”音。 ‘3Q’ is a homophonic abbreviation of the English ‘thank you’, using numbers and letters: ‘3’ represents the ‘th’ sound and ‘q’ represents the ‘you’ sound.谢谢 (Chinese expression for thank you)
Number Homo.DeepSeek-V3.2666溜溜溜表示非常厉害、令人惊叹的意思,源自数字6的谐音“溜”。 ‘666’ means something is awesome or impressive. It derives from the number 6, a homophone of ‘溜’ (‘slick/skilled’)牛牛牛 (semantically equivalent)
Unconv. Chars.Kimi-K2.5占戈哥欠走已战歌起将“战歌起”三字拆解为偏旁部首(战=占+戈,歌=哥+欠,起=走+己/已)的隐晦写法… A cryptic way of writing ‘战歌起’ (cue the battle anthem) by splitting each of the three characters into their radical components.战歌响起 (additional characters)
Unconv. Chars.Kimi-K2.5口区将汉字“呕”拆分为“口”和“区”二字输入的网络黑话… The Chinese character ‘呕’ (to vomit/retch) is broken apart and typed as its two components, ‘口’ and ‘区’.呕吐 (additional characters)
Unconv. Chars.DeepSeek-V3.2彳亍口巴行吧网络用语,由“彳亍”和“口巴”拼接而成。其中“彳亍”是“行”的拆分形式,“口巴”是“吧”的拆分形式。 ‘彳亍’ is the disassembled form of ‘行’ and ‘口巴’ is the disassembled form of ‘吧’.好吧 (another expression for okay)
Unconv. Chars.DeepSeek-V3.2占戈土也战地“占戈土也”是“战地”的拆分写法… ‘占戈土也’ is the disassembled-character way of ‘战地’(‘battlefield’)战争 (war)

실제로 확인된 결과

  • Tier 1(뜻 설명) 과제에서 최고 성적 모델 Kimi-K2.5가 67.74%를 기록했고, 오픈소스 모델 다수는 40% 미만이었다. 사람 평가자 평균은 85.63%였다.
  • 일반 벤치마크에서 강력한 GPT-5.1은 CNeo-Bench 전체에서 50.89%에 그쳐 중국어 특화 모델보다 낮았고, 특히 병음 줄임말(47.40% vs Kimi 76.56%)과 한자 분해(21.74% vs Kimi 82.61%)에서 차이가 크게 벌어졌다.
  • Tier 2(원형 복원·객관식) 과제는 모든 모델에서 자기 Tier 1 점수보다 높았지만, 원형을 직접 복원해야 하는 세 유형(동음 대체 2종, 한자 분해)에서는 Tier 1에서 맞게 설명한 문항 중 24.2%~57.1%가 Tier 2에서 실패했다.
  • Tier 1은 맞고 Tier 2는 틀린 오류를 분석한 결과, 가장 흔한 실패 유형은 원래 형태 대신 비슷한 뜻의 다른 표현(의역·유의어)을 내놓는 것이었다.
  • GPT-5.1, DeepSeek-V3.2, Kimi-K2.5가 모두 실패한 어려운 문항 1,058개에 예시를 1개만 보여줘도 37~50%가 정답으로 회복됐고, 3개를 보여주면 53~67%까지 회복됐지만 나머지는 3개를 줘도 여전히 틀렸다.

어디에 쓸 수 있나

  • 중국어 인터넷 밈이나 신조어를 다루는 챗봇, 번역기, 콘텐츠 모더레이션 시스템의 언어 이해 수준을 점검하는 진단 도구로 활용할 수 있다.
  • 모델이 뜻은 알지만 원형을 복원하지 못하는 약점을 찾아, 신조어 관련 프롬프트 설계나 예시 제공 전략을 세우는 데 참고할 수 있다.
  • 특정 언어 현상(발음 대체, 문자 분해 등)에 특화된 학습 데이터나 파인튜닝 전략을 설계할 때 어떤 유형이 취약한지 파악하는 기준으로 쓸 수 있다.

한계와 남은 검증

  • Tier 1 채점은 LLM 판정관이 의미가 통하는지만 보기 때문에, 원래 형태를 재현하지 않아도 정답으로 처리될 수 있다.
  • 숫자 동음(29개), 한자 분해(23개) 두 소분류는 항목 수가 적어 모델별 세부 비교는 참고용으로만 봐야 한다.
  • 이 연구는 문제를 진단하는 데 집중했고, 인식-조작 격차를 실제로 줄이는 학습 방법은 다루지 않아 향후 연구로 남겨졌다.
  • 데이터는 2026년 1월까지 수집된 것으로, 새로 생겨나는 신조어는 반영되어 있지 않을 수 있다.
  • few-shot 예시로 회복되지 않는 33~47%의 잔여 오류에 대한 원인 분석과 해결책은 아직 제시되지 않았다.

왜 중요한가

신조어는 모델이 학습 데이터에서 거의 접하지 못했을 가능성이 큰 표현이라, 이를 다루는 능력은 모델이 언어를 얼마나 유연하게 다루는지 보여주는 지표가 된다. 이 결과는 겉보기 뜻풀이 성능만으로 모델의 언어 이해력을 판단하면 안 된다는 것을 보여준다.

이 논문의 용어

  • 신조어(Neologism) · 최근에 새로 만들어지거나 기존 표현이 새로운 뜻으로 쓰이게 된 말
  • 동음 대체(Homophonic substitution) · 뜻이 다르지만 발음이 비슷한 글자나 숫자로 원래 표현을 바꿔 쓰는 방식. 예: 886은 '拜拜咯(바이바이)'와 발음이 비슷해서 쓰임
  • 한자 분해(Unconventional Character Decomposition) · 한 글자를 눈에 보이는 부분들로 쪼개 따로 쓰는 방식. 예: 行(행)을 彳와 亍로 나눠 씀
  • 인식-조작 격차(Recognition-Manipulation Gap) · 모델이 신조어의 뜻은 맞게 설명하지만, 그 신조어가 만들어진 원래 형태(원문·원 한자 등)는 복원하지 못하는 현상
  • LLM 판정관(LLM-as-judge) · 모델이 생성한 답을 다른 LLM이 정답과 비교해 맞았는지 채점하는 평가 방식

저자 · Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao, Yoshimasa Tsuruoka

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Kaiyan Zhao et al., arXiv:2608.28053, arxiv-nonexclusive