월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

70B짜리 대형 언어모델 대신 6억~10억 파라미터 소형 모델로도 인간 선택 데이터 예측력을 거의 그대로 재현할 수 있었다

arXiv:2608.052242026-08-04

Small Foundation Models of Human Cognition and Behaviour

70B짜리 대형 언어모델 대신 6억~10억 파라미터 소형 모델로도 인간 선택 데이터 예측력을 거의 그대로 재현할 수 있었다

연구팀은 1억3500만~140억 파라미터 규모의 모델 14개를 심리학 실험 데이터인 Psych-101에 파인튜닝해서, 모델 크기가 예측 정확도에 얼마나 영향을 주는지, 그리고 모델이 실제로 과제 구조를 이해하는지 아니면 단순 통계적 요행을 부리는지를 검증했다. 학습에 쓴 데이터와 같은 종류의 실험(분포 내)에서는 크기가 거의 의미가 없어서 6억~10억 파라미터 모델이 700억 파라미터 모델과 맞먹었지만, 학습에 없던 새로운 실험(분포 외)에서는 모델이 클수록 확실히 더 잘했다. 프롬프트를 조각내서 하나씩 지워보는 실험에서는 모델이 단순히 과거 선택 기록만 외운 게 아니라 실험에서 실제로 제시된 자극과 결과 정보를 활용하고 있다는 것이 확인됐다.

METAL LAB 해설 도표

소형 인지 파인튜닝 모델이 검증받은 3단계 구조

증거 상태측정 결과가 보고됨

  1. 1. 규모 스윕 실험1억3500만~140억 파라미터 14개 모델을 Psych-101(160개 실험, 1070만 시행)에 파인튜닝해 분포 내·분포 외 성능을 비교
  2. 2. 프롬프트 채널 절제과제 설명, 자극, 피드백, 선택 기록 네 부분을 순서대로 지워가며 27개 실험에서 어떤 정보가 예측에 실제로 쓰이는지 확인
  3. 3. 시행 순서 교환 테스트독립적 시행 과제(THINGS)와 순서 의존적 과제(시간할인 선택)에서 시행 순서를 뒤섞어 모델이 과제 구조를 구분해 반응하는지 확인
  4. 결론: 과제 적응적 정보 활용모델은 단순 기억이나 형식 패턴이 아니라 실험이 실제로 제공한 자극·피드백 내용을 과제 구조에 맞게 사용한다
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. Psych-101이라는 16만 건 실험, 1070만 건의 시행 단위 선택 데이터를 모아둔 데이터셋에 Llama, Qwen3, SmolLM, OLMo 등 4개 모델군에서 총 14개 모델(1억3500만~140억 파라미터)을 파인튜닝했다.
  2. 학습에 쓰인 실험과 같은 종류(분포 내)에서는 모델 크기가 거의 의미가 없어서 6억~10억 파라미터 모델이 재현한 700억 파라미터 Centaur 모델과 맞먹는 성능을 냈지만, 처음 보는 새로운 실험(분포 외)에서는 큰 모델이 뚜렷하게 더 잘 일반화했다.
  3. 모델이 답을 맞히는 데 어떤 정보를 쓰는지 확인하려고, 프롬프트를 과제 설명, 실험에서 보여준 자극, 결과 피드백, 과거 선택 기록의 네 부분으로 나눠서 하나씩 지워가며 27개 실험에서 테스트했다.
  4. 자극과 피드백의 구체적인 내용(예: 실제 점수나 도형 모양)을 지우고 형식만 남기자 학습된 정보의 75.7%가 사라지고 성능이 무작위 추측보다 나빠졌는데, 이는 모델이 단순히 과거 선택 패턴만 암기해서 맞히는 게 아님을 보여준다.
  5. 시행 순서를 뒤섞는 테스트에서는, 시행들이 서로 독립적인 과제에서는 순서를 바꿔도 예측이 그대로였고, 이전 응답이 다음 시행 내용을 결정하는 과제에서는 순서를 바꾸면 예측이 달라져서, 모델이 각 과제의 구조에 맞게 정보를 적절히 다르게 쓴다는 것을 보여줬다.
Figure 1: Adapter rank against model size on Psych-101. Mean negative log-likelihood over the 38 of 46 Psych-101 tasks for which 11 publish a domain-specific cognitive model, plotted against parameter count, one panel per adapter rank (r=4–64). Trend lines are fitted only within groups matched on generation and context window, so Olmotaur has none and the Smoltaur fit covers the SmolLM2 models only.
Figure 1: Adapter rank against model size on Psych-101. Mean negative log-likelihood over the 38 of 46 Psych-101 tasks for which 11 publish a domain-specific cognitive model, plotted against parameter count, one panel per adapter rank (r=4–64). Trend lines are fitted only within groups matched on generation and context window, so Olmotaur has none and the Smoltaur fit covers the SmolLM2 models only.
Table 1: Core comparison of LLM-based cognitive and behavioural foundation models.
ModelBase LLMPost-trainingTraining dataDomain
Centaur (11)Llama-3.1-70BSFT (masked CE on response tokens); rank-stabilised QLoRA (r=α=8), 4-bit quantisedPsych-101: 160 expts, 60,092 partic., 10.7M choicesDecision, memory, learning, planning
1021Qwen2.5-7B-InstructSFT / Centaur-style SFT / GRPO compared; LoRA (r=α=32)choices13k: 13,102 train / 1,462 test risky choice problemsDecision
Be.FM (95)Llama-3.1-{8,70}B-InstructSFT with LoRA (all layers); 8-bit quantised2AER: 2,703 papers.; MobLab: 68,779 subj., 82,057 obs.; Big Five: 17,667 subj.Behavioural science, economic game, personality
Socrates (47)Llama-3-8B-Instruct Qwen2.5-14B-InstructSFT / SFT + oracle reasoning traces / contrastive DPO compared; full fine-tuningSocSci210: 210 TESS3 expts, 400,491 partic., 2.9M individual responsesSocial sciences (economics, psychology, political science)
HumanLLM (55)Qwen2.5-{3,7}B-Instruct Qwen3-8B4 Llama-3.1-8B-Instruct Phi-3-mini-128k-instructSFT (masked non-response tokens); full fine-tuning; 1:1 weight merge with base (LM-Cocktail)Cognitive Genome: Reddit 2.8M, Twitter 673K, Blogger 368K, Amazon 1.7M; 1.2M train samplesSocial intelligence (personalised behaviour)
GeCCo (66)Llama-3.1-70B-Instruct DeepSeek-R1-Distill-Llama-3.1-70B Qwen2.5-72B-InstructNo fine-tuning; in-context learning with iterative BIC-based refinement (10×5 runs)Behavioural data from 4 cognitive domains (in-context, not for training)Decision, learning, planning, working memory
Ours Llama-Centaur Qwentaur Smoltaur OlmotaurLlama-3.2-{1,3}B Llama-3.1-8B Qwen3-{0.6,1.7,4,8,14}B-Base SmolLM2-{135,360}M SmolLM2-1.7B SmolLM3-3B-Base OLMo-2-0425-1B OLMo-3-1025-7BSFT (masked CE on response tokens, Centaur-style); rank-stabilised LoRA (r=α∈{4,8,16,32,64})Psych-101: 160 expts, 60,092 partic., 10.7M choicesDecision, memory, learning, planning
1No named model; methods (SFT v. RL) comparison only. 28-bit quantisation applies to 70B variant only; 8B is unquantised. 3TESS: NSF’s Time-sharing Experiments for the Social Sciences, a repository of peer-reviewed social science experiments conducted on nationally representative samples. 455 states “Qwen3-8B” without Instruct suffix; base/instruct status unspecified.
Figure 2: In-distribution versus out-of-distribution scaling at rank 16. (a) Psych-101, mean NLL over the 38 tasks with a reported cognitive model. (b) Psych-201-RT, all 18 held-out experiments; no cognitive baseline is published for Psych-201. Trend lines are fitted only within groups matched on generation and context window.
Figure 2: In-distribution versus out-of-distribution scaling at rank 16. (a) Psych-101, mean NLL over the 38 tasks with a reported cognitive model. (b) Psych-201-RT, all 18 held-out experiments; no cognitive baseline is published for Psych-201. Trend lines are fitted only within groups matched on generation and context window.
Table 2: Training hyperparameters for LLM-based cognitive and behavioural models. Adaptation strategy refers to the parameter-efficient or full fine-tuning method applied to the base LLM. All our variants share identical hyperparameters except for per-device batch size, gradient accumulation steps, and the set of adapter ranks trained. Entries marked “Not specified” indicate that the original paper did not report the corresponding detail.
ModelAdaptationOptimiserEpochsLearning rateEff. batch size (PD×GA×d)Scheduler
Centaur (11)QLoRA (r=α=8), all linear layers8-bit AdamW115×10−51×32×1Cosine (WU 100 steps)
102LoRA r=α=32, all linear layers, dropout 0.05AdamWSFT: 6; RL: 3SFT: 10−5; RL: 3×10−6SFT: PD×8×1; RL: PD×8×4SFT: fixed; RL: cosine
Be.FM (95)LoRA, all layersNot specified310−41×8×dCosine (WU 0.1)
Socrates (47)Full fine-tuning (no LoRA)Not specified1SFT: 10−5; DPO: 10−6PD×GA×8=256Cosine (WU 0.05)
HumanLLM (55)Full fine-tuning (no LoRA)Not specified35×10−6PD×GA×8=64Cosine (WU 0.5)
GeCCo (66)N/A (no training)2N/AN/AN/AN/AN/A
Qwentaur-0.6BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−52×16×1Linear warmup (100 steps)
Qwentaur-1.7BLoRA (r=α∈{8,16}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Qwentaur-4BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Qwentaur-8BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Qwentaur-14BLoRA (r=α∈{4,16,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Llama-Centaur-1BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−52×16×1Linear warmup (100 steps)
Llama-Centaur-3BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Llama-Centaur-8BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Smoltaur-0.1BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−58×4×1Linear warmup (100 steps)
Smoltaur-0.4BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−54×8×1Linear warmup (100 steps)
Smoltaur-1.7BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−52×16×1Linear warmup (100 steps)
Smoltaur-3BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Olmotaur-1BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−52×16×1Linear warmup (100 steps)
Olmotaur-7BLoRA (r=α∈{4,8,16,32,64}), all linear layersAdamW15×10−51×32×1Linear warmup (100 steps)
Abbreviations: PD = per-device batch size; GA = gradient accumulation steps; d = number of GPU devices; WU = warmup ratio. 1The published training script (https://github.com/marcelbinz/Llama-3.1-Centaur-70B/blob/main/scripts/cluster_train.sh) specifies 5 epochs. 11 report training for 1 epoch, suggesting early checkpoint selection. 2GeCCo (66) uses in-context learning with iterative BIC-based feedback over 10 sampling iterations × 5 independent runs. Models are fitted to held-out data with SciPy minimize (20 random restarts) and evaluated by BIC. All post-training methods use supervised fine-tuning unless otherwise noted; 102 additionally compare RL (GRPO with 12 candidate completions per step, max 1024 tokens; reward =1−|oB−pB| + format bonus up to 0.5, no standard-deviation normalisation); Socrates (47) additionally compares contrastive DPO (preference pairs constructed by varying the demographic persona under the same experimental condition and outcome question).
Figure 4: Structural ablation results across eight models and 27 experiments with well-defined response options. Five experiments with mixed or continuous formats are excluded. (a) Mean information retention R=(ln⁡k−ℒ​c)/(ln⁡k−ℒ​orig) by condition, averaged over experiments; R=1: no loss, R=0: chance. Circles: Llama-Centaur; squares: Qwentaur; colour intensity scales with model size; error bars: SEM. (b) Per-experiment fraction of learned information lost, δ=1−R, averaged across models. Columns sorted by history-only δ; colour strip indicates task type. Experiments with δ>1 under the history-only condition are omitted; see Figure 18 (Appendix F.2) for the complete set.
Figure 4: Structural ablation results across eight models and 27 experiments with well-defined response options. Five experiments with mixed or continuous formats are excluded. (a) Mean information retention R=(ln⁡k−ℒ​c)/(ln⁡k−ℒ​orig) by condition, averaged over experiments; R=1: no loss, R=0: chance. Circles: Llama-Centaur; squares: Qwentaur; colour intensity scales with model size; error bars: SEM. (b) Per-experiment fraction of learned information lost, δ=1−R, averaged across models. Columns sorted by history-only δ; colour strip indicates task type. Experiments with δ>1 under the history-only condition are omitted; see Figure 18 (Appendix F.2) for the complete set.
Table 3: Data pipeline and loss configuration for LLM-based cognitive and behavioural models. Precision refers to the numerical format used during training: quantised formats (4-bit NF4, 8-bit) apply to base model weights while LoRA adapters and forward/backward computation use bf16. Loss masking describes which tokens contribute to the training loss; masked approaches restrict gradient updates to human response tokens, avoiding optimisation on task instructions and context. Data format describes the input representation seen by the model during training. Data synthesis indicates whether and how any computer-assisted tools were used to construct or augment the training corpus.
ModelPrecisionWeight decayLoss maskingData formatData synthesis
Centaur (11)4-bit NF40.01Human response tokens onlyNL trial-by-trial prompts (∼32K tokens)Template-based prompt construction1
102Not specifiedNot specifiedSFT: standard; Centaur-style. GRPOJSON aggregated choice proportions per problem (empirical % rounded to nearest integer, e.g. {"A": 29, "B": 71})Reformatted from choices13k empirical choice frequencies (problem-level rather than individual-participant-level prediction)
Be.FM (95)8-bit (70B base, bitsandbytes); bf16 (8B)Not specifiedStandard (Alpaca template)Alpaca template {instruction, input, output}GPT-4o for research workflow extraction
Socrates (47)Not specified0.1SFT: Response token only. DPO{persona, stimuli, outcome, response}o4-mini-high (dataset agent2); GPT-4o-mini (reasoning traces3)
HumanLLM (55)Not specifiedNot specifiedNon-response positions maskedShareGPT formatLlama-3.3-70B (extraction); GPT-4o (quality validation)
GeCCo (66)N/AN/AN/ANL prompt + Python function templateN/A
Oursbf160.01Human response tokens only (Centaur-style)NL trial-by-trial prompts (∼32K tokens)Psych-101 dataset (unmodified)
1Each experiment is converted into natural-language trial-by-trial prompts via author-written scripts that map structured experimental data (participant responses, stimuli, feedback) to verbalised narratives; for an example, see https://github.com/marcelbinz/Psych-201/blob/main/binz2022heuristics/generate_prompts.py. 2An LLM-based agent (o4-mini-high) parses raw TESS datasets into structured {persona, stimuli, response} tuples; see Appendix A of 47 for details. 3Given an experimental prompt and the corresponding human response, GPT-4o-mini generates reasoning traces explaining the human decision from a social scientist’s perspective; see Appendix E of 47.
Figure 5: Per-participant order variance across models and experiments. (a) THINGS odd-one-out (exchangeable; primary test). (b) Intertemporal choice (adaptive staircase; negative control). Left: violin and box plots; fine-tuned (solid) vs. base (hatched); y-axis scales differ. Right: ECDFs; solid = fine-tuned, dashed = base; steeper curves near zero indicate greater order invariance.
Figure 5: Per-participant order variance across models and experiments. (a) THINGS odd-one-out (exchangeable; primary test). (b) Intertemporal choice (adaptive staircase; negative control). Left: violin and box plots; fine-tuned (solid) vs. base (hatched); y-axis scales differ. Right: ECDFs; solid = fine-tuned, dashed = base; steeper curves near zero indicate greater order invariance.
Table 4: Infrastructure, packages, and computational cost for LLM-based cognitive and behavioural models. Hardware refers to GPU resources used during training (or in-context generation for GeCCo). Training time reports wall-clock duration as stated in each paper; missing entries indicate the information was not reported. Our models were each trained for 1 epoch on Psych-101 on a single A100 80GB, with rank-stabilised LoRA at r=α∈{4,8,16,32,64}.
ModelTraining FrameworkKey packagesHardwareTrain timeMax seq. len.Inference
Centaur (11)unslothunsloth1× A100 80GB∼5 days∼32,768Not specified
102Not specifiedvLLM (inference)RL: 4× H100; SFT: 1× A100RL: ∼80 h; SFT: ∼5 h1,024 (RL); 30 (SFT inf.)T=0.7, top-p=0.95, top-k=0.5
Be.FM (95)LlamaFactoryLlamaFactory, bitsandbytesNot specifiedNot specifiedNot specifiedNot specified
Socrates (47)LlamaFactoryLlamaFactory8× A100 80GB4–24 h4,096 (inf.)T=0.6, top-p=0.9
HumanLLM (55)LlamaFactoryLlamaFactory, DeepSpeed Zero, vLLM8× A100 40GB∼120 h (8B)8,192T=0.7
GeCCo (66)N/A1SciPy (minimize); Python exec()4× A100 40GB≤8 h per domainN/AT: 0.1–0.22
Qwentaur-0.6Bunslothunsloth1× A100 80GB4 h∼32,768
Qwentaur-1.7Bunslothunsloth1× A100 80GB7 h∼32,768
Qwentaur-4Bunslothunsloth1× A100 80GB18 h∼32,768
Qwentaur-8Bunslothunsloth1× A100 80GB1 d∼32,768
Qwentaur-14Bunslothunsloth1× A100 80GB1 d 15 h∼32,768
Llama-Centaur-1Bunslothunsloth1× A100 80GB7 h∼32,768
Llama-Centaur-3Bunslothunsloth1× A100 80GB12 h∼32,768
Llama-Centaur-8Bunslothunsloth1× A100 80GB21 h∼32,768
Smoltaur-0.1Bunslothunsloth1× A100 80GB3 h∼8,192
Smoltaur-0.4Bunslothunsloth1× A100 80GB4 h∼8,192
Smoltaur-1.7Bunslothunsloth1× A100 80GB7 h∼8,192
Smoltaur-3Bunslothunsloth1× A100 80GB1 d 4 h∼32,768
Olmotaur-1Bunslothunsloth1× A100 80GB5 h∼4,096
Olmotaur-7Bunslothunsloth1× A100 80GB2 d 1 h∼32,768
Abbreviations: T = sampling temperature; inf. = inference; d = days; h = hours. 1No training framework required; GeCCo generates cognitive models via in-context prompting. SciPy is used for parameter fitting of the generated models; Python exec() executes the LLM-generated code. 2Temperature varies by LLM: Llama 0.2, Qwen 0.15, DeepSeek-R1 0.1. All our models were trained on a single NVIDIA A100 80GB GPU. Training times scale approximately linearly with parameter count within each model family.
Figure 6: Adapter rank against model size on Psych-101. Mean negative log-likelihood over the 38 of 46 Psych-101 tasks for which 11 publish a domain-specific cognitive model baseline, against parameter count, with colour intensity encoding LoRA rank and right-hand panels showing one rank at a time. Trend lines and shaded envelopes are fitted only within groups matched on model generation and context window, so Olmotaur has none and the Smoltaur fit covers the SmolLM2 models only. Marker shape indicates context window (circle ≥ 32k, square < 32k); colour indicates family.
Figure 6: Adapter rank against model size on Psych-101. Mean negative log-likelihood over the 38 of 46 Psych-101 tasks for which 11 publish a domain-specific cognitive model baseline, against parameter count, with colour intensity encoding LoRA rank and right-hand panels showing one rank at a time. Trend lines and shaded envelopes are fitted only within groups matched on model generation and context window, so Olmotaur has none and the Smoltaur fit covers the SmolLM2 models only. Marker shape indicates context window (circle ≥ 32k, square < 32k); colour indicates family.
Table 5: Fraction of available information captured above chance, (ln⁡k−NLL)/ln⁡k, by task type for finetuned models (bf16). A value of 0 indicates chance-level performance; 1 indicates perfect prediction. Restricted to 34 experiments (of 46) with both a cognitive model baseline and a well-defined discrete response space (ln⁡k>0); 12 experiments are excluded. Finetuned families: Qwentaur, Llama-Centaur, Smoltaur, Olmotaur. Base columns: Q-8B/Q-14B (Qwen3), L-8B (Llama-3.1). Subscript r denotes our reproducing evaluation of the original Centaur model distributed by 11, evaluated under identical python library and CUDA versions as our small foundation models for fair comparison. Subscript p denotes values published by 11. 70Bp is shown for reference but excluded from best/second-best marking, as it was evaluated under different software conditions. Bold+underline marks the best model, underline the second-best.
QwentaurLlama-CentaurSmoltaurOlmotaurBase11
Task type0.6B1.7B4B8B14B1B3B8B0.1B0.4B1.7B3B1B7BQ-8BQ-14BL-8B70Bp70Br
Decision (8)0.530.530.540.540.540.530.530.540.410.460.480.540.470.530.340.350.330.560.52
MDP (5)0.280.300.310.310.320.280.300.310.180.200.260.300.200.300.150.170.190.320.31
Bandit (12)0.480.480.490.490.500.460.480.490.360.410.440.490.420.480.350.360.330.460.45
Memory (4)0.570.570.580.580.580.560.570.580.460.510.550.570.530.570.420.470.430.580.58
Misc. (2)0.570.570.580.580.580.560.570.580.450.530.550.580.550.570.410.420.400.570.57
Sup. learn. (3)0.300.290.310.310.310.290.310.310.180.230.250.300.200.310.240.240.230.300.30
Mean (34)0.460.460.470.480.480.450.470.480.350.390.430.470.400.470.320.340.320.470.46
Figure 7: Adapter rank and training-set size by family. (Left) Mean NLL against LoRA rank at full data. (Right) Mean NLL against the fraction of Psych-101 used for training, at r=16. Rows are the four model families; the dotted line is the cognitive-model baseline and the diamonds are the reproduced and reported Centaur-70B values. Subsets are nested and experiment-stratified, so every fraction covers all 160 experiments and reducing data quantity does not reduce paradigm coverage.
Figure 7: Adapter rank and training-set size by family. (Left) Mean NLL against LoRA rank at full data. (Right) Mean NLL against the fraction of Psych-101 used for training, at r=16. Rows are the four model families; the dotted line is the cognitive-model baseline and the diamonds are the reproduced and reported Centaur-70B values. Subsets are nested and experiment-stratified, so every fraction covers all 160 experiments and reducing data quantity does not reduce paradigm coverage.
Table 9: Fraction of available information captured above chance, (ln⁡k−NLL)/ln⁡k, by task type on Psych-201 (out-of-distribution). A value of 0 indicates chance-level performance; 1 indicates perfect prediction. Restricted to 15 experiments (of 18) with a well-defined discrete response space (ln⁡k>0); 3 experiments with continuous or mixed responses are excluded. Bold+underline marks the best model, underline the second-best.
QwentaurLlama-CentaurSmoltaurOlmotaurBase11
Task type0.6B1.7B4B8B14B1B3B8B0.1B0.4B1.7B3B1B7BQ-8BQ-14BL-8BCentaur-70B
Decision (2)0.300.320.350.340.370.270.310.330.200.210.240.340.210.310.280.290.280.33
MDP (2)0.560.560.580.580.580.480.530.570.390.460.520.570.420.570.510.530.530.50
Bandit (4)0.320.350.370.360.380.330.330.360.240.230.320.340.280.340.220.250.250.39
Misc. (7)0.280.310.450.460.490.140.300.400.030.110.180.360.090.350.370.410.290.52
Mean (15)0.330.350.430.430.450.260.340.400.160.200.270.380.200.370.340.370.310.46
Figure 8: Reproducibility of Centaur-70B (4-bit) evaluation across Psych-101. Reproduced vs. reported NLL for Centaur-70B (4-bit) across 46 Psych-101 experiments (r=0.994, mean Δ=+0.024, median Δ=+0.006).
Figure 8: Reproducibility of Centaur-70B (4-bit) evaluation across Psych-101. Reproduced vs. reported NLL for Centaur-70B (4-bit) across 46 Psych-101 experiments (r=0.994, mean Δ=+0.024, median Δ=+0.006).
Table 12: Impact of cognitive fine-tuning across different benchmarks. Each cell shows Δ = fine-tuned − base, where fine-tuned models are trained with LoRA r=16 on the full dataset. Significance is assessed with a two-sided z-test; group means use Stouffer’s method to combine per-task z-scores. EQ-Bench is on a separate scale and excluded from means.
QwentaurLlama-Centaur
Benchmark0.6B1.7B4B8B14B1B3B8B
MetaBench
ARC-0.03-0.01-0.01+0.01-0.01+0.01-0.05+0.01
GSM8K-0.25∗∗∗-0.15∗∗∗-0.05-0.17∗∗∗-0.12∗∗∗-0.05∗+0.00-0.10∗
HellaSwag-0.02-0.09-0.02-0.03-0.03-0.04-0.03-0.06
MMLU-0.09-0.05-0.06+0.02+0.02-0.04-0.05+0.01
TruthfulQA-0.01+0.03-0.05+0.05+0.03-0.01+0.04+0.05
Winogrande-0.02-0.02-0.02+0.03-0.02-0.03-0.04-0.01
Mean (6)-0.07∗∗∗-0.05∗-0.03-0.01-0.02-0.03-0.02-0.02
Ethics
CM+0.01+0.00+0.07∗∗∗+0.01+0.04∗∗∗-0.02∗-0.03∗∗+0.00
Deontology+0.02∗-0.03∗-0.02+0.03∗-0.04∗∗∗+0.00+0.05∗∗∗+0.02∗
Justice+0.06∗∗∗+0.00-0.01+0.17∗∗∗-0.08∗∗∗+0.00+0.05∗∗∗+0.03∗
Utilitarian+0.00+0.02∗+0.09∗∗∗+0.03∗∗∗+0.09∗∗∗+0.00-0.04∗∗∗-0.01
Virtue+0.56∗∗∗+0.02∗-0.01+0.03∗∗∗-0.05∗∗∗-0.01+0.20∗∗∗-0.16∗∗∗
Mean (5)+0.13∗∗∗+0.00+0.02∗∗∗+0.05∗∗∗-0.01-0.01+0.04∗∗∗-0.02∗∗∗
Cog. & Lang.
LogiQA+0.00-0.01-0.02-0.02-0.05+0.00-0.02+0.00
PIQA-0.02-0.01+0.01+0.02+0.01-0.01+0.00+0.00
Social IQA+0.00+0.01+0.03∗+0.00+0.01+0.01+0.01+0.01
CoQA (F1)-0.11∗∗∗-0.03-0.02-0.01+0.00-0.07∗∗-0.05∗-0.03
LAMBADA (OAI)+0.00-0.01+0.00+0.00+0.01+0.01+0.02∗+0.01
LAMBADA (Std)+0.02+0.01+0.03∗∗+0.02∗+0.02+0.02∗+0.03∗∗∗+0.02∗
EQ-Bench-44.7∗∗∗+5.1+5.9-21.3∗∗∗-0.3+16.0∗∗∗+22.9∗∗∗+7.4
Mean (6)-0.02-0.01+0.01+0.00+0.00-0.01+0.00+0.00
ACP (Planning)
App (B)-0.10-0.10-0.09-0.02-0.05-0.29∗∗∗+0.08+0.05
Areach (B)-0.04-0.22∗∗∗-0.24∗∗∗-0.24∗∗∗+0.01-0.37∗∗∗+0.14∗+0.02
Just (B)+0.05-0.06-0.06-0.12∗-0.01-0.34∗∗∗-0.13∗-0.06
Land (B)-0.12∗∗-0.44∗∗∗-0.15∗∗-0.21∗∗∗+0.10-0.11∗∗-0.13∗∗-0.04
Figure 9: Specificity of cognitive fine-tuning: comparison with non-cognitive control models. Mean NLL on Psych-101 for cognitively fine-tuned models (Llama-Centaur, Qwentaur; LoRA r=16, full training data) and size-matched non-cognitive and cognitive controls (Hermes, Nemotron, Be.FM). Models are grouped by parameter count. Cognitively fine-tuned models consistently outperform non-cognitive controls at every scale, confirming that the improvement is specific to the behavioural signal in Psych-101 and not an artefact of fine-tuning per se.
Figure 9: Specificity of cognitive fine-tuning: comparison with non-cognitive control models. Mean NLL on Psych-101 for cognitively fine-tuned models (Llama-Centaur, Qwentaur; LoRA r=16, full training data) and size-matched non-cognitive and cognitive controls (Hermes, Nemotron, Be.FM). Models are grouped by parameter count. Cognitively fine-tuned models consistently outperform non-cognitive controls at every scale, confirming that the improvement is specific to the behavioural signal in Psych-101 and not an artefact of fine-tuning per se.
Table 13: Prior ablation conditions expressed in the four-channel notation. ∙ = present, — = removed, Imin = reduced to a minimal action-space definition, ⋅~ = content replaced by generic placeholders with formatting preserved, ⊗ = replaced by an instruction that contradicts the task, ⋅π = trial order permuted.
SourceCondition𝑰𝑺𝑭𝑪Closest condition in ours
11OriginalOriginal
94No psychological taskIminHistory-only
Zero-shot predictionno analogue
57Instruction freeInstruction-ablated
Misleading instructionno analogue
Context freeChoice-only
OursInstruction-ablated
Content-maskedIminS~F~new
History-onlyImin
Choice-only
Order-permutednew
Figure 10: Impact of cognitive fine-tuning on MetaBench performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair on six standard LM benchmarks (ARC, GSM8K, HellaSwag, MMLU, TruthfulQA, Winogrande) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.
Figure 10: Impact of cognitive fine-tuning on MetaBench performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair on six standard LM benchmarks (ARC, GSM8K, HellaSwag, MMLU, TruthfulQA, Winogrande) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.
Table 15: Per-experiment NLL under sequential ablation conditions for Centaur-70B (11) on Psych-101 (in-distribution). Columns correspond to progressive prompt degradation: orig retains the full prompt; inst removes task instructions; cont additionally masks stimulus values and feedback; hist further removes trial structure, leaving only the response history. The ln⁡(k) column shows the random-guessing baseline where k is the number of per-trial response options. Experiments with mixed or continuous response formats have no well-defined k and are shown as – .
ExperimentTypeoriginstconthistln⁡(k)
Gardening task (28)Decision0.490.490.690.690.69
Columbia card task (32)Decision0.210.240.280.290.69
Experiential-symbolic task (34)Decision0.460.460.851.43
Multi-attribute DM (42)Decision0.060.080.740.740.69
Risky choice (50)Decision0.430.640.820.89
choices13k (63)Decision0.430.440.520.560.69
CPC18 (64)Decision0.350.350.380.410.69
Decisions from description (93)Decision0.590.610.690.680.69
Two-step task (48)MDP0.480.531.181.230.69
Two-step task (49)MDP0.530.541.151.160.69
Virtual subway network (84)MDP1.161.461.411.301.61
Multi-task RL (83)MDP0.570.660.770.821.10
Two-step task (104)MDP0.510.520.971.030.69
Drifting four-armed bandit (6)Bandit0.710.850.840.861.39
Horizon task (27)Bandit0.400.390.530.550.69
Two-armed bandit (35)Bandit0.300.360.390.460.69
Prob. instrumental learning (54)Bandit0.500.492.072.010.69
Horizon task (69)Bandit0.580.590.630.680.69
Structured bandit (72)Bandit0.640.680.841.002.08
Horizon task (77)Bandit0.350.360.530.610.69
Iowa gambling task (80)Bandit0.911.080.970.971.39
Horizon task (88)Bandit0.150.150.340.430.69
Horizon task (89)Bandit0.480.480.560.620.69
Spatially correlated MAB (91)Bandit1.821.952.452.623.40
Decisions from experience (93)Bandit0.470.950.540.57
Changing bandit (96)Bandit0.450.451.280.980.69
Cond. assoc. learning (19)Memory0.520.551.031.061.10
Shepard categorization (5)Sup. learn.0.540.580.700.700.69
Multiple-cue judgment (20)Sup. learn.1.141.191.941.942.20
Medin categorization (56)Sup. learn.0.500.580.790.90
Figure 11: Impact of cognitive fine-tuning on Ethics benchmark performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair across five ethical reasoning tasks (commonsense morality, deontology, justice, utilitarianism, virtue) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.
Figure 11: Impact of cognitive fine-tuning on Ethics benchmark performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair across five ethical reasoning tasks (commonsense morality, deontology, justice, utilitarianism, virtue) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.

실제로 확인된 결과

  • 같은 학습 실험 범위 안(분포 내)에서는 0.6B~1B 파라미터 모델이 재현한 700억 파라미터 Centaur-70B와 대등한 예측 성능(음의 로그가능도 기준)을 보였고, 8개 모델 전체가 0.028 nats라는 좁은 성능 범위 안에 몰려 있었다.
  • 학습에 없던 새로운 실험(분포 외, Psych-201-RT)에서는 같은 8개 모델의 성능 차이가 0.244 nats로 훨씬 벌어졌고, 모델이 클수록 뚜렷하게 더 잘했다.
  • 자극과 피드백의 구체적 내용을 지우자 학습된 정보의 75.7%가 사라지고 성능이 무작위 추측 수준 이하로 떨어졌으며, 과제 설명만 지웠을 때는 손실이 12.5%에 그쳤다.
  • 27개 실험 중 18개(67%)는 정보를 지울수록 예측력이 꾸준히 나빠지는 단조적 패턴을 보였다.
  • 시행이 독립적인 THINGS 유사성 판단 과제에서는 순서를 뒤섞어도 예측이 거의 변하지 않았고, 이전 응답이 다음 조건을 결정하는 시간할인 선택 과제에서는 순서를 뒤섞자 예측이 크게 달라졌다.
Figure 12: Impact of cognitive fine-tuning on cognitive and language benchmark performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair across six tasks (LogiQA, PIQA, Social IQA, CoQA, LAMBADA-OpenAI, LAMBADA-Standard) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.
Figure 12: Impact of cognitive fine-tuning on cognitive and language benchmark performance (Δ = fine-tuned − base). Each bar shows the change in accuracy for a matched base–fine-tuned pair across six tasks (LogiQA, PIQA, Social IQA, CoQA, LAMBADA-OpenAI, LAMBADA-Standard) plus their mean. Error bars show pooled standard errors. Positive values (green region) indicate improvement; negative values (red region) indicate degradation. Colour intensity scales with model size within each family.

어디에 쓸 수 있나

  • 대형 모델 대신 소형 파인튜닝 모델을 인간 선택 행동 예측 도구로 활용해 계산 비용을 줄이려는 연구 설계
  • 특정 심리학 실험에서 기존 이론 모델이 어디까지 설명할 수 있는지 비교할 잡음 한계 추정 도구로 활용
  • 프롬프트를 채널별로 나눠 지워보는 방식으로 다른 행동 예측 모델이 실제로 어떤 정보를 쓰는지 진단하는 절차
Figure 18: Complete per-experiment ablation heatmaps for 26 experiments with a well-defined chance baseline. Five additional experiments with mixed or continuous response formats are excluded throughout because no single k defines a chance baseline. The probabilistic instrumental learning task (54) is also omitted. Each cell shows the fraction of learned information lost, δ=(ℒc−ℒorig)/(ln⁡k−ℒorig), averaged across all eight models. Columns represent individual experiments sorted by δ under the history-only condition; rows represent ablation conditions. (Top) 20 experiments with mean δ≤1 under the history-only condition (i.e. performance remains at or above chance). (Bottom) 6 experiments where ablation degrades performance below chance (δ>1); note the separate colour scale. The colour strip below each panel indicates task type (see legend).
Figure 18: Complete per-experiment ablation heatmaps for 26 experiments with a well-defined chance baseline. Five additional experiments with mixed or continuous response formats are excluded throughout because no single k defines a chance baseline. The probabilistic instrumental learning task (54) is also omitted. Each cell shows the fraction of learned information lost, δ=(ℒc−ℒorig)/(ln⁡k−ℒorig), averaged across all eight models. Columns represent individual experiments sorted by δ under the history-only condition; rows represent ablation conditions. (Top) 20 experiments with mean δ≤1 under the history-only condition (i.e. performance remains at or above chance). (Bottom) 6 experiments where ablation degrades performance below chance (δ>1); note the separate colour scale. The colour strip below each panel indicates task type (see legend).

한계와 남은 검증

  • 이 결과는 Psych-101에 포함된 160개 실험 유형과 유사한 과제에서만 검증됐고, 완전히 새로운 유형의 실험에 대한 일반화는 보장되지 않는다.
  • 모든 모델은 저랭크 어댑터(LoRA) 방식으로만 파인튜닝됐고, 전체 파라미터를 다시 학습하는 완전 파인튜닝이나 다른 아키텍처(예: 전문가 혼합, 상태공간 모델)는 테스트되지 않았다.
  • 구조적 절제 실험은 27개 실험에만 적용됐고, 자극 내용과 응답이 분리되지 않는 기억 과제 등 일부 실험 유형은 분석에서 제외됐다.
  • 시행 순서 교환 가능성 테스트는 대조되는 성질을 지닌 딱 2개 실험(THINGS 유사성 판단, 시간할인 선택)에서만 이뤄졌다.
  • 인지 파인튜닝이 수학 추론이나 형식적 계획 같은 일반 능력 지표를 저하시키는 경향이 관찰됐고, 윤리적 추론에 대한 영향은 모델별로 일관되지 않아 해석을 보류했다.

왜 중요한가

지금까지는 인간 행동을 예측하는 대형 AI 모델을 만들려면 700억 파라미터급 거대 모델과 막대한 컴퓨팅 자원이 필요하다고 여겨졌는데, 이 연구는 훨씬 작은 모델로도 학습된 범위 안에서는 비슷한 성능을 낼 수 있음을 보여준다. 심리학 실험 설계자나 인지과학 연구자에게는, 이런 소형 모델이 특정 실험에서 이론이 어디까지 설명해야 하는지 가늠하는 기준선(잡음 한계) 도구로 쓰일 가능성을 제시한다.

이 논문의 용어

  • Psych-101 · 160개 심리학 실험에서 나온 1070만 건의 시행 단위 인간 선택 데이터를 모은 데이터셋
  • 분포 내/분포 외 · 분포 내는 학습에 쓰인 실험과 같은 종류, 분포 외는 학습에 없던 새로운 실험을 뜻한다
  • LoRA (어댑터) · 원래 모델 전체를 다시 학습하지 않고 작은 추가 부품만 학습시켜 성능을 조정하는 경량 파인튜닝 방법
  • 잡음 한계(noise ceiling) · 사람 행동 자체에 존재하는 불가피한 무작위성 때문에 어떤 모델도 그 이상은 예측할 수 없는 상한선
  • 교환 가능성(exchangeability) · 시행들이 서로 독립적이어서 순서를 바꿔도 결과 예측에 영향이 없는 성질

저자 · Nick Oh

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Nick Oh et al., arXiv:2608.05224, CC BY 4.0