매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study

arXiv:2608.197902026-08-21

신소재 후보를 고를 때, 챗GPT 같은 AI가 전문 통계기법을 대신할 수 있을까 실험해봤다

연구진은 값비싼 실험 없이도 좋은 신소재를 빨리 찾아야 하는 문제에서, 오픈소스 대형언어모델(LLM) 5종이 다음에 어떤 후보를 실험할지 스스로 고르게 해봤다. 무작위로 고르는 것보다는 훨씬 나았지만, 이 분야에서 오래 쓰여온 통계적 방법인 가우시안 프로세스보다는 대체로 못했다. 후보를 어떻게 나열해서 보여주는지, 소재 관련 설명을 곁들이는지에 따라 성능이 크게 흔들려서, 아직은 안정적으로 믿고 쓰기는 이르다는 결론이다.

무엇을 했나

  1. 강철, 자석 합금, 압전 세라믹 등 4가지 소재 최적화 문제에서 정답(최적 후보)을 찾는 실험을 재현하는 방식으로, Gemma·DeepSeek·Qwen 계열 등 5개 오픈소스 LLM이 매번 후보 목록 중 하나를 고르게 하고 몇 번 만에 최적값을 찾는지 측정했다
  2. 후보를 한 번에 다 보여주는 방식과, 후보를 여러 조로 나눠 토너먼트처럼 승자를 뽑아 올라가는 방식(배치 토너먼트) 두 가지 제시 방법을 비교했다
  3. LLM은 무작위 선택보다는 항상 더 적은 시도로 정답을 찾았지만, 전통적인 통계기법인 가우시안 프로세스보다는 4개 과제 중 3개에서 뒤처졌고, 나머지 1개 과제에서만 앞서거나 비슷했다
  4. 같은 LLM이라도 후보 목록의 순서, 배치 크기, 소재 관련 설명(원소 이름 등) 유무에 따라 성적이 크게 달라졌고, 특히 목록의 특정 위치에 있는 후보를 유독 자주 고르는 '위치 편향' 현상도 모델마다 다르게 나타났다
  5. LLM의 선택 패턴을 뜯어봤더니 가우시안 프로세스처럼 규칙적으로 탐색과 활용 사이를 오가는 방식이 아니라 훨씬 불규칙했다
Figure 1: Best-so-far optimization trajectories using Gemma 4 31B for the LLM policies. At iteration t, each curve shows the mean percentage gap 100​(y⋆−bt)/|y⋆| across the 25 shared initialization seeds. Shading denotes pointwise normal-approximation 95% confidence intervals, computed as g¯t±1.96​st/25. We show Gemma 4 31B because it ranks among the strongest LLMs across all four tasks.
Figure 1: Best-so-far optimization trajectories using Gemma 4 31B for the LLM policies. At iteration t, each curve shows the mean percentage gap 100​(y⋆−bt)/|y⋆| across the 25 shared initialization seeds. Shading denotes pointwise normal-approximation 95% confidence intervals, computed as g¯t±1.96​st/25. We show Gemma 4 31B because it ranks among the strongest LLMs across all four tasks.
Table 1: Retrospective optimization tasks.
TaskCandidatesFeaturesTarget
Kerr Rotation9213Kerr rotation (mrad)
Coercivity9213coercivity (mT)
Matbench Steels31214yield strength (MPa)
Electrostrain8115electrostrain (%)
Figure 2: Batch-size sensitivity of tournament selection on Matbench Steels and Kerr Rotation for runs of 25 seeds. Curves report the mean number of iterations to the optimum across seeds, with uncertainty shown shaded as the mean ± std/nseeds.
Figure 2: Batch-size sensitivity of tournament selection on Matbench Steels and Kerr Rotation for runs of 25 seeds. Curves report the mean number of iterations to the optimum across seeds, with uncertainty shown shaded as the mean ± std/nseeds.
Table 2: Iterations to the global optimum over 25 initializations (mean ± standard deviation), except for random selection which is initialized over 1000 seeds. Lower is better. Baselines are shown once because they do not depend on an LLM presentation protocol.
ModelKerr RotationCoercivityMatbench SteelsElectrostrain
Random462.48 ± 262.24448.09 ± 269.63159.31 ± 91.2839.89 ± 23.00
GP-EI8.40 ± 2.7891.80 ± 62.7742.72 ± 27.2418.48 ± 14.53
Whole-pool selection
Gemma431.28 ± 17.24209.88 ± 134.0544.16 ± 23.3614.36 ± 7.71
Qw27B27.08 ± 13.59195.20 ± 114.3948.24 ± 23.4319.44 ± 10.40
Qw35B45.72 ± 20.28289.88 ± 127.4450.36 ± 24.8927.64 ± 18.15
Qw397B23.48 ± 10.40∗258.24 ± 151.68∗53.32 ± 32.90∗16.84 ± 5.93
DS110.64 ± 85.31236.28 ± 210.7287.60 ± 52.1013.88 ± 9.98
Batch-tournament (b=20)
Gemma413.12 ± 5.12389.64 ± 92.0078.08 ± 49.8314.80 ± 6.50
Qw27B14.28 ± 7.33468.64 ± 50.6864.52 ± 32.9418.04 ± 5.33
Qw35B22.00 ± 8.73379.80 ± 110.2752.68 ± 20.1220.36 ± 9.89
Qw397B14.48 ± 7.54178.80 ± 82.87†81.96 ± 41.1015.08 ± 6.92
DS16.88 ± 5.52241.84 ± 52.1032.88 ± 31.7112.36 ± 4.71
(a) Position distributions under batch-tournament and whole-pool selection.
(a) Position distributions under batch-tournament and whole-pool selection.
Table 3: Position-bias metrics for all model-task combinations. We report DKL for whole-pool and batch-tournament selection with b=20, except for the proxy configurations marked with an asterisk. Missing measurements are denoted by ’-’.
ElectrostrainMatbench SteelsCoercivityKerr Rotation
ModelsWholeBatchWholeBatchWholeBatchWholeBatch
Gemma40.9982.271.022.111.602.082.212.02
DS1.362.291.472.132.482.352.602.15
Qw27B0.9292.100.8522.042.182.062.172.04
Qw35B1.262.200.7472.121.672.242.772.08
Qw397B1.042.201.36∗2.081.84†2.06
Figure 4: Optimization performance on Matbench Steels, with and without materials context across 25 initializations. Bars report the mean iteration at which the optimum is reached, with normal-approximation 95% confidence intervals, computed as mean ± 1.96​std/nseeds. We also report the mean iterations to optimum of random selection with a dotted line.
Figure 4: Optimization performance on Matbench Steels, with and without materials context across 25 initializations. Bars report the mean iteration at which the optimum is reached, with normal-approximation 95% confidence intervals, computed as mean ± 1.96​std/nseeds. We also report the mean iterations to optimum of random selection with a dotted line.
Table 4: Generation-and-matching performance on Matbench Steels compared with random, GP-EI, and whole-pool selection. We report iterations to the global optimum over 25 initializations (mean ± standard deviation), except for random selection which is initialized over 1000 seeds.
ModelRandomGP-EIWhole-poolGeneration-and-matching
Qw27B159.31 ± 91.2842.72 ± 27.2448.24 ± 23.43123.60 ± 63.60
Gemma444.16 ± 23.3648.80 ± 26.14
Figure 5: Explorative behavior of two selection protocols on Matbench Steels for seed 4 with Gemma 4 31B. Dots represent points of the pool projected onto the first two principal components, with the PCA being scaled over every pool entry. On this seed, GP-EI and batch-tournament reach the optimum after 60 and 56 iterations, respectively.
Figure 5: Explorative behavior of two selection protocols on Matbench Steels for seed 4 with Gemma 4 31B. Dots represent points of the pool projected onto the first two principal components, with the PCA being scaled over every pool entry. On this seed, GP-EI and batch-tournament reach the optimum after 60 and 56 iterations, respectively.

왜 중요한가

신소재 개발처럼 시험 한 번에 비용과 시간이 많이 드는 분야에서, 별도의 훈련 없이 범용 AI 모델만으로 다음 실험 대상을 고를 수 있는지는 실무적으로 큰 의미가 있다. 다만 이번 연구는 그 가능성과 한계를 동시에 보여주므로, 지금 단계에서 LLM을 실제 소재 탐색 파이프라인에 무작정 투입하기보다는 신중한 검증이 필요함을 시사한다.

Figure 7: Batch-size sensitivity of tournament selection on Coercivity and Electrostrain. Curves report the mean number of iterations to the optimum across seeds, with uncertainty shown shaded as the mean ± std/nseeds.
Figure 7: Batch-size sensitivity of tournament selection on Coercivity and Electrostrain. Curves report the mean number of iterations to the optimum across seeds, with uncertainty shown shaded as the mean ± std/nseeds.

이 논문의 용어

  • 능동학습(Active Learning) · 적은 수의 관측 결과를 보고 다음에 어떤 데이터를 확인할지 스스로 정하는 학습 방식
  • 가우시안 프로세스(Gaussian Process, GP) · 불확실성까지 함께 예측해주는 확률적 통계 모델로, 신소재 탐색에서 널리 쓰이는 대표적 기법
  • 획득 함수(Acquisition Function) · 다음에 어떤 후보를 실험할지 점수를 매겨 결정하는 규칙, 대표적으로 기대개선량(EI)이 있다
  • 오픈웨이트(Open-weight) LLM · 모델 내부 가중치가 공개되어 누구나 내려받아 실행할 수 있는 대형언어모델
  • 위치 편향(Position Bias) · AI가 목록에서 특정 위치(예: 맨 앞)에 있는 항목을 실제 중요도와 상관없이 더 자주 고르는 경향
Figure 8: Comparison of evolution of fitted β^t on Matbench Steels between Gemma 4 31B, Random and GP-EI. Results are averaged across seeds for each iteration, when at least half of the seeds have not converged yet.
Figure 8: Comparison of evolution of fitted β^t on Matbench Steels between Gemma 4 31B, Random and GP-EI. Results are averaged across seeds for each iteration, when at least half of the seeds have not converged yet.

논문 원문 초록 (영문)

Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this setting. We evaluate five LLMs across four retrospective finite-pool materials optimization tasks under different candidate-presentation strategies and compare them with random selection and conventional Gaussian-process methods. LLM policies generally reach the global optimum in fewer iterations than random selection, indicating that they provide a useful acquisition signal without task-specific training. Their performance relative to Gaussian-process methods is mixed: conventional acquisition performs better on most tasks, while LLMs match or outperform it in some settings. Performance varies substantially across tasks, models, initializations, and candidate presentations, with no LLM approach performing best across all tasks. Overall, open-weight LLMs show potential as acquisition policies for finite-pool materials search, although their reliability remains sensitive to the task and to how candidates and scientific context are presented.

저자 · Dino-Rober Demir, Florian Le Bronnec, Rio Yokota

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Dino-Rober Demir et al., arXiv:2608.19790, CC BY 4.0