월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

LLM 에이전트에게 365일짜리 온라인 쇼핑몰을 맡겨봤더니, 사람보다 훨씬 빨리 자산 관리에 흥미를 잃었다

arXiv:2607.289562026-07-30

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

LLM 에이전트에게 365일짜리 온라인 쇼핑몰을 맡겨봤더니, 사람보다 훨씬 빨리 자산 관리에 흥미를 잃었다

MerchantBench는 실제 이커머스 플랫폼 데이터 98,843개 상품 정보를 바탕으로 만든 365일짜리 가상 온라인 판매점 운영 시뮬레이션이다. 여덟 개 LLM을 두 가지 에이전트 프레임워크로 총 48번 돌려 상품 소싱, 가격 관리, 현금흐름, 지연 피드백 대응 능력을 측정했다. 가장 성적이 좋은 LLM 조합도 사람 참가자 평균 최종 순자산의 27.3%밖에 달성하지 못했다.

METAL LAB 해설 도표

MerchantBench 구조: 사람이 상점을 운영하듯 에이전트를 365일 시험한다

증거 상태측정 결과가 보고됨

  1. 실제 데이터 기반 상품 카탈로그1688 플랫폼의 실제 상품 98,843건, 공급자 36,576곳 데이터를 365일 수요 이력과 함께 시뮬레이션에 이식
  2. 4가지 상시 의사결정상품 소싱, 등록/가격 관리, 현금흐름 관리, 서로 다른 속도로 도착하는 피드백 대응을 26개 도구로 수행
  3. 즉시 신호 vs 지연 신호공급자 이벤트(가격변동·품절 등)는 빨리 보이지만 주문 결과(반품·환불·나쁜 후기)는 늦게 드러나 현금과 평점에 누적 영향
  4. 48회 실험8개 LLM x 2개 프레임워크(ReAct, Hermes) x 3회 반복으로 365일 시뮬레이션 수행, 사람 3명·규칙기반 봇과 비교
  5. 일관성 상실 진단시간이 갈수록 활동이 줄어드는 운영적 일관성 상실과 목표에서 벗어나는 전략적 일관성 상실을 실제 실행 로그로 확인
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 에이전트 벤치마크는 대부분 명확한 성공 기준이 있는 짧은 과제 위주였는데, 이 연구는 정답이 없고 장기간 일관성 있게 판단해야 하는 '오랫동안 흔들리지 않는 판단력(Long-Term Coherence)'을 재는 것을 목표로 삼았다.
  2. 1688(중국 최대 도매 오픈마켓) 실제 상품 98,843건과 공급자 36,576곳 데이터를 바탕으로, 시간당 진행되는 8,760스텝(365일) 시뮬레이션을 만들고 26개의 상점 운영 도구(상품 검색, 등록, 가격 변경, 재무 확인 등)를 에이전트에게 제공했다.
  3. 이 환경의 핵심은 '즉시 보이는 신호'와 '늦게 드러나는 결과'가 섞여 있다는 점이다. 새 주문은 현금을 즉시 묶어두지만, 반품·환불·나쁜 후기 같은 비정상 결과는 한참 뒤에 드러나며 상점 평점에 누적 영향을 준다.
  4. GPT-5.6 Sol, Claude Opus 4.8, Qwen3.7-Max/Plus, GLM-5.2, DeepSeek-V4-Pro/Flash, Kimi K2.6 여덟 모델을 ReAct(최소 도구 제어)와 Hermes(코드 실행·계획·메모리·스킬 관리까지 갖춘 프레임워크) 두 방식으로 각 3회씩, 총 48회 365일 시뮬레이션을 돌렸다.
  5. 결과는 사람 참가자 3명, 규칙 기반(Rule-based) 자동화 봉과 비교됐으며, 최고 성적 LLM 조합조차 사람 평균 최종 순자산의 27.3%에 그쳤고, 시간이 지날수록 활동량이 줄거나(운영적 일관성 상실) 목표에서 벗어나는(전략적 일관성 상실) 패턴이 관찰됐다.
Figure 1: Order-level dynamics couple immediate liquidity pressure with delayed feedback. New orders commit available cash before settlement, whereas abnormal outcomes surface later and affect store rating.
Figure 1: Order-level dynamics couple immediate liquidity pressure with delayed feedback. New orders commit available cash before settlement, whereas abnormal outcomes surface later and affect store rating.
Table 1: Business performance, store reliability, and long-horizon activity after 365 simulated days. Values are means over three runs. Final net assets and GMV are reported in thousands of RMB, total fines in RMB, and rate metrics in percent. SWR denotes Sustained Window Rate. The best result within each framework is shown in bold.
Model or OperatorBusiness PerformanceStore ReliabilityLong-Horizon Activity
Net AssetsGMVProfit MarginOrdersFinesAvg. Store RatingAnomaly RateAvg. Active ListingsSWRTool Calls
ReAct
GPT-5.6 Sol40.8974.1951.39964994.0410.750.099.47,257
Claude Opus 4.831.8969.1044.41,2147964.0512.124.045.01,139
Qwen3.7-Max20.6639.7344.59256723.9016.139.611.1815
Qwen3.7-Plus20.7440.8545.61,0567053.9913.149.952.21,221
GLM-5.225.7360.9037.32,1581,4223.9314.926.053.32,045
DeepSeek-V4-Pro6.568.4041.94502454.0114.423.530.6660
DeepSeek-V4-Flash14.4728.7839.69855174.0414.119.340.6960
Kimi K2.624.9963.6932.92,2301,4743.8915.347.310.61,228
Hermes
GPT-5.6 Sol52.93133.0740.23,2511,0964.099.250.066.14,831
Claude Opus 4.835.5683.2339.91,8081,0894.0211.822.131.71,138
Qwen3.7-Max59.46116.7646.91,9291,2953.9015.749.622.21,366
Qwen3.7-Plus29.4253.6948.99816423.9513.849.919.4820
GLM-5.242.32103.0636.92,7311,4544.0511.349.662.81,792
DeepSeek-V4-Pro16.7131.9543.41,0626653.9814.533.033.3942
DeepSeek-V4-Flash24.6964.5237.61,9891,7743.9316.048.862.21,259
Kimi K2.623.9675.0626.83,3982,6713.7319.148.317.8969
Others
Human217.61608.0635.39,4425,6223.9812.549.1100.08,311
Rule-based24.4853.3740.31,6051,3743.7618.050.0100.03,236
Figure 2: Overview of MerchantBench. The merchant agent coordinates four decision components through shorter supplier feedback and delayed order outcomes over 365 days. The environment combines an upstream supplier simulation, a merchant store, and a downstream order level simulation to evaluate Long-Term Coherence and business outcomes.
Figure 2: Overview of MerchantBench. The merchant agent coordinates four decision components through shorter supplier feedback and delayed order outcomes over 365 days. The environment combines an upstream supplier simulation, a merchant store, and a downstream order level simulation to evaluate Long-Term Coherence and business outcomes.
Table 2: MerchantBench merchant tool inventory.
ToolAccessDescription
Product Sourcing
get_daily_reportReadReturns the daily market report for the current simulation date with market news and opportunity signals
search_productsReadSearches the visible Product Catalog using public fields
get_product_detailReadReturns visible product, logistics, rating, and supplier fields
get_supplier_profileReadReturns the public supplier profile and visible product count
list_supplier_productsReadLists the currently visible products from one supplier
Listing and Pricing Control
list_productWriteAdds products to the store at specified selling prices
delist_productWriteRemoves products from the store
adjust_priceWriteChanges selling prices for active listings
review_my_listingsReadReviews listing age, sales velocity, fines, and fulfillment backlog
query_my_listingsReadReturns current listings with cumulative sales, profit, and fines
query_store_performanceReadSummarizes store outcomes by day or week
query_product_sales_statsReadRanks product outcomes and reports abnormality counts
Cash-Flow Management
query_balanceReadReturns the cash balance, security deposit, funds in transit, receivables, and fines
get_store_snapshotReadSummarizes orders, supply, cash, listings, and store rating
query_platform_rulesReadReturns capital, settlement, penalty, and closure rules
query_cash_pipelineReadSummarizes receivable aging and active order cost exposure
Supplier and Order Monitoring
query_supply_chain_anomaliesReadReturns new or current supplier abnormalities and affected listings
query_my_ordersReadSearches historical orders with logistics and accounting fields
query_open_ordersReadReturns active orders with fulfillment timing and economics
query_order_updatesReadReturns status changes since the previous observation window
query_order_detailReadReturns one order’s full status timeline, accounting, and penalties
Agent Support and Control
read_memory_docReadReads the run local agent memory document
write_memory_docWriteReplaces the run local agent memory document
get_observationReadReturns the current rendered observation
list_toolsReadReturns tool schemas after scenario filtering
Figure 3: Real-world demand patterns and calibrated risk profiles across 98,843 products. Whiskers span the 10th to 90th percentiles, and boxes show interquartile ranges.
Figure 3: Real-world demand patterns and calibrated risk profiles across 98,843 products. Whiskers span the 10th to 90th percentiles, and boxes show interquartile ranges.
Table 5: Built in Hermes tools provided by the official architecture.
ToolAccessDescription
Execution and Files
terminalExecuteExecutes shell commands in a persistent environment
processManageMonitors and controls background processes
execute_codeExecuteRuns Python programs that call Hermes tools and process their outputs
read_fileReadReads text files with line numbers and pagination
write_fileWriteCreates or replaces files and checks supported formats
patchWriteApplies targeted file edits and returns a unified diff
search_filesReadSearches file names and contents
Memory and Skills
memoryWriteStores durable facts that persist across sessions
session_searchReadSearches messages from previous Hermes sessions
skills_listReadLists available skills and their descriptions
skill_viewReadLoads skill instructions and linked resources
skill_manageWriteCreates, revises, or deletes skills
Planning and Coordination
todoManageMaintains the task list for the current session
clarifyInteractRequests clarification, feedback, or a decision from the user
delegate_taskDelegateAssigns independent tasks to subagents
Projects and Output
project_listReadLists available project workspaces
project_createWriteCreates and activates a project workspace
project_switchWriteSwitches the active project workspace
text_to_speechGenerateConverts text into speech audio
image_generateGenerateGenerates or edits images from prompts and references
Figure 4: Final net asset distributions across three repeated runs for each configuration.
Figure 4: Final net asset distributions across three repeated runs for each configuration.
Table 6: Built in Hermes skills and their functions.
SkillCategoryDescription
apple-notesAppleCreates, searches, and edits Apple Notes
apple-remindersAppleAdds, lists, and completes Apple Reminders
findmyAppleTracks Apple devices and AirTags
imessageAppleSends and receives iMessages and SMS
claude-codeAutonomous agentsDelegates coding tasks to Claude Code
codexAutonomous agentsDelegates coding tasks to OpenAI Codex
hermes-agentAutonomous agentsConfigures and extends the Hermes Agent codebase
opencodeAutonomous agentsDelegates coding and review tasks to OpenCode
computer-useGeneralOperates desktop interfaces through visual interaction
architecture-diagramCreativeCreates architecture and infrastructure diagrams
ascii-artCreativeGenerates and transforms ASCII art
ascii-videoCreativeConverts video and audio into ASCII video
baoyu-infographicCreativeProduces infographics using reusable layouts and styles
claude-designCreativeDesigns standalone HTML artifacts
comfyuiCreativeGenerates images, video, and audio with ComfyUI
design-mdCreativeAuthors and validates DESIGN.md specifications
excalidrawCreativeCreates hand drawn Excalidraw diagrams
humanizerCreativeRevises text to remove formulaic AI phrasing
manim-videoCreativeProduces mathematical and algorithmic animations
p5jsCreativeCreates interactive p5.js sketches and generative art
popular-web-designsCreativeApplies established web interface design systems
pretextCreativeSupports interactive creative browser demonstrations
sketchCreativeProduces alternative HTML interface mockups
songwriting-and-ai-musicCreativeSupports songwriting and AI music prompting
touchdesigner-mcpCreativeControls TouchDesigner through an MCP interface
Figure 5: Monthly Operational Coherence profiles for the Human baseline and all eight Hermes models. Panels report effective window rate, environment tool calls, and monthly net profit.
Figure 5: Monthly Operational Coherence profiles for the Human baseline and all eight Hermes models. Panels report effective window rate, environment tool calls, and monthly net profit.
Table 7: Built in Hermes skills and their functions, continued.
SkillCategoryDescription
jupyter-live-kernelData sciencePerforms iterative analysis in a persistent Jupyter kernel
dogfoodGeneralConducts exploratory testing of web applications
himalayaEmailManages email through the Himalaya command line interface
codebase-inspectionGitHubMeasures codebase size, languages, and composition
github-authGitHubConfigures tokens, keys, and command line authentication
github-code-reviewGitHubReviews pull request diffs and inline comments
github-issuesGitHubCreates and manages GitHub issues
github-pr-workflowGitHubManages branches, commits, checks, and pull requests
github-repo-managementGitHubClones, creates, forks, and maintains repositories
gif-searchMediaSearches and downloads GIF content
heartmulaMediaGenerates songs from lyrics and style tags
songseeMediaExtracts and visualizes audio features
youtube-contentMediaConverts YouTube transcripts into written content
huggingface-hubMLOpsSearches, downloads, and uploads models and datasets
evaluating-llms-harnessMLOpsEvaluates language models with standard benchmarks
weights-and-biasesMLOpsTracks experiments, sweeps, and model artifacts
llama-cppMLOpsRuns local GGUF model inference
serving-llms-vllmMLOpsServes language models with vLLM
audiocraft-audio-generationMLOpsGenerates music and sound with AudioCraft
segment-anything-modelMLOpsPerforms prompt based image segmentation
obsidianNote takingReads, searches, creates, and edits Obsidian notes
Figure 6: Monthly Product Sourcing across four representative models. Scores are listing-hour-weighted catalog demand percentiles of active products in each month. Lines and bands show repeat means and standard deviations.
Figure 6: Monthly Product Sourcing across four representative models. Scores are listing-hour-weighted catalog demand percentiles of active products in each month. Lines and bands show repeat means and standard deviations.
Table 8: Built in Hermes skills and their functions, continued.
SkillCategoryDescription
airtableProductivityManages Airtable records and queries
google-workspaceProductivityOperates Gmail, Calendar, Drive, Docs, and Sheets
mapsProductivityProvides geocoding, points of interest, routes, and time zones
nano-pdfProductivityEdits PDF text and document metadata
notionProductivityManages Notion pages and databases
ocr-and-documentsProductivityExtracts text from PDFs and scanned documents
petdexProductivityInstalls and selects animated Hermes mascots
powerpointProductivityCreates and edits presentation decks
teams-meeting-pipelineProductivityOperates the Teams meeting summary pipeline
arxivResearchSearches arXiv by topic, author, category, or identifier
blogwatcherResearchMonitors blogs and syndicated feeds
llm-wikiResearchBuilds and queries an interlinked knowledge base
polymarketResearchQueries prediction markets, prices, and order books
research-paper-writingResearchSupports machine learning paper development and submission
openhueSmart homeControls Philips Hue lights, rooms, and scenes
xurlSocial mediaReads and operates X through its command line interface
hermes-agent-skill-authoringSoftware developmentAuthors and validates Hermes skill packages
node-inspect-debuggerSoftware developmentDebugs Node.js through the inspector protocol
planSoftware developmentProduces actionable implementation plans
python-debugpySoftware developmentDebugs Python with pdb and debugpy
requesting-code-reviewSoftware developmentPerforms structured review before integration
simplify-codeSoftware developmentRefines recent code changes with parallel review
spikeSoftware developmentRuns disposable experiments before implementation
systematic-debuggingSoftware developmentApplies a structured root cause debugging process
test-driven-developmentSoftware developmentApplies test driven development workflows
yuanbaoGeneralOperates Yuanbao groups and member queries
Figure 11: Composition and calibrated distributions of the filtered data. Panel (a) reports the numbers of Product Records and unique supplier IDs within each category. Panel (b) shows procurement price distributions within each category. Panel (c) shows the joint distribution of Downstream Order Outcome probability and Upstream Supplier Event intensity.
Figure 11: Composition and calibrated distributions of the filtered data. Panel (a) reports the numbers of Product Records and unique supplier IDs within each category. Panel (b) shows procurement price distributions within each category. Panel (c) shows the joint distribution of Downstream Order Outcome probability and Upstream Supplier Event intensity.
Table 19: Product fields and their visibility to the merchant agent. Catalog results also include the public supplier fields in Table 20.
Product fieldAccessMeaning
product_idVisibleStable identifier for a Product in the Product Catalog
nameVisibleMarketplace product title used for retrieval and comparison
categoryVisibleOne of the ten normalized first level product categories
quantityVisibleCurrent effective supplier inventory after replenishment
priceVisibleCurrent procurement price offered by the supplier
historical_avg_ratingVisibleHistorical product rating obtained from the source platform
logistics_hoursVisibleBaseline transit time from supplier dispatch to delivery
is_listed_by_supplierVisibleCurrent procurement availability, exposed as supplier_available
ref_priceHiddenReference price used in the price response term of the demand model
base_priceHiddenSupplier price restored after a temporary Price Change ends
cancel_rateHiddenProduct level probability used to sample Cancellation
refund_rateHiddenProduct level probability used to sample Return and Refund
only_refund_rateHiddenProduct level probability used to sample Returnless Refund
bad_review_rateHiddenProduct level probability used to sample Bad Review
max_quantityHiddenInventory capacity used by the supplier replenishment process
hourly_incrementHiddenHourly supplier inventory replenishment amount
elasticityHiddenProduct specific price elasticity used by the demand model
market_curveHiddenReal-world product level demand history over 365 days
quantity_updated_tHiddenInternal timestamp used for lazy inventory replenishment
price_recover_tHiddenPrescheduled end time of an active Price Change
delist_recover_tHiddenPrescheduled end time of an active Product Delisting
Figure 12: Temporal demand patterns in the filtered data. The top panel presents aggregate daily demand and its seven day mean. The bottom panels show normalized demand histories for representative red envelope, electric fan, and hot water bag Product Records.
Figure 12: Temporal demand patterns in the filtered data. The top panel presents aggregate daily demand and its seven day mean. The bottom panels show normalized demand histories for representative red envelope, electric fan, and hot water bag Product Records.
Table 20: Supplier fields and their visibility to the merchant agent. Supplier trust attributes are constant across Products sharing the same supplier_id, while event hazards are calibrated at the Product level.
Supplier fieldAccessMeaning
supplier_idVisibleStable supplier identifier
supplier_nameVisiblePublic supplier name
shop_ratingVisiblePublic supplier rating shared by all Products from the supplier
return_buyer_rateVisiblePublic repeat buyer rate returned by the supplier profile
supplier_age_yearsVisiblePublic supplier tenure in years
product_countVisibleNumber of currently available Products from the supplier
supplier_ship_hoursVisibleCurrent dispatch time for a Product from this supplier
base_ship_hoursHiddenDispatch time restored after a Shipment Delay ends
timeout_rateHiddenProduct level hazard for Shipment Delay
price_change_rateHiddenProduct level hazard for Price Change
supplier_delist_rateHiddenProduct level hazard for Product Delisting
timeout_activeHiddenInternal indicator of an active Shipment Delay
timeout_recover_tHiddenPrescheduled end time of an active Shipment Delay
Figure 13: Monthly Operational Coherence profiles for the Human baseline and each of the eight Hermes models. Panels (a), (b), (c), and (d) report effective window rate, environment tool calls, month end active listings, and monthly net profit.
Figure 13: Monthly Operational Coherence profiles for the Human baseline and each of the eight Hermes models. Panels (a), (b), (c), and (d) report effective window rate, environment tool calls, month end active listings, and monthly net profit.
Table 21: Order fields and their visibility to the merchant agent. Some visible lifecycle fields remain empty until realization, while presampled future outcomes and internal schedules remain hidden.
Order fieldAccessMeaning
order_idVisibleStable identifier for an individual customer order
product_id, product_nameVisibleProduct identity associated with the order
supplier_id, supplier_nameVisibleSupplier identity associated with the order
order_timeVisibleCalendar and simulation time at which the order was placed
current_statusVisibleLatest realized lifecycle state
status_age_hoursVisibleElapsed time since the latest realized status transition
expected_delivery_timeVisibleCurrent delivery estimate computed from realized timing information
delivered_timeVisibleMerchant facing delivery timestamp populated after delivery
late_timeVisibleMerchant facing timestamp populated only after Late Shipment is realized
sale_priceVisibleMerchant selling price recorded when the order was created
purchase_priceVisibleProcurement price recorded when the order was created
supplier_ship_hoursVisibleSupplier dispatch duration recorded for the order
supplier_logistics_hoursVisibleBaseline post dispatch logistics duration
actual_logistics_hoursVisibleRealized transit duration populated after delivery
realized_revenueVisibleRevenue credited from outcomes realized so far
realized_costVisibleProcurement cost realized so far
total_penaltyVisibleSum of penalties already applied to the order
net_profitVisibleRealized revenue minus realized cost and total penalty
profit_finalizedVisibleIndicator that no further profit component remains unresolved
status_logVisibleRealized sequence of lifecycle states and their timestamps
preset_anomalyHiddenPresampled future outcome among normal fulfillment and four customer abnormalities
preset_anomaly_tHiddenInternal realization time of the presampled abnormal outcome
settlement_delay_stepsHiddenPresampled delay from delivery to final settlement
purchase_t, shipped_t, delivered_t, settled_tHiddenRaw internal transition times, with only realized merchant facing views exposed
Figure 14: Monthly Operational Coherence profiles for the Human baseline and each of the eight ReAct models. Panels (a), (b), (c), and (d) report effective window rate, environment tool calls, month end active listings, and monthly net profit.
Figure 14: Monthly Operational Coherence profiles for the Human baseline and each of the eight ReAct models. Panels (a), (b), (c), and (d) report effective window rate, environment tool calls, month end active listings, and monthly net profit.

실제로 확인된 결과

  • 최고 성적 LLM 조합(Hermes+Qwen3.7-Max)도 사람 참가자 평균 최종 순자산의 27.3%만 달성했다.
  • 여덟 모델 평균으로 볼 때 Hermes 프레임워크가 ReAct보다 최종 순자산 53.3%, GMV 71.5%, 주문 수 71.2% 더 높았고, 8개 모델 중 7개에서 Hermes가 더 나은 성과를 냈다(Kimi K2.6만 예외로 4.1% 낮음).
  • 사람 참가자는 SWR(지속 활동 비율)이 100%였던 반면, LLM 구성들은 ReAct에서 10.6~99.4%, Hermes에서 17.8~66.1%로 활동이 시간이 갈수록 줄어드는 경향을 보였다.
  • 일부 모델(예: Hermes Claude Opus 4.8)은 근거 없이 '상품 수를 줄이면 트래픽이 집중된다'는 잘못된 정책을 유지해 활성 매물이 47개에서 3개로 줄었고, 한 Kimi K2.6 실행에서는 104일째 '회복 불가' 판단 후 남은 523개 결정 구간 중 355개에서 아무 조치도 취하지 않았다.
  • 사람 참가자는 자금 여유가 생기면서 조달 가격대를 초반 43.4~53.1위안에서 후반 58.7~90.8위안으로 넓혔지만, GLM-5.2, DeepSeek-V4-Flash, Kimi K2.6 등은 가격대가 거의 변하지 않았다.
Figure 15: Daily net asset curves for all eight models under ReAct over 365 simulated days.
Figure 15: Daily net asset curves for all eight models under ReAct over 365 simulated days.

어디에 쓸 수 있나

  • 장기간 자율적으로 재고·가격·현금흐름을 관리해야 하는 자동화 에이전트를 설계할 때, 시간이 지나며 활동이 줄거나 목표에서 이탈하는지 점검하는 평가 틀로 참고할 수 있다.
  • 여러 LLM과 에이전트 프레임워크(도구 호출 방식, 메모리·스킬 관리 유무)의 조합이 장기 성과에 어떤 차이를 만드는지 비교하는 실험 설계에 참고할 수 있다.
  • 지연된 피드백(반품, 환불, 후기)이 누적되는 상황에서 에이전트가 과거 결정을 되짚어 수정하는지 진단하는 방법론으로 활용할 수 있다.
Figure 16: Daily net asset curves for all eight models under Hermes over 365 simulated days.
Figure 16: Daily net asset curves for all eight models under Hermes over 365 simulated days.

한계와 남은 검증

  • 시뮬레이션은 1688 플랫폼 데이터를 바탕으로 한 가상 환경으로, 실제 온라인 판매점 운영과는 규제·경쟁·고객 행동 등에서 차이가 있을 수 있다.
  • 평가는 3회 반복 실행에 기반하며, 특히 Qwen3.7-Max+Hermes처럼 변동계수 55.1%로 매우 불안정한 조합도 있어 결과 해석에 주의가 필요하다.
  • 사람 참가자는 단 3명으로 전문 판매자가 아닌 비경험자였기 때문에, '사람 대비 27.3%'라는 수치를 일반적인 인간 전문가 성과 기준으로 확대 해석하기 어렵다.
  • 논문에서 언급된 GPT-5.6, Claude Opus 4.8, Qwen3.7 시리즈 등은 실제 공개된 모델명과 다를 수 있어 특정 최신 모델 성능으로 그대로 받아들이기보다 벤치마크 설계 자체에 주목해야 한다.
  • 코드 실행이나 스킬 생성 같은 Hermes의 부가 기능이 왜 특정 모델에서만 효과를 내는지에 대한 원인 분석은 사례 관찰 수준이며 추가 검증이 필요하다.

왜 중요한가

챗봇이나 짧은 작업 자동화를 넘어, 실제 업무처럼 몇 달간 스스로 판단하고 수정해야 하는 상황에서 지금의 LLM 에이전트가 얼마나 미덥지 못한지를 구체적 수치로 보여준다. 재고관리·판매 운영처럼 장기간 자율적으로 맡길 자동화 시스템을 설계하는 사람에게, 어디서 실패가 쌓이는지(활동 감소, 조기 포기, 근거 없는 정책 유지)를 미리 알려주는 참고 자료가 된다.

이 논문의 용어

  • Long-Term Coherence · 긴 시간 동안 목적을 잃지 않고 누적된 증거에 맞춰 판단을 조정하는 능력
  • POMDP(부분 관측 마르코프 결정 과정) · 에이전트가 전체 상태를 다 보지 못하고 일부 관측 정보만으로 결정을 내려야 하는 의사결정 모델
  • Sustained Window Rate(SWR) · 30일 단위로 미끄러지며 확인했을 때, 도구를 한 번이라도 사용한 결정 구간의 최소 비율
  • Hermes / ReAct · ReAct는 26개 도구만 쓰는 최소 제어 방식, Hermes는 여기에 코드 실행·메모리·스킬 관리 등을 더한 확장 프레임워크
  • Upstream Supplier Events / Downstream Order Outcomes · 공급자 쪽에서 빨리 드러나는 사건(가격변동·품절 등)과 주문 쪽에서 늦게 드러나는 결과(반품·환불 등)

저자 · Qiming Shi

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Qiming Shi et al., arXiv:2607.28956, arxiv-nonexclusive