MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
让LLM智能体经营365天网店,结果大多数比人类更快就撒手不管了
MerchantBench是一个基于98,843条真实电商商品数据构建的365天模拟网店环境,用来测试LLM智能体能否在长期持续运营中保持连贯合理的判断。研究让8个LLM在两种智能体框架下各跑三次,共48次365天模拟,考察商品选品、定价、现金流和延迟反馈应对能力。即便表现最好的LLM配置,最终净资产也只达到人类参与者平均水平的27.3%。
METAL LAB 解读图
MerchantBench如何用模拟一年来考验智能体的持久力
证据状态已报告实测结果
- 真实数据商品目录来自1688平台的98,843条真实商品记录与36,576家供应商数据,每条记录附带365天需求历史,被移植进模拟环境
- 四项持续性决策商品选品、上架与定价管理、现金流管理、应对不同延迟的反馈,均通过26种工具完成
- 快信号与慢信号供应商事件(涨价、缺货)很快可见,而订单结果(退货、退款、差评)延迟显现,逐渐侵蚀现金和店铺评分
- 48次评测运行8个LLM x 2种框架(ReAct、Hermes)x 3次重复,共48次365天模拟,并与3名人类参与者及规则型机器人对比
- 连贯性丧失诊断通过决策轨迹记录揭示运营连贯性丧失(活动逐渐减少)与策略连贯性丧失(目标漂移)两种失败模式
他们做了什么
- 以往的智能体评测大多是有明确成败标准的短任务,这项研究关注的是长期连贯性(Long-Term Coherence),即能否在漫长过程中持续追求目标并根据累积证据调整决策。
- 环境基于1688(中国大型批发电商平台)的98,843条真实商品记录和36,576家供应商数据构建,按小时推进共8,760步(相当于365天),智能体可使用26种店铺管理工具(搜索、上架、改价、查财务等)。
- 环境的核心设计是反馈延迟不一:新订单会立刻占用现金,但取消、退款、差评等异常结果要延后才会显现,并逐渐侵蚀店铺评分。
- 研究对GPT-5.6 Sol、Claude Opus 4.8、Qwen3.7-Max/Plus、GLM-5.2、DeepSeek-V4-Pro/Flash、Kimi K2.6共8个模型,分别在ReAct(仅用26种工具的最简控制器)和Hermes(附带代码执行、规划、记忆、技能管理的完整框架)下各跑三次,共48次365天模拟。
- 结果与三名人类参与者和一个规则型机器人对比后发现,表现最好的LLM配置最终净资产也只达到人类平均水平的27.3%,且轨迹显示出运营连贯性丧失(活动逐渐减少)和策略连贯性丧失(偏离盈利目标)两种模式。
| Model or Operator | Business Performance | Store Reliability | Long-Horizon Activity | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Net Assets | GMV | Profit Margin | Orders | Fines | Avg. Store Rating | Anomaly Rate | Avg. Active Listings | SWR | Tool Calls | |
| ReAct | ||||||||||
| GPT-5.6 Sol | 40.89 | 74.19 | 51.3 | 996 | 499 | 4.04 | 10.7 | 50.0 | 99.4 | 7,257 |
| Claude Opus 4.8 | 31.89 | 69.10 | 44.4 | 1,214 | 796 | 4.05 | 12.1 | 24.0 | 45.0 | 1,139 |
| Qwen3.7-Max | 20.66 | 39.73 | 44.5 | 925 | 672 | 3.90 | 16.1 | 39.6 | 11.1 | 815 |
| Qwen3.7-Plus | 20.74 | 40.85 | 45.6 | 1,056 | 705 | 3.99 | 13.1 | 49.9 | 52.2 | 1,221 |
| GLM-5.2 | 25.73 | 60.90 | 37.3 | 2,158 | 1,422 | 3.93 | 14.9 | 26.0 | 53.3 | 2,045 |
| DeepSeek-V4-Pro | 6.56 | 8.40 | 41.9 | 450 | 245 | 4.01 | 14.4 | 23.5 | 30.6 | 660 |
| DeepSeek-V4-Flash | 14.47 | 28.78 | 39.6 | 985 | 517 | 4.04 | 14.1 | 19.3 | 40.6 | 960 |
| Kimi K2.6 | 24.99 | 63.69 | 32.9 | 2,230 | 1,474 | 3.89 | 15.3 | 47.3 | 10.6 | 1,228 |
| Hermes | ||||||||||
| GPT-5.6 Sol | 52.93 | 133.07 | 40.2 | 3,251 | 1,096 | 4.09 | 9.2 | 50.0 | 66.1 | 4,831 |
| Claude Opus 4.8 | 35.56 | 83.23 | 39.9 | 1,808 | 1,089 | 4.02 | 11.8 | 22.1 | 31.7 | 1,138 |
| Qwen3.7-Max | 59.46 | 116.76 | 46.9 | 1,929 | 1,295 | 3.90 | 15.7 | 49.6 | 22.2 | 1,366 |
| Qwen3.7-Plus | 29.42 | 53.69 | 48.9 | 981 | 642 | 3.95 | 13.8 | 49.9 | 19.4 | 820 |
| GLM-5.2 | 42.32 | 103.06 | 36.9 | 2,731 | 1,454 | 4.05 | 11.3 | 49.6 | 62.8 | 1,792 |
| DeepSeek-V4-Pro | 16.71 | 31.95 | 43.4 | 1,062 | 665 | 3.98 | 14.5 | 33.0 | 33.3 | 942 |
| DeepSeek-V4-Flash | 24.69 | 64.52 | 37.6 | 1,989 | 1,774 | 3.93 | 16.0 | 48.8 | 62.2 | 1,259 |
| Kimi K2.6 | 23.96 | 75.06 | 26.8 | 3,398 | 2,671 | 3.73 | 19.1 | 48.3 | 17.8 | 969 |
| Others | ||||||||||
| Human | 217.61 | 608.06 | 35.3 | 9,442 | 5,622 | 3.98 | 12.5 | 49.1 | 100.0 | 8,311 |
| Rule-based | 24.48 | 53.37 | 40.3 | 1,605 | 1,374 | 3.76 | 18.0 | 50.0 | 100.0 | 3,236 |
| Tool | Access | Description |
|---|---|---|
| Product Sourcing | ||
| get_daily_report | Read | Returns the daily market report for the current simulation date with market news and opportunity signals |
| search_products | Read | Searches the visible Product Catalog using public fields |
| get_product_detail | Read | Returns visible product, logistics, rating, and supplier fields |
| get_supplier_profile | Read | Returns the public supplier profile and visible product count |
| list_supplier_products | Read | Lists the currently visible products from one supplier |
| Listing and Pricing Control | ||
| list_product | Write | Adds products to the store at specified selling prices |
| delist_product | Write | Removes products from the store |
| adjust_price | Write | Changes selling prices for active listings |
| review_my_listings | Read | Reviews listing age, sales velocity, fines, and fulfillment backlog |
| query_my_listings | Read | Returns current listings with cumulative sales, profit, and fines |
| query_store_performance | Read | Summarizes store outcomes by day or week |
| query_product_sales_stats | Read | Ranks product outcomes and reports abnormality counts |
| Cash-Flow Management | ||
| query_balance | Read | Returns the cash balance, security deposit, funds in transit, receivables, and fines |
| get_store_snapshot | Read | Summarizes orders, supply, cash, listings, and store rating |
| query_platform_rules | Read | Returns capital, settlement, penalty, and closure rules |
| query_cash_pipeline | Read | Summarizes receivable aging and active order cost exposure |
| Supplier and Order Monitoring | ||
| query_supply_chain_anomalies | Read | Returns new or current supplier abnormalities and affected listings |
| query_my_orders | Read | Searches historical orders with logistics and accounting fields |
| query_open_orders | Read | Returns active orders with fulfillment timing and economics |
| query_order_updates | Read | Returns status changes since the previous observation window |
| query_order_detail | Read | Returns one order’s full status timeline, accounting, and penalties |
| Agent Support and Control | ||
| read_memory_doc | Read | Reads the run local agent memory document |
| write_memory_doc | Write | Replaces the run local agent memory document |
| get_observation | Read | Returns the current rendered observation |
| list_tools | Read | Returns tool schemas after scenario filtering |
| Tool | Access | Description |
|---|---|---|
| Execution and Files | ||
| terminal | Execute | Executes shell commands in a persistent environment |
| process | Manage | Monitors and controls background processes |
| execute_code | Execute | Runs Python programs that call Hermes tools and process their outputs |
| read_file | Read | Reads text files with line numbers and pagination |
| write_file | Write | Creates or replaces files and checks supported formats |
| patch | Write | Applies targeted file edits and returns a unified diff |
| search_files | Read | Searches file names and contents |
| Memory and Skills | ||
| memory | Write | Stores durable facts that persist across sessions |
| session_search | Read | Searches messages from previous Hermes sessions |
| skills_list | Read | Lists available skills and their descriptions |
| skill_view | Read | Loads skill instructions and linked resources |
| skill_manage | Write | Creates, revises, or deletes skills |
| Planning and Coordination | ||
| todo | Manage | Maintains the task list for the current session |
| clarify | Interact | Requests clarification, feedback, or a decision from the user |
| delegate_task | Delegate | Assigns independent tasks to subagents |
| Projects and Output | ||
| project_list | Read | Lists available project workspaces |
| project_create | Write | Creates and activates a project workspace |
| project_switch | Write | Switches the active project workspace |
| text_to_speech | Generate | Converts text into speech audio |
| image_generate | Generate | Generates or edits images from prompts and references |
| Skill | Category | Description |
|---|---|---|
| apple-notes | Apple | Creates, searches, and edits Apple Notes |
| apple-reminders | Apple | Adds, lists, and completes Apple Reminders |
| findmy | Apple | Tracks Apple devices and AirTags |
| imessage | Apple | Sends and receives iMessages and SMS |
| claude-code | Autonomous agents | Delegates coding tasks to Claude Code |
| codex | Autonomous agents | Delegates coding tasks to OpenAI Codex |
| hermes-agent | Autonomous agents | Configures and extends the Hermes Agent codebase |
| opencode | Autonomous agents | Delegates coding and review tasks to OpenCode |
| computer-use | General | Operates desktop interfaces through visual interaction |
| architecture-diagram | Creative | Creates architecture and infrastructure diagrams |
| ascii-art | Creative | Generates and transforms ASCII art |
| ascii-video | Creative | Converts video and audio into ASCII video |
| baoyu-infographic | Creative | Produces infographics using reusable layouts and styles |
| claude-design | Creative | Designs standalone HTML artifacts |
| comfyui | Creative | Generates images, video, and audio with ComfyUI |
| design-md | Creative | Authors and validates DESIGN.md specifications |
| excalidraw | Creative | Creates hand drawn Excalidraw diagrams |
| humanizer | Creative | Revises text to remove formulaic AI phrasing |
| manim-video | Creative | Produces mathematical and algorithmic animations |
| p5js | Creative | Creates interactive p5.js sketches and generative art |
| popular-web-designs | Creative | Applies established web interface design systems |
| pretext | Creative | Supports interactive creative browser demonstrations |
| sketch | Creative | Produces alternative HTML interface mockups |
| songwriting-and-ai-music | Creative | Supports songwriting and AI music prompting |
| touchdesigner-mcp | Creative | Controls TouchDesigner through an MCP interface |
| Skill | Category | Description |
|---|---|---|
| jupyter-live-kernel | Data science | Performs iterative analysis in a persistent Jupyter kernel |
| dogfood | General | Conducts exploratory testing of web applications |
| himalaya | Manages email through the Himalaya command line interface | |
| codebase-inspection | GitHub | Measures codebase size, languages, and composition |
| github-auth | GitHub | Configures tokens, keys, and command line authentication |
| github-code-review | GitHub | Reviews pull request diffs and inline comments |
| github-issues | GitHub | Creates and manages GitHub issues |
| github-pr-workflow | GitHub | Manages branches, commits, checks, and pull requests |
| github-repo-management | GitHub | Clones, creates, forks, and maintains repositories |
| gif-search | Media | Searches and downloads GIF content |
| heartmula | Media | Generates songs from lyrics and style tags |
| songsee | Media | Extracts and visualizes audio features |
| youtube-content | Media | Converts YouTube transcripts into written content |
| huggingface-hub | MLOps | Searches, downloads, and uploads models and datasets |
| evaluating-llms-harness | MLOps | Evaluates language models with standard benchmarks |
| weights-and-biases | MLOps | Tracks experiments, sweeps, and model artifacts |
| llama-cpp | MLOps | Runs local GGUF model inference |
| serving-llms-vllm | MLOps | Serves language models with vLLM |
| audiocraft-audio-generation | MLOps | Generates music and sound with AudioCraft |
| segment-anything-model | MLOps | Performs prompt based image segmentation |
| obsidian | Note taking | Reads, searches, creates, and edits Obsidian notes |
| Skill | Category | Description |
|---|---|---|
| airtable | Productivity | Manages Airtable records and queries |
| google-workspace | Productivity | Operates Gmail, Calendar, Drive, Docs, and Sheets |
| maps | Productivity | Provides geocoding, points of interest, routes, and time zones |
| nano-pdf | Productivity | Edits PDF text and document metadata |
| notion | Productivity | Manages Notion pages and databases |
| ocr-and-documents | Productivity | Extracts text from PDFs and scanned documents |
| petdex | Productivity | Installs and selects animated Hermes mascots |
| powerpoint | Productivity | Creates and edits presentation decks |
| teams-meeting-pipeline | Productivity | Operates the Teams meeting summary pipeline |
| arxiv | Research | Searches arXiv by topic, author, category, or identifier |
| blogwatcher | Research | Monitors blogs and syndicated feeds |
| llm-wiki | Research | Builds and queries an interlinked knowledge base |
| polymarket | Research | Queries prediction markets, prices, and order books |
| research-paper-writing | Research | Supports machine learning paper development and submission |
| openhue | Smart home | Controls Philips Hue lights, rooms, and scenes |
| xurl | Social media | Reads and operates X through its command line interface |
| hermes-agent-skill-authoring | Software development | Authors and validates Hermes skill packages |
| node-inspect-debugger | Software development | Debugs Node.js through the inspector protocol |
| plan | Software development | Produces actionable implementation plans |
| python-debugpy | Software development | Debugs Python with pdb and debugpy |
| requesting-code-review | Software development | Performs structured review before integration |
| simplify-code | Software development | Refines recent code changes with parallel review |
| spike | Software development | Runs disposable experiments before implementation |
| systematic-debugging | Software development | Applies a structured root cause debugging process |
| test-driven-development | Software development | Applies test driven development workflows |
| yuanbao | General | Operates Yuanbao groups and member queries |
| Product field | Access | Meaning |
|---|---|---|
| product_id | Visible | Stable identifier for a Product in the Product Catalog |
| name | Visible | Marketplace product title used for retrieval and comparison |
| category | Visible | One of the ten normalized first level product categories |
| quantity | Visible | Current effective supplier inventory after replenishment |
| price | Visible | Current procurement price offered by the supplier |
| historical_avg_rating | Visible | Historical product rating obtained from the source platform |
| logistics_hours | Visible | Baseline transit time from supplier dispatch to delivery |
| is_listed_by_supplier | Visible | Current procurement availability, exposed as supplier_available |
| ref_price | Hidden | Reference price used in the price response term of the demand model |
| base_price | Hidden | Supplier price restored after a temporary Price Change ends |
| cancel_rate | Hidden | Product level probability used to sample Cancellation |
| refund_rate | Hidden | Product level probability used to sample Return and Refund |
| only_refund_rate | Hidden | Product level probability used to sample Returnless Refund |
| bad_review_rate | Hidden | Product level probability used to sample Bad Review |
| max_quantity | Hidden | Inventory capacity used by the supplier replenishment process |
| hourly_increment | Hidden | Hourly supplier inventory replenishment amount |
| elasticity | Hidden | Product specific price elasticity used by the demand model |
| market_curve | Hidden | Real-world product level demand history over 365 days |
| quantity_updated_t | Hidden | Internal timestamp used for lazy inventory replenishment |
| price_recover_t | Hidden | Prescheduled end time of an active Price Change |
| delist_recover_t | Hidden | Prescheduled end time of an active Product Delisting |
| Supplier field | Access | Meaning |
|---|---|---|
| supplier_id | Visible | Stable supplier identifier |
| supplier_name | Visible | Public supplier name |
| shop_rating | Visible | Public supplier rating shared by all Products from the supplier |
| return_buyer_rate | Visible | Public repeat buyer rate returned by the supplier profile |
| supplier_age_years | Visible | Public supplier tenure in years |
| product_count | Visible | Number of currently available Products from the supplier |
| supplier_ship_hours | Visible | Current dispatch time for a Product from this supplier |
| base_ship_hours | Hidden | Dispatch time restored after a Shipment Delay ends |
| timeout_rate | Hidden | Product level hazard for Shipment Delay |
| price_change_rate | Hidden | Product level hazard for Price Change |
| supplier_delist_rate | Hidden | Product level hazard for Product Delisting |
| timeout_active | Hidden | Internal indicator of an active Shipment Delay |
| timeout_recover_t | Hidden | Prescheduled end time of an active Shipment Delay |
| Order field | Access | Meaning |
|---|---|---|
| order_id | Visible | Stable identifier for an individual customer order |
| product_id, product_name | Visible | Product identity associated with the order |
| supplier_id, supplier_name | Visible | Supplier identity associated with the order |
| order_time | Visible | Calendar and simulation time at which the order was placed |
| current_status | Visible | Latest realized lifecycle state |
| status_age_hours | Visible | Elapsed time since the latest realized status transition |
| expected_delivery_time | Visible | Current delivery estimate computed from realized timing information |
| delivered_time | Visible | Merchant facing delivery timestamp populated after delivery |
| late_time | Visible | Merchant facing timestamp populated only after Late Shipment is realized |
| sale_price | Visible | Merchant selling price recorded when the order was created |
| purchase_price | Visible | Procurement price recorded when the order was created |
| supplier_ship_hours | Visible | Supplier dispatch duration recorded for the order |
| supplier_logistics_hours | Visible | Baseline post dispatch logistics duration |
| actual_logistics_hours | Visible | Realized transit duration populated after delivery |
| realized_revenue | Visible | Revenue credited from outcomes realized so far |
| realized_cost | Visible | Procurement cost realized so far |
| total_penalty | Visible | Sum of penalties already applied to the order |
| net_profit | Visible | Realized revenue minus realized cost and total penalty |
| profit_finalized | Visible | Indicator that no further profit component remains unresolved |
| status_log | Visible | Realized sequence of lifecycle states and their timestamps |
| preset_anomaly | Hidden | Presampled future outcome among normal fulfillment and four customer abnormalities |
| preset_anomaly_t | Hidden | Internal realization time of the presampled abnormal outcome |
| settlement_delay_steps | Hidden | Presampled delay from delivery to final settlement |
| purchase_t, shipped_t, delivered_t, settled_t | Hidden | Raw internal transition times, with only realized merchant facing views exposed |
研究结果
- 表现最好的LLM配置(Hermes框架下的Qwen3.7-Max)最终净资产也只达到人类参与者平均水平的27.3%。
- 在8个模型的平均水平上,Hermes框架比ReAct框架的最终净资产高53.3%、GMV高71.5%、订单量多71.2%,8个模型中有7个在Hermes下表现更好(仅Kimi K2.6例外,低4.1%)。
- 人类参与者的持续窗口率(SWR)达到100%,而LLM配置在ReAct下为10.6%到99.4%,在Hermes下为17.8%到66.1%,显示活动量随时间明显衰减。
- 部分模型在没有依据的情况下固守错误策略:一次Hermes框架下的Claude Opus 4.8运行中,智能体错误地认为减少上架商品能让流量集中,导致活跃商品从47个减少到3个;一次Hermes框架下的Kimi K2.6运行在第104天判定店铺无法挽救后,在剩余523个决策窗口中有355个未采取任何行动。
- 人类参与者随着资金增加逐步扩大采购价格区间,从前三个月的43.4至53.1元人民币扩大到后三个月的58.7至90.8元人民币,而GLM-5.2、DeepSeek-V4-Flash和Kimi K2.6的定价区间则相对停滞不变。
可应用场景
- 可作为评测框架,用于检验需要长期自主管理库存、定价、现金流的智能体在运营中是否会活动衰减或目标漂移。
- 可用于比较不同LLM与智能体框架(仅工具调用 vs 附带记忆/技能管理)在长期业务表现上的差异。
- 可作为诊断方法,检验智能体能否将延迟反馈(退货、退款、差评)追溯到此前的决策并加以修正。
局限与待验证事项
- 该环境是基于1688平台数据构建的模拟系统,与真实网店经营在监管、竞争、客户行为等方面仍存在差异。
- 结果基于每个配置三次重复运行,部分配置(如Hermes下的Qwen3.7-Max)波动很大(变异系数达55.1%),解读结果时需考虑这种不稳定性。
- 人类基线仅有三名无电商经营经验的参与者,因此'达到人类27.3%'这一数字不宜直接推广为专业卖家的表现基准。
- 论文中提到的GPT-5.6、Claude Opus 4.8、Qwen3.7系列等模型名称可能与公开已知版本不完全对应,读者应关注基准设计本身而非将其视为对当前最新模型的权威排名。
- 关于Hermes附加能力(代码执行、技能创建)为何只对部分模型有效,论文仅通过案例观察进行讨论,尚缺乏系统性的因果分析。
为什么重要
这项研究用具体数字证实了一个普遍担忧:LLM智能体在短任务上表现不错,但在长达数月的自主经营任务中容易失去方向。对于正在设计需要长期自主运行的库存管理、定价决策等自动化系统的人来说,这提前指出了失败可能出现的地方——活动衰减、过早放弃、以及证据积累后仍不修正的僵化策略。
本文术语
- 长期连贯性(Long-Term Coherence) · 在长时间跨度中持续保持有目的的行为,并随着证据积累调整决策的能力
- 部分可观测马尔可夫决策过程(POMDP) · 智能体只能看到部分真实状态、必须基于不完整信息做决策的建模框架
- 持续窗口率(Sustained Window Rate, SWR) · 在任意连续30天的滚动周期内,至少发起一次环境工具调用的决策窗口所占的最低比例
- Hermes / ReAct · ReAct是仅用26种工具的最简控制器,Hermes在此基础上加入了代码执行、记忆和技能管理等功能
- 上游供应商事件 / 下游订单结果 · 供应商端很快显现的事件(涨价、下架等)与订单端延迟显现的结果(退货、退款、差评等)
论文原文摘要(英文)
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit一款把单词意义变化拆解到语法细节的开源分析工具
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)用AI总结股市新闻发现:简单的摘要方法反而比时髦的检索增强技术更靠谱
METAL LAB 最新报道
图片来源: Qiming Shi et al., arXiv:2607.28956, arxiv-nonexclusive