매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents

arXiv:2608.196622026-08-21

AI 에이전트가 도구 설명서를 매번 다시 읽지 않게 만들어 속도를 3배 이상 높이는 방법

AI 에이전트는 요청마다 도구나 기능 설명(스키마)을 다시 조합해 처음부터 다시 읽어야 해서 기존의 캐시 재사용 기법이 잘 통하지 않았다. ReCache는 각 도구 설명을 서로 독립적으로 미리 계산해두고, 예측에 중요한 부분만 골라 저장하는 방식으로 이 문제를 해결한다. 그 결과 응답 속도는 최대 3.655배 빨라지고 메모리 사용량은 92.43% 줄었으면서도 정확도는 기존 방식과 거의 같았다.

무엇을 했나

  1. 문제: 에이전트가 쓰는 도구·기능 설명(리소스)은 매 요청마다 다른 조합·순서로 등장해서, 문장 앞부분이 완전히 같아야 재사용되는 기존 '프리픽스 캐싱'이 통하지 않았다
  2. 해결책 1: '리소스별 어텐션'을 도입해 서로 다른 도구 설명끼리는 서로 참조하지 않게 하고, 각 설명 내부 위치 번호를 0부터 새로 매겨서 어떤 조합·순서로 와도 똑같이 계산되는 독립적인 캐시 블록을 만들었다
  3. 해결책 2: 모델의 어떤 층(layer)과 어떤 헤드 그룹이 실제로 도구 정보를 활용하는지 측정해서, 기여도가 높은 통로만 남기고(구조적 가지치기) 나머지는 아예 도구 정보를 못 보게 막았다
  4. 해결책 3: 도구 이름, 인자 이름, 인자 설명, 마지막 요약 토큰 등 호출에 꼭 필요한 부분만 남기고 나머지 텍스트는 지우는 '의미적 가지치기'를 추가했다
  5. 결과: 7개의 공개 도구·기능 사용 데이터셋으로 만든 벤치마크에서, 리소스별 어텐션만으로 정확도는 밀집 어텐션 대비 82.3% vs 82.4%로 거의 동일했고 첫 토큰 생성 속도는 3.655배 빨라졌으며, 전체 프레임워크는 캐시 메모리를 92.43% 줄이고 어텐션 계산을 1.423배 빠르게 했다
Figure 1: Cross-resource attention on Qwen3-4B.
Figure 1: Cross-resource attention on Qwen3-4B.
Table 1: Selected structural budgets. CC and ρ are percentages; ΔInv-F1 is measured relative to Ωfull after full-data training.
ModelKL/LKG/GCCLCCGρΔInv-F1 (%)
Ql20/363/897.747.379.2−0.2
Qs20/287/899.899.337.5−0.7
(b) Attention from subsequent conversation tokens onto resources, normalized and aligned by relative position.
(b) Attention from subsequent conversation tokens onto resources, normalized and aligned by relative position.
Table 2: Resource-wise attention ablation.
MethodInv-F1ID-F1Halluc.TTFT (ms)
Dense82.496.00.026.319
Block82.295.80.1
Ωfull82.396.00.27.200
Figure 2: ReCache constructs reusable KV blocks for each Ri and progressively reduces their subsequent cache retention and access. (a) Resource-wise attention removes inter-resource attention and resets positions from zero within each Ri. (b) Structural pruning exposes resource KV states (using Mresource) only through selected layer (Li)–KV-head-group (Gi) routes Ω⋆ (highlighted in blue). The remaining routes use Mcontext. (c) Semantic pruning retains resource names, argument (arg) names and descriptions (desc.), and the final suffix token. The final mask combines these decisions, ensuring that Y attends exclusively to retained semantic fields within assigned routes during decoding.
Figure 2: ReCache constructs reusable KV blocks for each Ri and progressively reduces their subsequent cache retention and access. (a) Resource-wise attention removes inter-resource attention and resets positions from zero within each Ri. (b) Structural pruning exposes resource KV states (using Mresource) only through selected layer (Li)–KV-head-group (Gi) routes Ω⋆ (highlighted in blue). The remaining routes use Mcontext. (c) Semantic pruning retains resource names, argument (arg) names and descriptions (desc.), and the final suffix token. The final mask combines these decisions, ensuring that Y attends exclusively to retained semantic fields within assigned routes during decoding.
Table 3: Results on 𝒯IND. Efficiency is reported relative to Dense. SMP denotes Semantic Pruning.
MethodAttn. ×Mem. ↓Inv-F1ID-F1Halluc.
Dense82.496.00.0
Ωfull1.001×0.47%82.396.00.2
Structural pruning
Ω20,G1.016×44.71%82.595.70.3
Ωfull + SPEED1.013×44.71%79.392.92.8
Ω20,31.314×79.27%82.195.60.2
Ωfull + SA20,31.302×79.27%79.193.71.0
Semantic pruning
Ωfull + Gist1.029×99.22%39.246.951.4
Ωfull + Beacon1.020×75.42%78.692.51.4
Ωfull + SMP1.021×63.57%81.695.20.2
ReCache1.423×92.43%80.394.90.2
Figure 3: Structural contribution coverage in Ql and Qs. Cumulative leave-one-in contributions of layers (top) and KV head groups (bottom); vertical lines mark the selected budgets.
Figure 3: Structural contribution coverage in Ql and Qs. Cumulative leave-one-in contributions of layers (top) and KV head groups (bottom); vertical lines mark the selected budgets.
Table 4: 𝒯OOD Results. SMP denotes Semantic Pruning.
MethodInv-F1ID-F1Halluc.
Dense66.395.60.0
Ωfull64.794.40.4
Structural pruning
Ω20,G64.394.30.5
Ωfull + SPEED54.381.613.4
Ω20,363.293.30.5
Ωfull + SA20,358.288.45.1
Semantic pruning
Ωfull + Gist9.717.879.9
Ωfull + Beacon58.589.74.2
Ωfull + SMP62.894.10.3
ReCache60.892.80.6
Figure 4: Efficiency on varying resource lengths Dℛ. Requests are grouped by average resource length into Small (0≤Dℛ<1​K), Medium (1​K≤Dℛ<5​K), Large (5​K≤Dℛ<10​K), and XL (Dℛ≥10​K). SMP denotes semantic pruning.
Figure 4: Efficiency on varying resource lengths Dℛ. Requests are grouped by average resource length into Small (0≤Dℛ<1​K), Medium (1​K≤Dℛ<5​K), Large (5​K≤Dℛ<10​K), and XL (Dℛ≥10​K). SMP denotes semantic pruning.
Table 5: Dataset composition and split statistics. Source rows report eligible examples before final selection; the Toucan aggregate is not double-counted in the summary row. Distinct names counts source-level unique declared tools or skills, Subset dup. counts records beyond the first occurrence of an identical candidate configuration, and Recurring names counts resources appearing in multiple records. The summary row corresponds to 18.8% duplicated candidate records and 77.3% recurring resource names. IND cand. and OOD cand. denote eligible in-distribution and out-of-distribution candidates, and dashes mark unavailable statistics.
GroupSourceTrain poolDistinct namesSubset dup.Recurring namesIND cand.OOD cand.
ToolToolACE8,40014,5661787,144200195
APIGEN10,0002,8067342,805500146
Toucan (all generators)13,0714,3556,9014,355600493
- Kimi-K23,3331,1031,7721,103200145
- GPT-OSS-120B4,1841,4122,1621,412200130
- Qwen3-32B5,5541,8402,9671,840200218
ToolRet1,1501,593235431100150
ToolMind20,00014,5041,85314,497500190
WildToolBench256
SkillSkillRouter87
Sources with frequency statistics52,62137,8249,90129,2321,9001,174
Candidate pool52,9641,9001,174
Final dataset49,4241,0001,000
(b) TPOT
(b) TPOT
Table 6: Structural budget ablation for Qs across training data regimes with comparable budgets.
ConfigInv-F1Resource IDHalluc.
ID-PID-RID-F1
Ωfull72.192.090.890.90.4
ΩL,549.674.573.573.44.8
ΩL,653.479.677.978.33.2
ΩL,757.383.481.882.12.9
Ω15,G69.389.288.388.31.1
Ω20,G70.590.689.589.61.1
(c) Attn.
(c) Attn.
Table 7: Structural budget ablation trained on 𝒟.
ConfigInv-F1Resource IDHalluc.
ID-PID-RID-F1
𝑸𝒔
Ωfull80.794.994.394.40.1
Ω20,578.894.193.593.60.5
Ω20,780.094.794.094.10.3
𝑸𝒍
Ωfull82.396.495.896.00.2
Ω20,281.995.394.794.70.3
Ω20,382.196.195.595.60.2
(d) Mem.
(d) Mem.
Table 8: Semantic pruning ablation trained on 𝒟subset. Arg stands for argument, Desc. stands for description. Ω20,3 denotes the routing configuration with KL=20,KG=3.
ConfigInv-F1Resource IDHalluc.
ID-PID-RID-F1
Ω20,376.193.192.892.50.9
Schema representation ablation
Suffix-only Schema14.822.823.022.976.0
+ Resource Name45.289.988.888.91.8
+ Arg Name67.792.191.191.21.3
+ Resource Desc.67.992.291.491.41.1
+ Arg Desc.72.892.891.992.01.0
+ Both Desc.73.493.092.292.21.2
Table 9: Individual contribution weights and cumulative contribution coverage of structural units. Within each panel, units are ranked independently by decreasing wi, and CC at rank k is the cumulative positive contribution weight through rank k. Bold entries fall within the empirical budgets KL=20,KG=3 for Ql and KL=20,KG=7 for Qs.
Layers
𝑸𝒍(Qwen3-4B)𝑸𝒔(Qwen3-1.7B)
rankunit𝒘𝒊 (%)CC (%)rankunit𝒘𝒊 (%)CC (%)
1L3415.5615.61L2220.0920.1
2L3510.5326.12L2512.9033.0
3L239.3735.53L2012.1445.1
4L258.3143.84L2610.9056.0
5L328.0151.85L2710.3966.4
6L306.9558.76L176.8273.2
7L336.8965.67L216.7079.9
8L275.1770.88L194.6584.6
9L214.8775.79L244.0088.6
10L284.6480.310L232.1390.7
11L244.0284.311L181.8992.6
12L223.7488.112L151.7394.3
13L172.3090.413L131.3995.7
14L311.9692.314L101.1396.9
15L181.2993.615L91.0397.9
16L141.2694.916L40.8898.8
17L200.7895.717L160.5699.3
18L360.6996.318L120.1899.5
19L190.6897.019L110.1599.7
20L160.6897.720L80.1499.8
21L100.6098.321L280.1199.9
22L50.3398.622L50.06100.0
23L110.2998.923L20.02100.0
24L90.2799.224L10.00100.0
25L120.2699.525L30.00100.0
26L150.2699.726L60.00100.0
27L130.1399.827L70.00100.0
28L80.0799.928L140.00100.0

왜 중요한가

도구를 호출하는 AI 에이전트는 매번 방대한 도구 설명서를 다시 읽어야 해서 응답이 느려지고 서버 비용이 커지는데, ReCache는 이 반복 계산을 없애 실서비스에서 에이전트를 더 싸고 빠르게 돌릴 수 있는 실질적인 방법을 제시한다. 코드가 공개되어 있어 도구 사용형 AI 시스템을 만드는 개발자가 바로 참고할 수 있다.

이 논문의 용어

  • KV 캐시 · 언어모델이 이전에 처리한 토큰들의 계산 결과(키·값)를 저장해두어 다음 계산에 재사용하는 저장 공간
  • 프리픽스 캐싱 · 입력 문장의 앞부분이 완전히 동일할 때만 이전 계산 결과를 재사용하는 기존 캐시 방식
  • TTFT (첫 토큰까지 걸리는 시간) · 요청을 받고 모델이 첫 번째 답변 토큰을 내놓기까지 걸리는 시간, 응답 체감 속도를 좌우함
  • 구조적 가지치기 / 의미적 가지치기 · 각각 모델의 층·헤드 단위, 텍스트의 토큰 단위에서 덜 중요한 부분을 골라 없애는 압축 기법
  • 리소스(도구·기능 스키마) · 에이전트가 호출할 수 있는 도구나 기능을 설명하는 텍스트 정의(이름, 인자, 설명 등)

논문 원문 초록 (영문)

Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReCache}, a framework for independently caching resource representations while reducing their inference-time computational and memory overhead. Resource-wise attention removes cross-resource interactions and assigns resource-local positions, producing composition-invariant KV blocks. ReCache then restricts resource visibility to contribution-selected layer--KV-head-group routes and retains only invocation-critical fields through structural and semantic pruning. We evaluate ReCache on a benchmark assembled from seven public tool- and skill-use datasets, including resource-disjoint tests. Resource-wise attention matches dense invocation performance (82.3\% versus 82.4\% Inv-F1) while providing a 3.655$\times$ time-to-first-token speedup. The complete framework reduces allocated KV-tensor memory by 92.43\% and accelerates attention by 1.423$\times$. These results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss. The code is available at https://github.com/EIT-NLP/ReCache.

저자 · Yichu Fang, Sitong Wei, Haozhe Hu, Xiaoyu Shen

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Yichu Fang et al., arXiv:2608.19662, arxiv-nonexclusive