Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory Arbitration
arXiv:2608.197012026-08-21
여러 AI 에이전트가 남긴 기억들이 사실은 같은 출처에서 복제된 것일 때, AI가 '가짜 다수결'에 속지 않게 만드는 법
여러 AI 에이전트가 함께 일하며 공유 기억 저장소에 기록을 남기는 시스템에서는, 서로 다른 에이전트가 쓴 기억이라도 사실은 같은 원본에서 나온 복사본이라 답을 낼 때 그 복사본들이 마치 여러 개의 독립된 증거인 것처럼 잘못 세어지는 문제가 있다. 연구팀은 이를 '기억 상관관계 편향'이라 이름 붙이고, 기억들 사이의 숨은 연결을 찾아내 진짜 독립된 증거의 개수를 계산한 뒤, 부족하면 추가로 증거를 찾아오는 CAMA라는 방법을 제안했다. 여러 벤치마크 실험에서 CAMA는 기존 최고 성능 방법들보다 더 정확하게 판단했고, 가짜 다수결에 덜 흔들렸다.
무엇을 했나
문제 제기: 여러 에이전트가 만든 기억을 단순히 투표나 가중치로 합치면, 같은 원본에서 파생된 중복 기억들이 마치 독립된 증거처럼 여러 번 세어져 잘못된 결론이 다수인 것처럼 보이는 '기억 상관관계 편향'이 생긴다는 점을 지적했다
방법: CAMA는 검색된 기억들을 신경망으로 분석해 서로 얼마나 겹치는 근거를 담고 있는지 추정하고, 기억이 어디서 유래했는지 기록한 출처 정보(프로버넌스)를 함께 활용해 실제로 독립된 증거가 몇 개인지 계산한다. 증거가 부족하면 추가 검색을 하거나 기억의 원본 출처를 추적하는 정책을 학습해 스스로 증거를 보충한다
결과: MemoryAgentBench, LongMemEval, LoCoMo 세 개 벤치마크와 DeepSeek-V4-Flash, Qwen3.6-27B 두 개 모델에서 실험한 결과, CAMA는 단순 검색 기반 방법, 다수결, Mem0, HippoRAG, 멀티에이전트 방법(MAD, MADAM-RAG) 등 기존 방법들보다 일관되게 더 나은 성능을 보였고, 소수 의견이 맞는 상황에서도 정답을 더 잘 찾아내고 중복 기억이 늘어나도 결정이 덜 흔들리는 것으로 나타났다
효율성: 멀티에이전트 방법들에 비해 더 적은 LLM 호출과 토큰으로도 더 높은 정확도 향상을 보여, 반복적인 에이전트 대화 없이 구조적인 증거 분석만으로 성능을 끌어올렸다
Figure 1: An overview of our proposed CAMA. The diagram illustrates the overall workflow of memory arbitration, where retrieved memories are progressively processed through evidence decoupling, conflict arbitration, and evidence recovery.
Table 1: Overall performance comparison on three benchmark datasets. Best results are marked by bold.
Backbone
Methods
MemoryAgentBench
LongMemEval
LOCOMO
FC-SH
FC-MH
Overall
EM
F1
BERT
Judge
EM
F1
BERT
Judge
DeepSeek-V4-Flash
Vanilla RAG
68.4
34.2
51.3
34.1
45.7
84.2
51.6
29.7
40.3
83.5
47.2
Majority Voting
70.1
33.5
51.8
33.6
45.1
84.0
50.9
28.9
39.6
83.2
46.1
HippoRAG
72.6
42.8
57.7
38.5
49.8
85.3
56.4
33.4
44.1
84.6
51.8
Mem0
73.9
44.5
59.2
40.2
51.6
85.7
58.1
35.8
46.7
85.2
54.3
MAD
74.7
46.9
60.8
41.7
52.9
86.1
60.5
36.9
47.8
85.5
56.2
MADAM-RAG
75.2
48.8
63.4
44.1
54.3
86.8
62.4
37.6
49.5
85.7
59.4
CAMA (Ours)
78.9
55.7
67.3
49.8
59.1
87.9
69.2
43.6
53.8
87.1
64.7
Qwen3.6-27B
Vanilla RAG
65.2
31.4
48.3
31.5
42.9
83.4
48.7
27.3
37.8
82.7
44.5
Majority Voting
66.8
30.7
48.8
30.9
42.3
83.1
47.9
26.5
37.1
82.4
43.6
HippoRAG
69.5
39.6
54.6
35.6
46.8
84.5
53.2
30.9
41.4
83.8
49.1
Mem0
71.0
41.3
56.2
37.4
48.7
84.9
55.3
33.2
43.9
84.4
51.7
MAD
71.9
43.8
57.9
38.9
50.1
85.3
57.6
34.5
45.2
84.7
53.8
MADAM-RAG
73.6
46.7
60.2
41.6
52.6
85.9
61.2
36.8
47.5
85.3
57.1
CAMA (Ours)
76.5
53.2
64.9
47.4
56.9
87.2
67.0
41.5
51.6
86.5
62.4
Figure 2: Hyperparameter sensitivity analysis on MemoryAgentBench with DeepSeek-V4-Flash.
Table 2: Evaluation of memory correlation bias mitigation under DeepSeek-V4-Flash.
Methods
MemoryAgentBench
LongMemEval
LOCOMO
CMR ↑
RS ↓
IEG ↑
ERR ↑
CMR ↑
RS ↓
IEG ↑
ERR ↑
CMR ↑
RS ↓
IEG ↑
ERR ↑
Vanilla RAG
38.7
41.2
5.8
5.3
36.9
43.5
5.1
4.6
33.4
45.8
4.4
3.9
Majority Voting
33.5
44.8
4.7
4.9
31.8
46.9
4.2
4.2
28.7
49.3
3.5
3.5
HippoRAG
46.8
27.4
11.6
9.8
44.2
29.1
10.5
8.9
40.5
31.6
9.1
7.7
Mem0
49.6
24.1
13.7
11.2
47.1
25.8
12.6
10.3
43.4
28.2
11.0
9.0
MAD
54.3
19.2
15.9
12.6
51.8
20.7
14.5
11.4
47.6
22.9
12.8
10.1
MADAM-RAG
60.7
15.3
19.4
14.1
58.2
16.6
18.1
12.9
53.9
18.5
16.2
11.5
CAMA (Ours)
71.2
7.8
25.1
36.2
67.4
9.1
22.6
33.1
62.1
10.3
20.2
29.4
(b)
Table 3: Ablation study on the MemoryAgentBench and LongMemEval benchmarks under DeepSeek-V4-Flash.
Methods
MemoryAgentBench
LongMemEval
FC-SH
FC-MH
Overall
CMR
RS
IEG
ERR
EM
F1
BERT
Judge
CMR
RS
IEG
ERR
w/o Evi. Decoupling
71.5
45.3
58.4
48.2
32.7
14.6
28.9
42.1
51.8
86.2
59.7
46.5
34.1
13.8
26.4
w/o Prov. Prior
76.8
51.9
64.4
63.4
13.8
21.7
34.2
47.3
56.6
87.4
65.8
61.2
14.6
20.3
31.8
w/o Expand
76.2
51.4
63.8
66.7
9.5
17.2
15.3
46.1
55.7
87.3
64.9
63.8
10.2
15.8
13.6
w/o Trace
75.4
49.6
62.5
64.8
15.2
19.4
27.1
46.8
56.2
87.4
65.7
62.5
16.3
18.1
25.2
w/o Policy
77.1
52.3
64.7
65.9
11.4
16.8
22.7
47.6
57.0
87.5
66.4
64.1
12.5
15.2
20.8
CAMA
78.9
55.7
67.3
71.2
7.8
25.1
36.2
49.8
59.1
87.9
69.2
67.4
9.1
22.6
33.1
(c)
Table 4: Efficiency analysis on the MemoryAgentBench benchmark under DeepSeek-V4-Flash.
Methods
Avg. Latency
LLM Calls
Token Cost (k)
Context Len (k)
Δ Acc./ kToken
Vanilla RAG
1.8
1.0
3.2
3.1
–
Majority Voting
4.6
5.0
12.8
3.1
0.04
HippoRAG
3.2
2.0
5.9
4.4
1.08
MAD
9.8
8.4
24.3
9.6
0.39
MADAM-RAG
11.4
10.6
28.9
11.2
0.42
CAMA(Ours)
6.7
4.2
14.6
6.8
1.10
(d)
Table 5: Evaluation of memory correlation bias mitigation under Qwen3.6-27B.
Methods
MemoryAgentBench
LongMemEval
LOCOMO
CMR ↑
RS ↓
IEG ↑
ERR ↑
CMR ↑
RS ↓
IEG ↑
ERR ↑
CMR ↑
RS ↓
IEG ↑
ERR ↑
Vanilla RAG
36.2
43.1
5.1
4.7
34.5
45.4
4.6
4.1
31.1
47.5
4.0
3.5
Majority Voting
31.4
46.5
4.1
4.3
29.8
48.6
3.7
3.8
26.8
51.0
3.1
3.1
HippoRAG
43.9
29.1
10.4
8.7
41.5
30.8
9.5
7.9
38.0
33.4
8.2
6.9
Mem0
46.7
25.7
12.3
10.1
44.4
27.4
11.4
9.3
40.8
29.8
9.9
8.2
MAD
51.2
20.6
14.5
11.5
48.9
22.1
13.2
10.4
44.8
24.4
11.7
9.2
MADAM-RAG
57.5
16.7
17.7
13.0
55.0
18.1
16.5
11.8
50.8
20.1
14.8
10.5
CAMA (Ours)
68.1
8.6
23.4
34.0
64.3
10.0
21.1
31.0
59.0
11.2
18.8
27.5
Table 6: Ablation study on the MemoryAgentBench and LongMemEval benchmarks under Qwen3.6-27B.
Methods
MemoryAgentBench
LongMemEval
FC-SH
FC-MH
Overall
CMR
RS
IEG
ERR
EM
F1
BERT
Judge
CMR
RS
IEG
ERR
w/o Evi. Decoupling
69.3
43.5
56.4
46.0
34.1
13.2
27.0
40.3
50.1
85.6
57.8
44.7
35.4
12.5
24.8
w/o Prov. Prior
74.6
49.8
62.2
60.8
14.9
20.2
32.1
45.7
55.0
86.8
63.9
58.6
15.7
18.9
29.8
w/o Expand
74.1
49.3
61.7
63.8
10.6
15.9
14.2
44.6
54.1
86.7
63.0
60.9
11.3
14.7
12.5
w/o Trace
73.4
47.6
60.5
61.9
16.2
18.0
25.4
45.2
54.7
86.8
63.8
59.7
17.4
17.0
23.7
w/o Policy
75.0
50.5
62.8
63.1
12.4
15.6
21.2
46.1
55.6
86.9
64.5
61.4
13.6
14.2
19.6
CAMA
76.5
53.2
64.9
68.1
8.6
23.4
34.0
47.4
56.9
87.2
67.0
64.3
10.0
21.1
31.0
Table 7: Efficiency analysis on the MemoryAgentBench benchmark under Qwen3.6-27B.
Methods
Avg. Latency
LLM Calls
Token Cost (k)
Context Len (k)
Δ Acc./ kToken
Vanilla RAG
1.5
1.0
3.2
3.1
–
Majority Voting
3.9
5.0
12.8
3.1
0.04
HippoRAG
2.8
2.0
5.9
4.4
1.07
MAD
8.3
8.4
24.3
9.6
0.40
MADAM-RAG
9.7
10.6
28.9
11.2
0.41
CAMA(Ours)
5.8
4.2
14.6
6.8
1.14
왜 중요한가
여러 AI 에이전트가 협업하며 기억을 계속 쌓아가는 시스템에서, 같은 오류가 복제되어 쌓이면 시간이 지날수록 틀린 결론이 굳어지는 위험이 있다. 이 연구는 그런 자기강화식 오류를 방지하는 실질적인 방법을 제시해, 장기간 운영되는 멀티에이전트 서비스의 신뢰성을 높이는 데 참고할 수 있다.
이 논문의 용어
기억 상관관계 편향(Memory Correlation Bias) · 서로 다른 에이전트가 남긴 기억이 사실은 같은 출처에서 나온 것이어서, 독립된 증거처럼 중복 계산되어 잘못된 다수 의견을 만드는 현상
프로버넌스(provenance) · 어떤 기억이 어디서, 어떤 과정을 거쳐 만들어졌는지 기록한 출처 정보
CAMA(Correlation-Aware Memory Arbitration) · 기억들 사이의 숨은 연관성을 찾아 독립 증거 개수를 추정하고, 부족하면 추가로 증거를 찾아 최종 결론을 내리는 이 연구의 제안 방법
실효 독립 증거 수(Neff) · 검색된 기억들이 실제로 몇 개의 서로 다른 독립된 근거를 담고 있는지를 계산한 값
회수 정책(recovery policy) · 증거가 부족할 때 추가 검색을 할지, 기억의 원본 출처를 추적할지, 아니면 멈출지를 스스로 판단하도록 학습된 정책
논문 원문 초록 (영문)
Long-term multi-agent systems continuously accumulate the memories produced by different agents. Existing memory methods typically treat retrieved memories as independent evidence and combine them through voting or weighting. However, this independence assumption often fails in multi-agent settings: memories written by different agents may inherit the same upstream source or shared bias, causing correlated evidence to be repeatedly counted and creating a false majority. We term this failure mode \textit{Memory Correlation Bias}. To address the issue, we propose the \textbf{C}orrelation-\textbf{A}ware \textbf{M}emory \textbf{A}rbitration (CAMA) framework that jointly decouples retrieved memories and recovers missing independent evidence. We model the retrieved memories as query-conditioned evidence groups and combine neural dependency inference with provenance-based symbolic priors to estimate the effective number of independent evidence sources, thereby preventing correlated memories from forming a false majority. Since critical independent evidence may be absent from the initial retrieval set, \textsc{CAMA} further learns a sequential recovery policy that actively retrieves alternative evidence or traces upstream sources before making the final decision, aiming to recover sufficient independent evidence for reliable arbitration while minimizing retrieval cost. Experiments on multiple benchmarks demonstrate the superiority of our method over the state-of-the-art baseline methods, suppressing false majorities induced by correlated memories.