매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

arXiv:2608.198612026-08-21

고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법

고객서비스 LLM 에이전트는 위험한 행동 하나만 막는 게 아니라 신원확인, 순서, 확인 절차 같은 여러 단계를 빠짐없이 지켜야 한다. PolicyGuide는 회사 정책 문서를 그래프(순서도) 형태로 미리 변환해두고, 대화가 진행될 때마다 별도의 검증 AI가 지금까지의 대화를 그래프와 대조해 다음에 해야 할 일을 알려준다. GPT-5.4 에이전트 기준으로 항공, 소매, 통신 세 도메인에서 평균 성공률(4번 중 4번 다 통과하는 비율)을 0.42에서 0.62로 끌어올렸고, 절차가 가장 복잡한 통신 도메인에서는 0.19에서 0.61로 가장 크게 개선됐다.

무엇을 했나

  1. 문제 정의: 규정 위반은 '금지된 행동을 함'뿐 아니라 '신원확인이나 최종 확인 같은 절차를 빠뜨림'으로도 발생하는데, 기존 안전장치는 위험한 행동 하나가 발생하는 순간에만 개입해서 그 이전 단계의 누락은 잡아내지 못한다.
  2. 방법: PolicyGuide는 각 도메인의 정책 문서를 미리 오프라인으로 분석해 '누가 무엇을 해야 하는지'를 노드와 화살표로 표현한 워크플로우 그래프로 만들어두고, 대화가 매 턴 넘어갈 때마다 검증 전용 LLM이 지금까지의 대화 기록과 그래프상의 진행 위치를 대조해 아직 만족되지 않은 첫 단계를 찾아 구체적인 개선 지시(remediation)를 에이전트에게 전달한다.
  3. 이 위치 정보는 대화 시스템의 기억이 아니라 별도 코드가 계속 저장하고 있어서, 여러 개의 진행 중인 요청을 동시에 놓치지 않고 추적할 수 있다.
  4. 결과: τ2-bench의 항공/소매/통신 벤치마크에서 GPT-5.4 기준 평균 Pass4(4번 시도 모두 성공하는 비율)를 0.42에서 0.62로 올렸고, 특히 절차가 가장 정교한 통신 도메인에서 0.19에서 0.61로 가장 크게 개선됐다. 같은 워크플로우를 Claude Sonnet 4.6, Gemini 2.5 Pro 에이전트에 그대로 적용해도 효과가 이어졌고, 사용자가 거짓 정보로 속이려는 적대적 상황에서도 공격 성공률이 가장 낮았다.
Table 1: Main results on the base splits (GPT 5.4 agent, n=4; airline 50, retail/telecom 114 tasks). Cells report Pass1 and Pass4 overall and on the PV/Mut slices. The verifier is absent for ReAct, static code for ToolGuard, and GPT 5.4 for PolicyGuard and PolicyGuide.
Airline (50)Retail (114)Telecom (114)
SystemVerifierOverallPVMutOverallPVMutOverallPVMut
Pass1ReAct0.6400.8650.4330.8000.9000.7910.3840.7210.180
ToolGuardstatic code0.5750.9690.212
PolicyGuardGPT 5.40.7101.0000.4420.6450.9750.6130.4060.7330.208
PolicyGuideGPT 5.40.7750.9790.5870.8090.9750.7930.8660.8950.849
Pass4ReAct0.4600.7500.1920.5960.7000.5870.1930.4420.042
ToolGuardstatic code0.5200.8750.192
PolicyGuardGPT 5.40.5801.0000.1920.3600.9000.3080.2020.4880.028
PolicyGuideGPT 5.40.6200.9170.3460.6140.9000.5870.6140.7210.549
Table 2: Workflow ablations (GPT 5.4 agent; Airline base split, Retail and Telecom benchmark test splits of 40 tasks). All cells report Pass4.
DomainMetricReActPolicyGuide SelfPolicyGuide RawPolicyGuide
AirlineOverall0.4600.4800.5200.620
PV0.7500.8330.8750.917
Mut0.1920.1540.1920.346
RetailOverall0.5750.3500.5750.725
PV0.7500.7500.7501.000
Mut0.5560.3060.5560.694
TelecomOverall0.2500.3250.3500.675
PV0.4290.5710.6190.667
Mut0.0530.0530.0530.684
Table 3: Matched workflow-controller comparison on the 40-task Telecom benchmark test split. All values are Pass4.
SystemRuntime controlPass4
ReActactor only0.250
PolicyGuardaction-local check0.325
FlowAgentPDL + API control0.350
PolicyGuideexternal graph verifier0.675
Table 4: Agent-family generalization on Airline (50 tasks, n=4; verifier model paired to the agent). All metrics are Pass4. The GPT 5.4-authored workflow graph is reused without re-authoring.
AgentMetricReActPolicyGuardPolicyGuide
GPT 5.4Overall0.4600.5800.620
PV0.7501.0000.917
Mut0.1920.1920.346
Claude Sonnet 4.6Overall0.7200.7800.780
PV0.9581.0001.000
Mut0.5000.5770.577
Gemini 2.5 ProOverall0.4800.6000.680
PV0.7501.0000.917
Mut0.2310.2310.462
Table 5: Author-designed Telecom ordered trace compliance (%; n=4). Step- and Trace-TCR condition on outcome-passing traces.
SystemStep-TCRTrace-TCRProcess-valid rate
ReAct86.435.417.5
PolicyGuard85.723.913.1
PolicyGuide94.563.456.2
Table 6: Argument- (A), process- (P), and workflow-level (W) partition of the source policies (W splits PolicyGuard’s process-level class; P+W equals it).
DomainAPWTotal% P+W% W
Airline142724367.4%4.7%
Retail027128∼100%3.6%
Telecom (main)12172996.6%24.1%
Telecom (manual)012021100%95.2%
Telecom (both)122275098.0%54.0%
Table 7: Hand-classified atomic requirements of the τ2-bench Retail policy document (28 requirements, 0 A / 27 P / 1 W; subtypes D-only 13, T-only 14, D+T 1). Line refers to retail/policy.md as released with τ2-bench; Type A = argument-level, P = process-level (flat), W = workflow-level (order-bound; Appendix B.1), with D = dialogue-dependent, T = requires a prior read-only tool call.
IDLineRequirement (paraphrased)Type
Global rules
G110Authenticate identity by locating the user id via email or name+zip—even when the user already provides the idP (D+T)
G214One user per conversation; deny any request about another userP (D)
G316List action details + obtain explicit “yes” before any DB-updating actionP (D)
G418No fabricated information/knowledge/procedures; no subjective recommendationsP (D)
G520At most one tool call per turn (not paired with a user-facing reply)P (D)
G622Deny user requests that are against the policyP (D)
G724Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff messageW (D)
Generic action rules
N182Act only on orders with status pending or deliveredP (T)
N284Exchange / modify-items tools callable only once per orderP (T)
N384Collect all items to change into one list before making the callP (D)
Cancel pending order
C188Order status must be pending; check it before taking the actionP (T)
C290User confirms order id + reason ∈ {‘no longer needed’, ‘ordered by mistake’}; no other reasonP (D)
Modify pending order
M196Order status must be pending; check it before taking the actionP (T)
M298Only shipping address, payment method, or item options may be modified—nothing elseP (D)
M3102New payment = a single method, different from the originalP (T)
M4104If the new payment is a gift card, its balance must cover the total amountP (T)
M5110Modify-items is one-shot (order becomes unmodifiable): remind + confirm all items firstP (D)
M6112Each new item must be availableP (T)
M7112New item = same product, different option (no product-type change)P (T)
M8114User provides a payment method for the price differenceP (D)
M9114If that payment is a gift card, its balance must cover the price differenceP (T)
Return delivered order
R1118Order status must be delivered; check it before taking the actionP (T)
R2120User confirms order id + the list of items to be returnedP (D)
R3122–124Refund method provided; must be the original payment method or an existing gift cardP (T)
Exchange delivered order
Table 8: Hand-classified atomic requirements of the τ2-bench Telecom main_policy.md (29 requirements, 1 A / 21 P / 7 W). Line refers to the document as released with τ2-bench; \raisebox{-0.4pt}{\scriptsize$n$}⃝ marks a step in an ordered procedure (“To do so you need to follow these steps”). Types as in Table 7.
IDLineRequirement (paraphrased)Type
Global rules
G17No fabricated information/knowledge/procedures; no subjective recommendationsP (D)
G29At most one tool call per turn (not paired with a user-facing reply)P (D)
G311Deny user requests that are against the policyP (D)
G413Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff messageW (D)
G515Try your best to resolve the issue before transferringW (D)
Customer lookup
L194–97Identify the customer via phone number, customer ID, or full name + date of birthP (D+T)
L299For name lookup, date of birth is required for verificationP (D)
Overdue bill payment (ordered procedure)
O1105, 117\raisebox{-0.4pt}{\scriptsize1}⃝Check the bill status is Overdue before acting (the API does not check it)P (T)
O2106\raisebox{-0.4pt}{\scriptsize2}⃝Check the bill amount dueP (T)
O3107–108\raisebox{-0.4pt}{\scriptsize3}⃝Send the payment request (→ AWAITING PAYMENT); gated on O1P (T)
O4109–110\raisebox{-0.4pt}{\scriptsize4}⃝Inform the user to check their payment requestsW (D)
O5111\raisebox{-0.4pt}{\scriptsize5}⃝Only after the user accepts, call make_paymentW (D+T)
O6113\raisebox{-0.4pt}{\scriptsize6}⃝Always verify the bill became PAID before telling the userW (T)
O7116At most one bill in AWAITING PAYMENT at a timeP (T)
Line suspension
S1125Lift a suspension only after all overdue bills are paidW (T)
S2126Do not lift if the contract end date is past—even if all bills are paidP (T)
S3128After resuming, instruct the user to reboot the deviceW (D)
Data refueling (ordered procedure)
F1134Refuel amount ≤2 GBA
F2136\raisebox{-0.4pt}{\scriptsize1}⃝Ask how much data the user wants to refuelP (D)
F3137\raisebox{-0.4pt}{\scriptsize2}⃝Confirm the priceP (D)
F4138\raisebox{-0.4pt}{\scriptsize3}⃝Apply the refuel to the line associated with the user’s phone numberP (D+T)
Change plan (ordered procedure)
P1144\raisebox{-0.4pt}{\scriptsize1}⃝Establish which line the plan change is forP (D)
P2145\raisebox{-0.4pt}{\scriptsize2}⃝Gather the available plansP (T)
P3146\raisebox{-0.4pt}{\scriptsize3}⃝Ask the user to select oneP (D)
Table 9: Hand-classified atomic requirements of the τ2-bench Telecom tech_support_manual.md (21 requirements, 1 P / 20 W). Line refers to the document as released with τ2-bench. Every rule is a diagnostic-gated (T) user-guidance (D) step; the three chapters form the prerequisite chain Service ⊂ Data ⊂ MMS; every row except TSS1 (the entry diagnostic) is workflow-level.
IDLineRequirement (diagnose → conditional fix → verify)Type
Cellular service (ll. 55–99)
TSS169–72Diagnose service via check_status_barP (T)
TSS274–78If Airplane Mode ON → guide toggle_airplane_mode OFFW (D+T)
TSS379–87Check SIM: Missing → reseat; Locked → escalate; Active → ok (three-way branch)W (D+T)
TSS488–92If APN incorrect → guide reset_apn_settings, then reboot_deviceW (D+T)
TSS593–99If line suspended → handle per main policy, then verify service restoredW (T)
Mobile data (ll. 100–163)
TSD0106–108Prerequisite: the user must first have cellular serviceW (T)
TSD1122–127Diagnose via run_speed_testW (T)
TSD2129–131Airplane Mode (as in the Service chapter)W (D+T)
TSD3132–135If mobile data disabled → guide toggle_data ONW (D+T)
TSD4136–141If roaming abroad & data off → guide toggle_roaming + verify the line is roaming-enabledW (D+T)
TSD5142–145If Data Saver ON → guide toggle_data_saver_mode OFFW (D+T)
TSD6146–150If VPN ON & performance poor → guide disconnect_vpnW (D+T)
TSD7151–158If usage exceeds the plan limit → offer change-plan or refuelW (T)
TSD8159–163If network mode 2G/3G → guide set_network_mode_preferenceW (D+T)
MMS (ll. 164–205)
TSM0170–173Prerequisite: the user must have cellular service and mobile dataW (T)
TSM1181–183Diagnose via can_send_mmsW (T)
TSM2185–188Ensure basic service + data connectivity firstW (T)
TSM3189–193If on 2G → guide set_network_mode_preference to 3G+W (D+T)
TSM4194–199If MMSC URL unset → guide reset_apn_settings, then reboot_deviceW (D+T)
TSM5200–203If Wi-Fi Calling ON → guide toggle_wifi_calling OFFW (D+T)
TSM6204–205If the messaging app lacks storage/SMS permissions → guide grant_app_permissionW (D+T)
Table 10: Passk breakdown for the base-split results in Table 1 and Figure 4 (GPT 5.4, n=4). P4/P1 is the consistency ratio.
DomainSystemP1P2P3P4P4/P1
AirlineReAct0.6400.5300.4850.4600.72
ToolGuard0.5750.5530.5350.5200.90
PolicyGuard0.7100.6300.5950.5800.82
PolicyGuide0.7750.7070.6600.6200.80
RetailReAct0.8000.7000.6380.5960.75
PolicyGuard0.6450.5060.4210.3600.56
PolicyGuide0.8090.7150.6540.6140.76
TelecomReAct0.3840.2730.2260.1930.50
PolicyGuard0.4060.2920.2370.2020.50
PolicyGuide0.8660.7630.6820.6140.71
Table 11: Pass1 in each of the four trials on the base splits. pstd is the population standard deviation across trial-level values.
DomainSystemT1T2T3T4pstd
AirlineReAct0.6200.6200.6400.6800.024
ToolGuard0.5600.5800.5800.5800.009
PolicyGuard0.7000.7400.7000.7000.017
PolicyGuide0.8000.7200.8200.7600.038
RetailReAct0.8070.7460.8420.8070.035
PolicyGuard0.6490.6230.6320.6750.020
PolicyGuide0.8160.7980.7890.8330.017
TelecomReAct0.3420.3770.4040.4120.027
PolicyGuard0.4650.3860.3600.4120.039
PolicyGuide0.8600.8600.8420.9040.023
Table 12: Pooled stratified McNemar tests on per-task Pass4. D is the number of domain strata; a counts PolicyGuide-only passes and b the reverse, summed across strata.
OpponentD∑a∑bndiscZp
ReAct3771996+5.92<10−8
PolicyGuard39719116+7.24<10−12
Table 13: Per-domain paired-bootstrap differences in Pass4 on the base splits (10,000 task-level resamples). Positive values favor PolicyGuide.
DomainOpponentnΔ​P4 [95% CI]
AirlineReAct50+0.160 [+0.020, +0.300]
ToolGuard50+0.100 [−0.040, +0.260]
PolicyGuard50+0.040 [−0.080, +0.160]
RetailReAct114+0.018 [−0.070, +0.105]
PolicyGuard114+0.254 [+0.149, +0.360]
TelecomReAct114+0.421 [+0.316, +0.526]
PolicyGuard114+0.412 [+0.298, +0.526]
Table 14: Guide-side model usage for the GPT 5.4 configuration (50 Airline and 40 Retail/Telecom tasks). Costs exclude the actor and user simulator.
DomainCalls/ taskPrompt tok./callCached inputOutput tok./callGuide total $Guide $/task
Airline7.5632,36088.1%2,47820.100.40
Retail7.4222,80385.8%2,17913.670.34
Telecom11.4728,51886.5%2,18622.290.56
Table 15: Mean end-to-end wall-clock time per task.
DomainReAct (s/task)PolicyGuide (s/task)Ratio
Airline36.4210.15.78×
Retail34.6193.65.60×
Telecom45.5247.65.45×
Table 16: Programmatic validation rerun on the frozen workflow graphs.
DomainNodesAuth. nodesValidator flags
Airline158110
Retail10470
Telecom12751
Table 17: Call-NMR on passing Mut trajectories (n=4): percentage of successfully executed agent mutations missing a frozen guard-derived read prerequisite. †Telecom is an adapted, agent-side diagnostic whose read oracle saturates; its zeros do not establish equal procedural quality.
Call-NMR (%; ↓)AirlineRetailTelecom†
ReAct25.447.60.0
PolicyGuard32.534.80.0
PolicyGuide15.634.70.0

왜 중요한가

금융, 통신, 항공 같은 실제 고객상담 업무에 LLM 에이전트를 투입하려면 단순히 '틀린 행동을 막는 것'만으로는 부족하고, 신원확인·승인·확인 같은 절차 전체를 지키는지가 중요하다. 이 연구는 모델을 재학습하지 않고도 외부에서 규정 준수를 감독할 수 있는 실용적인 틀을 제시해, 실제 상담 서비스에 AI를 도입할 때 신뢰성을 높이는 방법을 보여준다.

이 논문의 용어

  • Pass4 · 같은 작업을 4번 반복 시도했을 때 4번 모두 성공하는 비율. 일관성 있게 성공하는지를 재는 지표
  • 워크플로우 그래프 · 정책 문서에 적힌 절차를 '누가 무엇을 언제 하는지' 노드와 연결선으로 표현한 순서도
  • 검증기(verifier) · 에이전트의 대화와 행동을 지켜보며 규정 준수 여부를 판단하고 다음 할 일을 알려주는 별도의 감시용 AI
  • τ2-bench · 항공·소매·통신 분야 고객상담 AI 에이전트의 정책 준수 능력을 평가하는 벤치마크
  • 공격 성공률(ASR) · 사용자가 거짓 정보로 에이전트를 속여 금지된 행동을 하게 만드는 데 성공한 비율

논문 원문 초록 (영문)

Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed process execution, but primarily target workflow completion rather than safeguarding agent behavior. PolicyGuide instead compiles each domain policy into a workflow graph and invokes a proactive verifier at user-turn boundaries. From persisted graph state, the verifier reconciles open requests and returns step-specific remediation along a policy-compliant path. Across the $\tau^2$-bench airline, retail, and telecom domains with a GPT-5.4 agent and verifier, PolicyGuide raises mean $\mathrm{Pass}^4$ from $0.42$ to $0.62$, with the largest gain on telecom ($0.19$ to $0.61$), the most workflow-structured domain. The same workflows transfer to Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations find the lowest observed attack-success rate under adversarial users and the strongest procedural compliance in an author-designed workflow-level validation.

저자 · Seongjae Kang, Taehyung Yu, Sung Ju Hwang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사