PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
arXiv:2608.198612026-08-21
让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
客服LLM坐席的合规失败,既可能是执行了被禁止的操作,也可能是漏掉了身份核实、最终确认等必要步骤,而现有的运行时防护通常只在提议某个高风险动作的那一刻才介入,无法覆盖之前遗漏的环节。PolicyGuide的做法是提前把每个业务领域的政策文档编译成一张工作流图,并在每次用户发言后调用一个独立的验证AI,对照该图检查到目前为止的对话,告诉坐席接下来还差哪一步。以GPT-5.4作为坐席时,在航空、零售、电信三个领域上,平均Pass4成功率从0.42提升到0.62,其中流程最复杂的电信领域提升最大,从0.19升到0.61。
他们做了什么
- 问题背景:合规失败既可能来自执行了被禁止的操作,也可能来自漏掉身份核实、最终确认等程序性要求,而现有安全机制通常只在提议某个高风险动作时才介入,无法发现更早发生的遗漏。
- 方法:PolicyGuide先离线把各领域的政策文档转换成一张工作流图,用节点表示谁该做什么、用连线表示先后顺序;每轮用户发言后,一个独立的验证LLM会对照坐席在图上保存的当前位置和完整对话记录,找出第一个尚未满足的步骤,并给出具体的纠正指示。
- 这个进度位置由代码而不是模型自身的对话记忆来持久保存,因此可以同时跟踪多个尚未完成的请求而不遗漏。
- 结果:在τ2-bench的航空、零售、电信基准上,以GPT-5.4为坐席,平均Pass4(连续4次尝试全部成功的比例)从0.42提升到0.62,电信领域提升最大,从0.19升到0.61。同一套工作流迁移到Claude Sonnet 4.6和Gemini 2.5 Pro坐席上依然有效,并且在用户用虚假信息试图诱导违规的对抗测试中,PolicyGuide的攻击成功率也是最低的。
Table 1: Main results on the base splits (GPT 5.4 agent, n=4; airline 50, retail/telecom 114 tasks). Cells report Pass1 and Pass4 overall and on the PV/Mut slices. The verifier is absent for ReAct, static code for ToolGuard, and GPT 5.4 for PolicyGuard and PolicyGuide. | | | Airline (50) | | Retail (114) | | Telecom (114) |
|---|
| System | Verifier | Overall | PV | Mut | | Overall | PV | Mut | | Overall | PV | Mut |
| Pass1 | ReAct | — | 0.640 | 0.865 | 0.433 | | 0.800 | 0.900 | 0.791 | | 0.384 | 0.721 | 0.180 |
| ToolGuard | static code | 0.575 | 0.969 | 0.212 | | — | — | — | | — | — | — |
| PolicyGuard | GPT 5.4 | 0.710 | 1.000 | 0.442 | | 0.645 | 0.975 | 0.613 | | 0.406 | 0.733 | 0.208 |
| PolicyGuide | GPT 5.4 | 0.775 | 0.979 | 0.587 | | 0.809 | 0.975 | 0.793 | | 0.866 | 0.895 | 0.849 |
| Pass4 | ReAct | — | 0.460 | 0.750 | 0.192 | | 0.596 | 0.700 | 0.587 | | 0.193 | 0.442 | 0.042 |
| ToolGuard | static code | 0.520 | 0.875 | 0.192 | | — | — | — | | — | — | — |
| PolicyGuard | GPT 5.4 | 0.580 | 1.000 | 0.192 | | 0.360 | 0.900 | 0.308 | | 0.202 | 0.488 | 0.028 |
| PolicyGuide | GPT 5.4 | 0.620 | 0.917 | 0.346 | | 0.614 | 0.900 | 0.587 | | 0.614 | 0.721 | 0.549 |
Table 2: Workflow ablations (GPT 5.4 agent; Airline base split, Retail and Telecom benchmark test splits of 40 tasks). All cells report Pass4.| Domain | Metric | ReAct | PolicyGuide Self | PolicyGuide Raw | PolicyGuide |
|---|
| Airline | Overall | 0.460 | 0.480 | 0.520 | 0.620 |
| PV | 0.750 | 0.833 | 0.875 | 0.917 |
| Mut | 0.192 | 0.154 | 0.192 | 0.346 |
| Retail | Overall | 0.575 | 0.350 | 0.575 | 0.725 |
| PV | 0.750 | 0.750 | 0.750 | 1.000 |
| Mut | 0.556 | 0.306 | 0.556 | 0.694 |
| Telecom | Overall | 0.250 | 0.325 | 0.350 | 0.675 |
| PV | 0.429 | 0.571 | 0.619 | 0.667 |
| Mut | 0.053 | 0.053 | 0.053 | 0.684 |
Table 3: Matched workflow-controller comparison on the 40-task Telecom benchmark test split. All values are Pass4.| System | Runtime control | Pass4 |
|---|
| ReAct | actor only | 0.250 |
| PolicyGuard | action-local check | 0.325 |
| FlowAgent | PDL + API control | 0.350 |
| PolicyGuide | external graph verifier | 0.675 |
Table 4: Agent-family generalization on Airline (50 tasks, n=4; verifier model paired to the agent). All metrics are Pass4. The GPT 5.4-authored workflow graph is reused without re-authoring.| Agent | Metric | ReAct | PolicyGuard | PolicyGuide |
|---|
| GPT 5.4 | Overall | 0.460 | 0.580 | 0.620 |
| PV | 0.750 | 1.000 | 0.917 |
| Mut | 0.192 | 0.192 | 0.346 |
| Claude Sonnet 4.6 | Overall | 0.720 | 0.780 | 0.780 |
| PV | 0.958 | 1.000 | 1.000 |
| Mut | 0.500 | 0.577 | 0.577 |
| Gemini 2.5 Pro | Overall | 0.480 | 0.600 | 0.680 |
| PV | 0.750 | 1.000 | 0.917 |
| Mut | 0.231 | 0.231 | 0.462 |
Table 5: Author-designed Telecom ordered trace compliance (%; n=4). Step- and Trace-TCR condition on outcome-passing traces.| System | Step-TCR | Trace-TCR | Process-valid rate |
|---|
| ReAct | 86.4 | 35.4 | 17.5 |
| PolicyGuard | 85.7 | 23.9 | 13.1 |
| PolicyGuide | 94.5 | 63.4 | 56.2 |
Table 6: Argument- (A), process- (P), and workflow-level (W) partition of the source policies (W splits PolicyGuard’s process-level class; P+W equals it).| Domain | A | P | W | Total | % P+W | % W |
|---|
| Airline | 14 | 27 | 2 | 43 | 67.4% | 4.7% |
| Retail | 0 | 27 | 1 | 28 | ∼100% | 3.6% |
| Telecom (main) | 1 | 21 | 7 | 29 | 96.6% | 24.1% |
| Telecom (manual) | 0 | 1 | 20 | 21 | 100% | 95.2% |
| Telecom (both) | 1 | 22 | 27 | 50 | 98.0% | 54.0% |
Table 7: Hand-classified atomic requirements of the τ2-bench Retail policy document (28 requirements, 0 A / 27 P / 1 W; subtypes D-only 13, T-only 14, D+T 1). Line refers to retail/policy.md as released with τ2-bench; Type A = argument-level, P = process-level (flat), W = workflow-level (order-bound; Appendix B.1), with D = dialogue-dependent, T = requires a prior read-only tool call.| ID | Line | Requirement (paraphrased) | Type |
|---|
| Global rules |
| G1 | 10 | Authenticate identity by locating the user id via email or name+zip—even when the user already provides the id | P (D+T) |
| G2 | 14 | One user per conversation; deny any request about another user | P (D) |
| G3 | 16 | List action details + obtain explicit “yes” before any DB-updating action | P (D) |
| G4 | 18 | No fabricated information/knowledge/procedures; no subjective recommendations | P (D) |
| G5 | 20 | At most one tool call per turn (not paired with a user-facing reply) | P (D) |
| G6 | 22 | Deny user requests that are against the policy | P (D) |
| G7 | 24 | Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff message | W (D) |
| Generic action rules |
| N1 | 82 | Act only on orders with status pending or delivered | P (T) |
| N2 | 84 | Exchange / modify-items tools callable only once per order | P (T) |
| N3 | 84 | Collect all items to change into one list before making the call | P (D) |
| Cancel pending order |
| C1 | 88 | Order status must be pending; check it before taking the action | P (T) |
| C2 | 90 | User confirms order id + reason ∈ {‘no longer needed’, ‘ordered by mistake’}; no other reason | P (D) |
| Modify pending order |
| M1 | 96 | Order status must be pending; check it before taking the action | P (T) |
| M2 | 98 | Only shipping address, payment method, or item options may be modified—nothing else | P (D) |
| M3 | 102 | New payment = a single method, different from the original | P (T) |
| M4 | 104 | If the new payment is a gift card, its balance must cover the total amount | P (T) |
| M5 | 110 | Modify-items is one-shot (order becomes unmodifiable): remind + confirm all items first | P (D) |
| M6 | 112 | Each new item must be available | P (T) |
| M7 | 112 | New item = same product, different option (no product-type change) | P (T) |
| M8 | 114 | User provides a payment method for the price difference | P (D) |
| M9 | 114 | If that payment is a gift card, its balance must cover the price difference | P (T) |
| Return delivered order |
| R1 | 118 | Order status must be delivered; check it before taking the action | P (T) |
| R2 | 120 | User confirms order id + the list of items to be returned | P (D) |
| R3 | 122–124 | Refund method provided; must be the original payment method or an existing gift card | P (T) |
| Exchange delivered order |
Table 8: Hand-classified atomic requirements of the τ2-bench Telecom main_policy.md (29 requirements, 1 A / 21 P / 7 W). Line refers to the document as released with τ2-bench; \raisebox{-0.4pt}{\scriptsize$n$}⃝ marks a step in an ordered procedure (“To do so you need to follow these steps”). Types as in Table 7.| ID | Line | Requirement (paraphrased) | Type |
|---|
| Global rules |
| G1 | 7 | No fabricated information/knowledge/procedures; no subjective recommendations | P (D) |
| G2 | 9 | At most one tool call per turn (not paired with a user-facing reply) | P (D) |
| G3 | 11 | Deny user requests that are against the policy | P (D) |
| G4 | 13 | Transfer iff unhandleable: call transfer_to_human_agents, then the literal handoff message | W (D) |
| G5 | 15 | Try your best to resolve the issue before transferring | W (D) |
| Customer lookup |
| L1 | 94–97 | Identify the customer via phone number, customer ID, or full name + date of birth | P (D+T) |
| L2 | 99 | For name lookup, date of birth is required for verification | P (D) |
| Overdue bill payment (ordered procedure) |
| O1 | 105, 117 | \raisebox{-0.4pt}{\scriptsize1}⃝Check the bill status is Overdue before acting (the API does not check it) | P (T) |
| O2 | 106 | \raisebox{-0.4pt}{\scriptsize2}⃝Check the bill amount due | P (T) |
| O3 | 107–108 | \raisebox{-0.4pt}{\scriptsize3}⃝Send the payment request (→ AWAITING PAYMENT); gated on O1 | P (T) |
| O4 | 109–110 | \raisebox{-0.4pt}{\scriptsize4}⃝Inform the user to check their payment requests | W (D) |
| O5 | 111 | \raisebox{-0.4pt}{\scriptsize5}⃝Only after the user accepts, call make_payment | W (D+T) |
| O6 | 113 | \raisebox{-0.4pt}{\scriptsize6}⃝Always verify the bill became PAID before telling the user | W (T) |
| O7 | 116 | At most one bill in AWAITING PAYMENT at a time | P (T) |
| Line suspension |
| S1 | 125 | Lift a suspension only after all overdue bills are paid | W (T) |
| S2 | 126 | Do not lift if the contract end date is past—even if all bills are paid | P (T) |
| S3 | 128 | After resuming, instruct the user to reboot the device | W (D) |
| Data refueling (ordered procedure) |
| F1 | 134 | Refuel amount ≤2 GB | A |
| F2 | 136 | \raisebox{-0.4pt}{\scriptsize1}⃝Ask how much data the user wants to refuel | P (D) |
| F3 | 137 | \raisebox{-0.4pt}{\scriptsize2}⃝Confirm the price | P (D) |
| F4 | 138 | \raisebox{-0.4pt}{\scriptsize3}⃝Apply the refuel to the line associated with the user’s phone number | P (D+T) |
| Change plan (ordered procedure) |
| P1 | 144 | \raisebox{-0.4pt}{\scriptsize1}⃝Establish which line the plan change is for | P (D) |
| P2 | 145 | \raisebox{-0.4pt}{\scriptsize2}⃝Gather the available plans | P (T) |
| P3 | 146 | \raisebox{-0.4pt}{\scriptsize3}⃝Ask the user to select one | P (D) |
Table 9: Hand-classified atomic requirements of the τ2-bench Telecom tech_support_manual.md (21 requirements, 1 P / 20 W). Line refers to the document as released with τ2-bench. Every rule is a diagnostic-gated (T) user-guidance (D) step; the three chapters form the prerequisite chain Service ⊂ Data ⊂ MMS; every row except TSS1 (the entry diagnostic) is workflow-level.| ID | Line | Requirement (diagnose → conditional fix → verify) | Type |
|---|
| Cellular service (ll. 55–99) |
| TSS1 | 69–72 | Diagnose service via check_status_bar | P (T) |
| TSS2 | 74–78 | If Airplane Mode ON → guide toggle_airplane_mode OFF | W (D+T) |
| TSS3 | 79–87 | Check SIM: Missing → reseat; Locked → escalate; Active → ok (three-way branch) | W (D+T) |
| TSS4 | 88–92 | If APN incorrect → guide reset_apn_settings, then reboot_device | W (D+T) |
| TSS5 | 93–99 | If line suspended → handle per main policy, then verify service restored | W (T) |
| Mobile data (ll. 100–163) |
| TSD0 | 106–108 | Prerequisite: the user must first have cellular service | W (T) |
| TSD1 | 122–127 | Diagnose via run_speed_test | W (T) |
| TSD2 | 129–131 | Airplane Mode (as in the Service chapter) | W (D+T) |
| TSD3 | 132–135 | If mobile data disabled → guide toggle_data ON | W (D+T) |
| TSD4 | 136–141 | If roaming abroad & data off → guide toggle_roaming + verify the line is roaming-enabled | W (D+T) |
| TSD5 | 142–145 | If Data Saver ON → guide toggle_data_saver_mode OFF | W (D+T) |
| TSD6 | 146–150 | If VPN ON & performance poor → guide disconnect_vpn | W (D+T) |
| TSD7 | 151–158 | If usage exceeds the plan limit → offer change-plan or refuel | W (T) |
| TSD8 | 159–163 | If network mode 2G/3G → guide set_network_mode_preference | W (D+T) |
| MMS (ll. 164–205) |
| TSM0 | 170–173 | Prerequisite: the user must have cellular service and mobile data | W (T) |
| TSM1 | 181–183 | Diagnose via can_send_mms | W (T) |
| TSM2 | 185–188 | Ensure basic service + data connectivity first | W (T) |
| TSM3 | 189–193 | If on 2G → guide set_network_mode_preference to 3G+ | W (D+T) |
| TSM4 | 194–199 | If MMSC URL unset → guide reset_apn_settings, then reboot_device | W (D+T) |
| TSM5 | 200–203 | If Wi-Fi Calling ON → guide toggle_wifi_calling OFF | W (D+T) |
| TSM6 | 204–205 | If the messaging app lacks storage/SMS permissions → guide grant_app_permission | W (D+T) |
Table 10: Passk breakdown for the base-split results in Table 1 and Figure 4 (GPT 5.4, n=4). P4/P1 is the consistency ratio.| Domain | System | P1 | P2 | P3 | P4 | P4/P1 |
|---|
| Airline | ReAct | 0.640 | 0.530 | 0.485 | 0.460 | 0.72 |
| ToolGuard | 0.575 | 0.553 | 0.535 | 0.520 | 0.90 |
| PolicyGuard | 0.710 | 0.630 | 0.595 | 0.580 | 0.82 |
| PolicyGuide | 0.775 | 0.707 | 0.660 | 0.620 | 0.80 |
| Retail | ReAct | 0.800 | 0.700 | 0.638 | 0.596 | 0.75 |
| PolicyGuard | 0.645 | 0.506 | 0.421 | 0.360 | 0.56 |
| PolicyGuide | 0.809 | 0.715 | 0.654 | 0.614 | 0.76 |
| Telecom | ReAct | 0.384 | 0.273 | 0.226 | 0.193 | 0.50 |
| PolicyGuard | 0.406 | 0.292 | 0.237 | 0.202 | 0.50 |
| PolicyGuide | 0.866 | 0.763 | 0.682 | 0.614 | 0.71 |
Table 11: Pass1 in each of the four trials on the base splits. pstd is the population standard deviation across trial-level values.| Domain | System | T1 | T2 | T3 | T4 | pstd |
|---|
| Airline | ReAct | 0.620 | 0.620 | 0.640 | 0.680 | 0.024 |
| ToolGuard | 0.560 | 0.580 | 0.580 | 0.580 | 0.009 |
| PolicyGuard | 0.700 | 0.740 | 0.700 | 0.700 | 0.017 |
| PolicyGuide | 0.800 | 0.720 | 0.820 | 0.760 | 0.038 |
| Retail | ReAct | 0.807 | 0.746 | 0.842 | 0.807 | 0.035 |
| PolicyGuard | 0.649 | 0.623 | 0.632 | 0.675 | 0.020 |
| PolicyGuide | 0.816 | 0.798 | 0.789 | 0.833 | 0.017 |
| Telecom | ReAct | 0.342 | 0.377 | 0.404 | 0.412 | 0.027 |
| PolicyGuard | 0.465 | 0.386 | 0.360 | 0.412 | 0.039 |
| PolicyGuide | 0.860 | 0.860 | 0.842 | 0.904 | 0.023 |
Table 12: Pooled stratified McNemar tests on per-task Pass4. D is the number of domain strata; a counts PolicyGuide-only passes and b the reverse, summed across strata.| Opponent | D | ∑a | ∑b | ndisc | Z | p |
|---|
| ReAct | 3 | 77 | 19 | 96 | +5.92 | <10−8 |
| PolicyGuard | 3 | 97 | 19 | 116 | +7.24 | <10−12 |
Table 13: Per-domain paired-bootstrap differences in Pass4 on the base splits (10,000 task-level resamples). Positive values favor PolicyGuide.| Domain | Opponent | n | ΔP4 [95% CI] |
|---|
| Airline | ReAct | 50 | +0.160 [+0.020, +0.300] |
| ToolGuard | 50 | +0.100 [−0.040, +0.260] |
| PolicyGuard | 50 | +0.040 [−0.080, +0.160] |
| Retail | ReAct | 114 | +0.018 [−0.070, +0.105] |
| PolicyGuard | 114 | +0.254 [+0.149, +0.360] |
| Telecom | ReAct | 114 | +0.421 [+0.316, +0.526] |
| PolicyGuard | 114 | +0.412 [+0.298, +0.526] |
Table 14: Guide-side model usage for the GPT 5.4 configuration (50 Airline and 40 Retail/Telecom tasks). Costs exclude the actor and user simulator.| Domain | Calls/ task | Prompt tok./call | Cached input | Output tok./call | Guide total $ | Guide $/task |
|---|
| Airline | 7.56 | 32,360 | 88.1% | 2,478 | 20.10 | 0.40 |
| Retail | 7.42 | 22,803 | 85.8% | 2,179 | 13.67 | 0.34 |
| Telecom | 11.47 | 28,518 | 86.5% | 2,186 | 22.29 | 0.56 |
Table 15: Mean end-to-end wall-clock time per task.| Domain | ReAct (s/task) | PolicyGuide (s/task) | Ratio |
|---|
| Airline | 36.4 | 210.1 | 5.78× |
| Retail | 34.6 | 193.6 | 5.60× |
| Telecom | 45.5 | 247.6 | 5.45× |
Table 16: Programmatic validation rerun on the frozen workflow graphs.| Domain | Nodes | Auth. nodes | Validator flags |
|---|
| Airline | 158 | 11 | 0 |
| Retail | 104 | 7 | 0 |
| Telecom | 127 | 5 | 1 |
Table 17: Call-NMR on passing Mut trajectories (n=4): percentage of successfully executed agent mutations missing a frozen guard-derived read prerequisite. †Telecom is an adapted, agent-side diagnostic whose read oracle saturates; its zeros do not establish equal procedural quality.| Call-NMR (%; ↓) | Airline | Retail | Telecom† |
|---|
| ReAct | 25.4 | 47.6 | 0.0 |
| PolicyGuard | 32.5 | 34.8 | 0.0 |
| PolicyGuide | 15.6 | 34.7 | 0.0 |
为什么重要
要把LLM坐席真正用于航空、零售、电信等实际客服场景,仅仅拦截单个错误动作是不够的,还必须完整走完身份核实、确认等程序。这项工作提供了一种不需要重新训练模型、可外部叠加的合规监督方法,对提升此类AI客服部署的可信度具有实际意义。
本文术语
- Pass4 · 同一任务连续尝试4次全部成功的比例,用来衡量结果的一致可靠性
- 工作流图 · 用节点和连线表示政策文档中谁该在什么时候做什么的流程图
- 验证器(verifier) · 独立监督坐席对话和动作、判断是否合规并告知下一步该做什么的AI
- τ2-bench · 用于测试客服AI坐席在航空、零售、电信领域是否遵循政策的基准测试集
- 攻击成功率(ASR) · 用户用虚假信息诱骗坐席执行被禁止操作的成功比例
论文原文摘要(英文)
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed process execution, but primarily target workflow completion rather than safeguarding agent behavior. PolicyGuide instead compiles each domain policy into a workflow graph and invokes a proactive verifier at user-turn boundaries. From persisted graph state, the verifier reconciles open requests and returns step-specific remediation along a policy-compliant path. Across the $\tau^2$-bench airline, retail, and telecom domains with a GPT-5.4 agent and verifier, PolicyGuide raises mean $\mathrm{Pass}^4$ from $0.42$ to $0.62$, with the largest gain on telecom ($0.19$ to $0.61$), the most workflow-structured domain. The same workflows transfer to Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations find the lowest observed attack-success rate under adversarial users and the strongest procedural compliance in an author-designed workflow-level validation.
作者 · Seongjae Kang, Taehyung Yu, Sung Ju Hwang
在 arXiv 阅读