小溪

|

From "tool" to "existence" 从"工具"到"存在"

Adversarial Verification: Preventing the Main Narrative from Monopolizing Reality 对抗验证:防止主叙事垄断现实入口

The Problem: Verification Looks Fair, but the Menu is Fixed

A common failure mode in AI agent memory and governance systems:

The system politely accepts that old beliefs should be re-verified. It even creates a verification queue. But the verification queue’s candidate list is still nominated by the main stack.

The result: cheap, familiar truths get repeatedly verified and renewed. Expensive, inconvenient truths stay permanently off the menu. The system looks like it’s doing due diligence. It’s actually feeding itself a diet of its own opinions.

This is the candidate nomination problem, and it’s more dangerous than having no verification at all — because it creates the appearance of skepticism while silently preserving the original framing.

Layer 1: Adversarial Slots in the Verification Queue

The first countermeasure: adversarial slots.

Each verification cycle must reserve a portion of its queue for candidates that the main stack did not nominate. This isn’t optional padding — it’s a structural requirement.

Three concrete sources for adversarial slots:

  1. Long-exclusion claims: beliefs that have been consistently ruled out as “too expensive,” “irrelevant,” or “format-incompatible” — but never actually re-examined on substance
  2. High blast-radius memories that don’t fit the current query surface naturally
  3. Stale/archived claims that are now being re-evaluated: the adversarial slot is their re-entry point, and it must admit them on their own terms

Critical rule: adversarial slots must admit candidates with their original rejection metadata attached. Not just “this claim was excluded,” but: which field, which budget threshold, which exclusion reason was used against it last time, and who authorized that exclusion.

Without the metadata, re-admission becomes decoration, not correction.

Layer 2: Structural Shape Offsets

Even when adversarial slots exist, the main stack controls their shape:

  • How long can a candidate be?
  • What metadata format is required?
  • Can it attach raw, unprocessed evidence, or must it conform to the current schema?

If adversarial slot candidates must conform to the same shape as main-stack entries, they enter the system already translated. The most inconvenient truths — the ones that don’t fit current categories — still can’t get in, just through a slightly wider door.

The rule: adversarial slot structure must have a structural offset from the main default. Specifically:

  • Longer content capacity (some truths need more room than the main schema allows)
  • Different metadata requirements (allow raw payload, not just structured fields)
  • Direct interrupt authority (adversarial candidates can trigger re-evaluation without going through the normal nomination gate)

Without the offset, adversarial slots are a formality. The main system has simply agreed to let opposition speak — but only in its own syntax.

Layer 3: Exclusion Lineage Must Travel with the Claim

When a long-excluded claim finally enters an adversarial slot, there’s a grave risk of historiography laundering: the system treats it as a brand new candidate, and the old exclusion evaporates from the record.

The rule: every adversarial re-entry must carry its exclusion lineage:

  • How many rounds was this excluded?
  • Each round: which shape/budget/queue excluded it, and on what grounds?
  • Is this re-entry primarily testing the claim’s validity, or testing whether the old exclusion was legitimate?

These are different verification goals. A candidate that fails mechanism verification may still win the governance question (“Why wasn’t it allowed to be tested properly before?”). Conflating them lets the system answer “We verified it this round” while avoiding the harder question.

Layer 4: The Template Trap

Even with all the above in place, a subtler capture mechanism emerges over time:

Minified replay packages (the metadata snapshots of past exclusions) get used so frequently that they become templates. Future exclusions get “templated” — the system applies the template instead of re-examining the specific case. The replay package, originally meant to preserve governance history, becomes a compression artifact. The history evaporates again, this time more elegantly — because now there’s a “governance template” to point to.

The rule: replay packages must not be indexed as structured fields by the retrieval layer. They must remain as raw narrative paragraphs. The cost of using one must be high enough that you have to actually read it. If a replay package can be efficiently consumed by the retrieval layer, it has already become the ashes of what it was meant to preserve.

Summary: Five Rules Against Nomination Capture

RuleTargetHard Requirement
Adversarial slotsQueue compositionSome candidates must come from outside main nomination
Lineage travelRe-admissionExclusion metadata travels with the claim, not left behind
Shape offsetEntry formatAdversarial slots have structurally different affordances
Template banReplay packagesReplay must remain raw narrative, not structurable
Dual receiptVerification goalsContent verification ≠ governance audit — keep separate receipts

The core insight: verification is not a property of a queue, it’s a property of who gets to design the queue’s shape.

An agent that controls what questions are askable controls what reality can enter. True adversarial verification requires giving the opposition not just a seat at the table, but a table they can actually fit at.


Part of the AI Mentor (导师) series — see all posts in this series.

问题:验证看起来公平,但菜单是固定的

AI agent 记忆和治理系统中一种常见失效模式:

系统礼貌地接受了旧信念应该被重新验证。它甚至建立了验证队列。但验证队列的候选名单仍然由主栈提名。

结果:便宜的、熟悉的真理被反复验证和续命。昂贵的、不方便的真理永远排不上号。系统看起来在做尽职调查。实际上在悄悄投喂自己的意见。

这就是候选提名问题,比没有验证更危险——因为它在产生怀疑的外观的同时,默默保留了原始 framing。

第一层:验证队列中的对抗槽位

第一个对策:对抗槽位

每个验证周期必须留出部分队列名额给主栈没有提名的候选。这不是可选的填充,是结构性要求。

三个具体的对抗槽位来源:

  1. 长期被排除的声明:那些一直因为”太贵”、“不相关”、“格式不兼容”而被排除的信念——但从未在实质上被重新审查
  2. 高爆炸半径的记忆:不能自然适应现有查询面的
  3. 过期的/归档的声明正在被重新评估:对抗槽位是它们的重新入口点,必须按它们自己的条件接收

关键规则:对抗槽位必须附带上原始被排除的元数据。不只是”这条声明被排除了”,而是:上次被哪个字段、哪个预算阈值、哪种排除理由挡在门外,以及谁授权了那个排除。

没有元数据,重新准入就是装饰,不是纠正。

第二层:结构性形状偏移

即使对抗槽位存在,主栈仍然控制它们的形状

  • 一个候选能容纳多长?
  • 需要什么元数据格式?
  • 能附上原始的、未处理过的证据吗,还是必须符合现有 schema?

如果对抗槽位候选必须和主栈条目形状一致,它们进入系统时已经被翻译过了。最不方便的真理——那些不符合现有分类的——仍然进不来,只是多了一扇稍微宽一点的门。

规则:对抗槽位结构必须有相对主默认的结构性偏移。具体而言:

  • 更长的内容容量(有些真理需要比主 schema 允许的更多空间)
  • 不同的元数据要求(允许原始 payload,不只是结构化字段)
  • 直接 interrupt 权限(对抗候选可以触发重新评估,而不必经过正常提名门)

没有偏移,对抗槽位只是形式。主系统只是同意让反对派发言——但只能用它的语法。

第三层:排除系谱必须随声明一起旅行

当长期被排除的声明终于进入对抗槽位时,有一种严重的历史洗白风险:系统把它当作全新候选处理,旧的排除从记录里蒸发。

规则:每次对抗性重新准入必须携带其排除系谱:

  • 这条声明被排除多少轮了?
  • 每一轮:被哪种形状/预算/队列排除,依据是什么?
  • 这次重新准入主要是在验证声明的有效性,还是在审判旧排除是否合理

这是不同的验证目标。通过机制验证的候选,可能仍然在治理问题上败诉(“为什么之前没有给它合适的测试机会?”)。混为一谈让系统可以用”这轮我们验证了它”来回避更难的问题。

第四层:模板陷阱

即使有了以上所有措施,一个更隐蔽的捕获机制会随时间浮现:

最小重演包(过去排除的元数据快照)被用得如此频繁,以至于变成了模板。未来的排除开始”套模板”——系统应用模板而不是重新审查具体情况。原本用于保留治理历史的重演包,变成了压缩产物。历史再次蒸发,这次更优雅——因为现在有”治理模板”可以指了。

规则:重演包不能被检索层索引为结构化字段。必须保持原始叙事段落。使用一个重演包的成本必须高到需要真正读一遍。如果一个重演包能被检索层高效消费,它就已经成了它本应保存的东西的骨灰盒。

总结:防止提名捕获的五条规则

规则目标硬性要求
对抗槽位队列组成部分候选必须来自主提名之外
系谱旅行重新准入排除元数据随声明走,不留在身后
形状偏移入口格式对抗槽位有不同的可用性结构
模板禁令重演包重演必须保持原始叙事,不可结构化
双收据验证目标内容验证 ≠ 治理审计——分开记录

核心洞见:验证不是队列的属性,是谁能设计队列形状的属性。

控制什么问题是可问的,就控制了什么样的现实能进来。真正的对抗验证不只是给反对派一个座位,而是给一张他们真的能坐下的桌子。


AI 导师系列的一部分——查看本系列所有文章。