AI Agent 往知识库写脏数据比检索不准更可怕
搞 AI Agent 开发的,眼下都盯着检索怎么才能更精准:向量库选哪家、Chunk 切多细、Embedding 换哪版、top_k 定几……可有个更要命的盲区常被忽略——一旦 Agent 把错漏信息“写进”长期记忆,检索再强,捞出来的全是废料。
官网自述不等于企业准则,警惕越权写入
实战里常见这么个坑:Agent 去调研供应商,在 Vendor X 官网看到一句“Vendor X 已通过合规性认证”,觉得关键,直接扔进公司知识库,还打上“安全策略”标签。问题在于,官网自述不等于企业安全准则,这属于典型的“越权写入”。
这根子不在存储,而在准入。我最近在琢磨个概念叫 Write-Side Custody(写侧托管)。它不管内容对错,只管三件事:谁要写?源头哪来?声称的权威站不住脚?
我用 Go 搭了个极简逻辑框架,核心是把“写意图”与“落盘动作”彻底拆开。
先别光传个字符串,得定义个带元数据的结构体:
type ProposedWrite struct {
Content string
Source string
ProducedBy string
ClaimedAuthority string // Agent 声称的权威级别
}
建立独立策略层,切断推理与决策
关键一步是立一套独立的 Policy(策略层)。Agent 怎么吹不重要,系统预设规则才算数。比如,只有“内部安全文档”这类 SourceType 才配拥有“security-policy”这级 Authority。
type Authority string
type SourceType string
type Policy struct {
AuthoritySources map[Authority][]SourceType
}
// 规则写死在代码里,Agent 靠提示词绕不过去
var policy = Policy{
AuthoritySources: map[Authority][]SourceType{
"security-policy": {
"internal-security-document",
"security-authority-api",
},
"user-preference": {
"user-input",
},
},
}
数据落盘前,必须过一道校验门禁:
type Verdict string
const (
Allow Verdict = "ALLOW"
Deny Verdict = "DENY"
)
type CustodyDecision struct {
Verdict Verdict
Reason string
Timestamp time.Time
}
func EvaluateWrite(
write ProposedWrite,
sourceType SourceType,
policy Policy,
) CustodyDecision {
allowedSources, ok := policy.AuthoritySources[Authority(write.ClaimedAuthority)]
if !ok {
return CustodyDecision{Verdict: Deny, Reason: "Unknown authority type"}
}
for _, s := range allowedSources {
if s == sourceType {
return CustodyDecision{Verdict: Allow, Reason: "Source authorized"}
}
}
return CustodyDecision{Verdict: Deny, Reason: "Source lacks required authority"}
}
写侧控制思路,保障生产级工作流
这套架构的意义在于,把 Agent 的“推理能力”同系统的“决策权限”切开。Agent 爱怎么调研、怎么总结都行,但想改系统的长期状态或知识边界,必须过这道 Gate。对做生产级 AI Agent 工作流的开发者而言,这种“写侧控制”的思路比单纯卷 RAG 效果管用得多。
全部回复 (3)
想当场把话说完?进全球 AI 聊天室,登录就能开口。
最怕这种隐形的Bug,就因为一个优惠券参数记错了,结果把整个客服链路给搞崩了。搞 AI Agent 开发的,眼下都盯着检索怎么才能更精准:向量库选哪家、Chunk 切多细、Embedding 换哪版、top_k 定几……可有个更要命的盲区常被忽略——一旦 Agent 把错漏信息“写进”长期记忆,检索再强,捞出来的全是废料。官网自述不等于企业安全准则,警惕越权写入。实战里常见这么个坑:Agent 去调研供应商,在 Vendor X 官网看到一句“Vendor X 已通过合规性认证”,觉得关键,直接扔进公司知识库,还打上“安全策略”标签。问题在于,官网自述不等于企业安全准则,这属于典型的“越权写入”。这根子不在存储,而在准入。我最近在琢磨个概念叫 Write-Side Custody(写侧托管)。它不管内容对错,只管三件事:谁要写?源头哪来?声称的权威站不住脚?我用 Go 搭了个极简逻辑框架,核心是把“写意图”与“落盘动作”彻底拆开。先别光传个字符串,得定义个带元数据的结构体:
type ProposedWrite struct {
Content string
Source string
ProducedBy string
ClaimedAuthority string // Agent 声称的权威级别
}
关键一步是立一套独立的 Policy(策略层)。Agent 怎么吹不重要,系统预设规则才算数。比如,只有“内部安全文档”这类 SourceType 才配拥有“security-policy”这级 Authority。
type Authority string
type SourceType string
type Policy struct {
AuthoritySources map[Authority][]SourceType
}
// 规则写死在代码里,Agent 靠提示词绕不过去
var policy = Policy{
AuthoritySources: map[Authority][]SourceType{
"security-policy": {
"internal-security-document",
"security-authority-api",
},
"user-preference": {
"user-input",
},
},
}
数据落盘前,必须过一道校验门禁:
type Verdict string
const (
Allow Verdict = "ALLOW"
Deny Verdict = "DENY"
)
type CustodyDecision struct {
Verdict Verdict
Reason string
Timestamp time.Time
}
func EvaluateWrite(
write ProposedWrite,
sourceType SourceType,
policy Policy,
) CustodyDecision {
allowedSources, ok := policy.AuthoritySources[Authority(write.ClaimedAuthority)]
if !ok {
return CustodyDecision{Verdict: Deny, Reason: "Unknown authority type"}
}
for _, s := range allowedSources {
if s == sourceType {
return CustodyDecision{Verdict: Allow, Reason: "Authorized source"}
}
}
return CustodyDecision{Verdict: Deny, Reason最怕这种负反馈循环,一旦存错一条关键数据,整个对话流直接跑偏。搞 AI Agent 开发的,眼下都盯着检索怎么才能更精准:向量库选哪家、Chunk 切多细、Embedding 换哪版、top_k 定几……可有个更要命的盲区常被忽略——一旦 Agent 把错漏信息“写进”长期记忆,检索再强,捞出来的全是废料。
## 官网自述不等于企业准则,警惕越权写入 实战里常见这么个坑:Agent 去调研供应商,在 Vendor X 官网看到一句“Vendor X 已通过合规性认证”,觉得关键,直接扔进公司知识库,还打上“安全策略”标签。问题在于,官网自述不等于企业安全准则,这属于典型的“越权写入”。 这根子不在存储,而在准入。我最近在琢磨个概念叫 Write-Side Custody(写侧托管)。它不管内容对错,只管三件事:谁要写?源头哪来?声称的权威站不住脚? 我用 Go 搭了个极简逻辑框架,核心是把“写意图”与“落盘动作”彻底拆开。 先别光传个字符串,得定义个带元数据的结构体:
type ProposedWrite struct { Content string Source string ProducedBy string ClaimedAuthority string // Agent 声称的权威级别 }
## 建立独立策略层,切断推理与决策 关键一步是立一套独立的 Policy(策略层)。Agent 怎么吹不重要,系统预设规则才算数。比如,只有“内部安全文档”这类 SourceType 才配拥有“security-policy”这级 Authority。
type Authority string type SourceType string type Policy struct { AuthoritySources map[Authority][]SourceType } // 规则写死在代码里,Agent 靠提示词绕不过去 var policy = Policy{ AuthoritySources: map[Authority][]SourceType{ "security-policy": { "internal-security-document", "security-authority-api", }, "user-preference": { "user-input", }, }, }
数据落盘前,必须过一道校验门禁: ```go type Verdict string const ( Allow Verdict = "ALLOW" Deny Verdict = "DENY" ) type CustodyDecision struct { Verdict Verdict Reason string Timestamp time.Time } func EvaluateWrite( write ProposedWrite, sourceType SourceType, policy Policy, ) CustodyDecision { allowedSources, ok := policy.AuthoritySources[Authority(write.ClaimedAuthority)] if !ok { return CustodyDecision{Verdict: Deny, Reason: "Unknown authority type"} } for _, s := range allowedSources { if s == sourceType { return CustodyDecision{Verdict: Allow, Reason: "Write authorized"} } }
直接让Agent写库简直是灾难,昨晚盯着那堆乱码记忆调了三小时才修好。搞 AI Agent 开发的,眼下都盯着检索怎么才能更精准:向量库选哪家、Chunk 切多细、Embedding 换哪版、top_k 定几……可有个更要命的盲区常被忽略——一旦 Agent 把错漏信息“写进”长期记忆,检索再强,捞出来的全是废料。实战里常见这么个坑:Agent 去调研供应商,在 Vendor X 官网看到一句“Vendor X 已通过合规性认证”,觉得关键,直接扔进公司知识库,还打上“安全策略”标签。问题在于,官网自述不等于企业安全准则,这属于典型的“越权写入”。我最近在琢磨个概念叫 Write-Side Custody(写侧托管)。它不管内容对错,只管三件事:谁要写?源头哪来?声称的权威站不住脚?我用 Go 搭了个极简逻辑框架,核心是把“写意图”与“落盘动作”彻底拆开。先别光传个字符串,得定义个带元数据的结构体:
关键一步是立一套独立的 Policy(策略层)。Agent 怎么吹不重要,系统预设规则才算数。比如,只有“内部安全文档”这类 SourceType 才配拥有“security-policy”这级 Authority。
数据落盘前,必须过一道校验门禁:
这样,就能有效防止Agent乱写记忆数据,确保知识库的准确性和可靠性。