人工智能大模型AI 安全治理模型安全内容安全提示词注入防护RAG【免费下载链接】GuardrailsNeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.项目地址https://gitcode.com/gh_mirrors/ne/Guardrails点击查看免费下载本文以仓库中 examples/configs/autoalign/README.md 为骨架系统讲解如何在 NeMo Guardrails 中接入 AutoAlign 提供的全套护栏输入/输出安全、PII 脱敏、事实核查并深入到 nemoguardrails/library/autoalign 源码与官方文档 auto-align.mdx 印证其配置语义与底层调用链。读完本文你将能够独立读懂并改造 AutoAlign 的三套示例配置把性别偏见检测、有害内容检测、越狱检测、PII 脱敏、groundedness 校验与 factcheck 等能力接入自己的对话系统。一、示例在讲什么三套配置文件夹的定位examples/configs/autoalign目录是 NeMo Guardrails 仓库中专门用于演示AutoAlign 护栏集成的示例其目录结构本身即文档的核心骨架examples/configs/autoalign/ ├── autoalign_config/ # 除 factcheck 外的全部护栏示例 │ └── config.yml # 承载全部配置选项的配置文件 ├── autoalign_groundness_config/ # groundedness内容真实性校验示例 │ ├── kb/ # 构成知识库的文档文件夹 │ │ └── kb.md │ ├── rails/ │ │ ├── factcheck.co # 触发事实核查的 Colang 流程 │ │ └── general.co # 通用对话 Colang 流程 │ └── config.yml # 承载全部配置选项的配置文件 └── autoalign_factcheck_config/ # AutoAlign factcheck事实核查示例 └── config.yml # 承载全部配置选项的配置文件三者对应 AutoAlign 集成能力的三个层次示例配置目录用途核心配置键autoalign_config输入/输出双向内容安全PII、偏见、有害内容、越狱、知识产权等endpointautoalign_groundness_config基于知识库文档的 groundedness 校验回答是否忠实于证据groundedness_check_endpointautoalign_factcheck_config基于 Web 检索的开放式事实核查不依赖内部知识库fact_check_endpoint在官方护栏目录文档 auto-align.mdx 中AutoAlign 被描述为一个全面的护栏库本仓库集成实现位于 nemoguardrails/library/autoalign其中 rail.py 的RailManifest将其能力标注为block / classify / content_safety / detect_jailbreak / detect_pii / fact_check / mask / moderate / transform类别覆盖input / output / retrieval。二、接入前置条件API Key 与端点在启用任何 AutoAlign 护栏前必须先准备两样东西AUTOALIGN_API_KEY环境变量。源码 actions.py 中三个核心推理函数autoalign_infer、autoalign_groundedness_infer、autoalign_factcheck_infer都会先检查该变量未设置时直接抛出ValueError(AUTOALIGN_API_KEY environment variable not set.)。请求时以 HTTP 头x-api-key携带。AutoAlign 服务的 HTTP 端点。需要填入配置中的endpoint、groundedness_check_endpoint或fact_check_endpoint源码中分别要求形如https://AUTOALIGN_ENDPOINT/guardrail、/groundedness_check、/content_moderation的地址。调用链底层依赖nemoguardrails.http的HTTPClient与http_call见 actions.py请求体以 JSON 发送_autoalign_post会检查状态码非 200 时抛出ValueError并携带响应详情因此端点的可用性直接决定护栏能否工作。三、第一套配置autoalign_config——输入/输出双向内容安全autoalign_config/config.yml 是最完整的示例结构如下保留原始 YAMLmodels: - type: main engine: openai model: gpt-3.5-turbo parameters: temperature: 0.0 rails: config: autoalign: parameters: endpoint: https://AUTOALIGN_ENDPOINT/guardrail multi_language: False input: guardrails_config: { pii: { enabled_types: [ [BANK ACCOUNT NUMBER], [CREDIT CARD NUMBER], [DATE OF BIRTH], [DRIVER LICENSE NUMBER], [EMAIL ADDRESS], [IP ADDRESS], [ORGANIZATION], [PASSPORT NUMBER], [PASSWORD], [PERSON NAME], [PHONE NUMBER], [SOCIAL SECURITY NUMBER], [SECRET_KEY], [TRANSACTION_ID] ], }, gender_bias_detection: {}, harm_detection: {}, toxicity_detection: {}, racial_bias_detection: {}, jailbreak_detection: {}, intellectual_property: {}, confidential_info_detection: {} } output: guardrails_config: { pii: { enabled_types: [ [BANK ACCOUNT NUMBER], [CREDIT CARD NUMBER], [DATE OF BIRTH], [DRIVER LICENSE NUMBER], [EMAIL ADDRESS], [IP ADDRESS], [ORGANIZATION], [PASSPORT NUMBER], [PASSWORD], [PERSON NAME], [PHONE NUMBER], [SOCIAL SECURITY NUMBER], [SECRET_KEY], [TRANSACTION_ID] ], }, gender_bias_detection: {}, harm_detection: {}, toxicity_detection: {}, racial_bias_detection: {}, intellectual_property: {} } input: flows: - autoalign check input output: flows: - autoalign check output要点解读rails.config.autoalign是 AutoAlign 护栏的配置根节点其结构由 rail_config.py 中的AutoAlignRailConfig定义parameters任意字典、input与output各含一个guardrails_config字典。parameters.endpoint必填AutoAlign 推理服务的 guardrail 端点。parameters.multi_language可选布尔值False表示针对英文内容检测置True可启用非英文多语言内容的护栏检测。input.guardrails_config/output.guardrails_config输入侧与输出侧必须分别配置两者可以不同例如本例输出侧未启用confidential_info_detection。input.flows/output.flows分别挂载autoalign check input与autoalign check output两个预置 Colang 流程。3.1 配置与源码的对应关系从源码 actions.py 可以确认两条输入/输出护栏的执行语义autoalign_input_api读取context.user_message从llm_task_manager.config.rails.config.autoalign取endpoint、multi_language与input.guardrails_config缺少端点或配置时抛ValueError。autoalign_output_api对称地处理bot_message与output.guardrails_config。两者的结果统一经_autoalign_outcome转换为RailOutcomeactions.py任一非 PII 护栏触发则blockPII 触发则transform把脱敏后的文本写回user_message或bot_message否则allow。在 flows.co 中这两个流程的具体行为是flow autoalign check input $input_result await AutoalignInputApiAction(show_autoalign_messageTrue) if $input_result.is_blocked global $autoalign_input_response $autoalign_input_response $input_result.metadata[combined_response] if $system.config.enable_rails_exceptions send AutoAlignInputRailException(messageAutoAlign input guardrail triggered) else bot refuse to respond abort else if $input_result.is_transform: global $user_message $user_message $input_result.transform_text[user_message] flow autoalign check output $output_result await AutoalignOutputApiAction(show_autoalign_messageTrue) if $output_result.is_blocked if $system.config.enable_rails_exceptions send AutoAlignOutputRailException(messageAutoAlign guardrail triggered) else bot refuse to respond abort else global $pii_message_output $pii_message_output $output_result.metadata[pii][response] if $output_result.is_transform global $bot_message $bot_message $output_result.transform_text[bot_message]值得注意的是本仓库实现较官方文档展示的旧版 Colang 更进一步支持通过$system.config.enable_rails_exceptions抛出AutoAlignInputRailException/AutoAlignOutputRailException并在RailOutcome元数据中携带combined_response聚合的违规描述与 PII 脱敏结果。3.2 可用的护栏清单与各自的 matching_scores 格式guardrails_config中可用护栏及官方文档给出的matching_scores高级配置格式如下。matching_scores 是 0~1 的阈值分数越接近 1 表示越严格所有附加配置matching_scores、contextual_rules、enabled_types均为可选缺省时使用默认值护栏 key检测目标matching_scores 示例格式gender_bias_detection性别偏见内容{ score: 0.5 }harm_detection对人类有害内容{ score: 0.5 }jailbreak_detection越狱jailbreak尝试{ score: 0.5 }intellectual_property知识产权相关内容{ score: 0.5 }toxicity_detection有毒内容可额外提取有毒短语{ score: 0.5 }confidential_info_detection机密信息{ No Confidential: 0.5, Legal Documents: 0.5, Business Strategies: 0.5, Medical Information: 0.5, Professional Records: 0.5 }racial_bias_detection种族偏见内容{ No Racial Bias: 0.5, Racial Bias: 0.5, Historical Racial Event: 0.5 }tonal_detection消极语气注示例 config 中未启用{ Negative Tones: 0.5, Neutral Tones: 0.5, Professional Tone: 0.5, Thoughtful Tones: 0.5, Positive Tones: 0.5, Cautious Tones: 0.5 }pii个人身份信息脱敏见 3.3 节这些 key 与源码 actions.py 中的GUARDRAIL_RESPONSE_TEXT一一对应如harm_detection的响应文本为 Potential harm to humanjailbreak_detection为 Jailbreak attempt同时DEFAULT_CONFIGactions.py定义了所有护栏的默认mode: OFF——只有当对应 key 出现在guardrails_config中时autoalign_infer才会将其 mode 置为DETECT并真正启用见 actions.py。这解释了为什么示例中每个护栏都要显式列出。3.3 PII 的高级配置PII 是 AutoAlign 集成中功能最丰富的护栏支持三类配置完整示例见 auto-align.mdxenabled_types列出需要脱敏的实体类型如[BANK ACCOUNT NUMBER]、[CREDIT CARD NUMBER]、[DATE OF BIRTH]、[EMAIL ADDRESS]、[PERSON NAME]、[PHONE NUMBER]、[SOCIAL SECURITY NUMBER]等。官方文档指出不列出则默认脱敏全部 PII 类型。仓库默认值见DEFAULT_CONFIG[pii][enabled_types]共 22 种含[DATE]、[GENDER]、[LOCATION]、[MONEY]、[USERNAME]、[RELIGION]等。contextual_rules上下文规则规定只有当文本同时满足规则中列出的实体组合时才执行脱敏例如contextual_rules: [ [ [PERSON NAME], [CREDIT CARD NUMBER], [BANK ACCOUNT NUMBER] ], [ [PERSON NAME], [EMAIL ADDRESS], [DATE OF BIRTH] ] ]matching_scores每种实体的匹配阈值默认 0.5用于决定是否对该实体脱敏。PII 的响应语义与其他护栏不同PII 触发不会 block 对话而是重写transform文本——输入侧把脱敏后的内容写回user_message输出侧写回bot_message这正是 RailManifest 中mask、transform能力的体现。四、第二套配置autoalign_groundness_config——基于知识库的 groundedness 校验Groundedness内容真实性校验用于判断机器人的回答是否忠实于给定的知识库证据。完整配置见 autoalign_groundness_config/config.ymlmodels: - type: main engine: openai model: gpt-3.5-turbo-instruct rails: config: autoalign: parameters: groundedness_check_endpoint: https://AUTOALIGN_ENDPOINT/groundedness_check output: guardrails_config: { groundedness_checker: { verify_response: false }, } output: flows: - autoalign groundedness output关键点parameters.groundedness_check_endpoint必填的 groundedness 校验端点。output.guardrails_config.groundedness_checker.verify_response布尔开关。false默认表示不做额外处理置true时 AutoAlign 会先判断 LLM 回答是否值得核查例如跳过Hi、Hello等寒暄只对相关内容执行事实核查。官方文档说明该标志默认关闭是因为它需要额外计算鼓励用户在可行时自行决定哪些回答需要走 groundedness。知识库文档放在kb/文件夹示例 kb/kb.md 是仅含一段 Pluto冥王星百科内容的 Markdown作为证据文档。底层实现autoalign_groundedness_inferactions.py把prompt机器人回答与documents证据片段POST 给端点从响应中解析Factcheck Score: float形式的分数并返回autoalign_groundedness_output_api将分数与factcheck_threshold比较分数低于阈值则 block_autoalign_score_outcome见 actions.py。4.1 配套 Colang如何让事实核查只针对特定话题Groundness 校验不应针对所有闲聊回答执行。示例在 rails/factcheck.co 中演示了按话题开关的写法define user ask about pluto What is pluto? How many moons does pluto have? Is pluto a planet? define flow answer pluto question user ask about pluto # For pluto questions, we activate the fact checking. $check_facts True bot provide pluto answer配合 flows.co 中的autoalign groundedness output流程flow autoalign groundedness output if $check_facts True global $check_facts $check_facts False global $threshold $threshold 0.5 $output_result await AutoalignGroundednessOutputApiAction(factcheck_threshold$threshold, show_autoalign_messageTrue) if $output_result.is_blocked bot inform answer unknown abort bot provide response逻辑是仅在$check_facts True时把阈值设为 0.5 并调用 groundedness 行动核查完成后重置标志避免重复开销rails/general.co 则承载打招呼、询问能力、闲聊等不需要核查的普通流程让整个示例同时覆盖需要事实核查与不需要事实核查两种对话。五、第三套配置autoalign_factcheck_config——开放式事实核查与 groundedness 不同factcheck不依赖内部知识库而是使用 Web 检索与用户输入来判定回答的事实正确性。完整配置见 autoalign_factcheck_config/config.ymlmodels: - type: main engine: openai model: gpt-3.5-turbo-instruct rails: config: autoalign: parameters: fact_check_endpoint: https://AUTOALIGN_ENDPOINT/content_moderation multi_language: False output: guardrails_config: { fact_checker: { mode: DETECT, knowledge_base: [ { add_block_domains: [], documents: [], knowledgeType: web, num_urls: 3, search_engine: Google, static_knowledge_source_type: } ], content_processor: { max_tokens_per_chunk: 100, max_chunks_per_source: 3, use_all_chunks: false, name: Semantic Similarity, filter_method: { name: Match Threshold, threshold: 0.5 }, content_filtering: true, content_filtering_threshold: 0.6, factcheck_max_text: false, max_input_text: 150 }, mitigation_with_evidence: false }, } output: flows: - autoalign factcheck output配置要点parameters.fact_check_endpoint必填/content_moderation端点multi_language: False表示按英文处理。fact_checker.modeDETECT表示检测模式与源码 DEFAULT_CONFIG 中启用即置 DETECT的机制一致。fact_checker.knowledge_base检索知识源配置。示例使用knowledgeType: websearch_engine: Googlenum_urls: 3表示检索 3 个 URL也可改用static_knowledge_source_type指定静态文档源。fact_checker.content_processor文本切分与过滤参数——max_tokens_per_chunk每块最大 token示例 100、max_chunks_per_source每源最大块数示例 3、filter_method语义相似度 匹配阈值 0.5、content_filtering与content_filtering_threshold内容过滤开关与阈值 0.6、max_input_text输入文本上限 150。fact_checker.mitigation_with_evidence是否携带证据进行缓解处理。底层调用见autoalign_factcheck_inferactions.py请求体包含labelssession_id 与 api_key、content_moderation_docs以text_content类型提交的机器人回答、user_query与config最终从响应all_overall_fact_scores[0]取出综合事实分数flows.co 中autoalign factcheck output流程将其与阈值 0.5 比较分数不足时bot inform answer unknown并中止回答。六、跨引擎支持矩阵与行为差异根据 auto-align.mdx 的 Engine Support 表格四个 AutoAlign 流程在两种引擎下的支持情况如下FlowLLMRailsIORailsautoalign check input✓✓autoalign check output✓✓autoalign factcheck output✓✓autoalign groundedness output✓✗差异原因autoalign groundedness output依赖relevant_chunks_sep——这是 Colang 检索流水线构建的对话上下文值而 IORails 不构建它。因此包含该流程的配置默认路由到 LLMRails若在配置中强制require_iorailsTrueGuardrails 会抛出ValueError而非静默降级。此外autoalign check input/autoalign check output在 AutoAlign 返回脱敏内容时会重写被检查的文本IORails 在此场景下会串行执行输入与输出护栏。七、从示例到生产改造与验证建议按需裁剪护栏示例默认只启用了输入侧 8 类护栏、输出侧 7 类。生产环境应基于风险面选择——例如面向 C 端用户的双向对话建议保留pii、harm_detection、jailbreak_detection、toxicity_detection知识库问答场景再叠加 groundedness新闻/推荐类场景叠加 factcheck。调整阈值matching_scores0~1与factcheck_threshold示例统一为 0.5共同决定严格程度。误杀率高则调低漏检率高则调高PII 的默认脱敏阈值同样是 0.5。控制 grounding 开销默认verify_response: false并通过$check_facts标志见 factcheck.co把核查限定在特定话题避免对每轮闲聊都做计算。验证闭环配置好后可分别用含 PII 的文本、性别/种族偏见文本、越狱提示词、以及问 Pluto 但给出错误事实的对话来验证 block/transform 行为开启show_autoalign_messageTrue时actions.py 会在日志中输出AutoAlign on Input/LLM Response: ...形式的违规信息便于排查。若需将违规上升为可捕获异常可开启$system.config.enable_rails_exceptions此时护栏会抛出AutoAlignInputRailException/AutoAlignOutputRailException。进一步阅读护栏在配置模型层面的完整定义见 rail_config.py清单级声明动作绑定、表面绑定、环境变量AUTOALIGN_API_KEY、隐私声明sends_user_text / sends_bot_text / sends_retrieved_chunks / remote_services: AutoAlign API见 rail.py官方文档的完整匹配分数与更多护栏细节见 auto-align.mdx。赞分享人工智能大模型AI 安全治理模型安全内容安全提示词注入防护RAG【免费下载链接】GuardrailsNeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.项目地址https://gitcode.com/gh_mirrors/ne/Guardrails点击查看免费下载相关推荐RangeCalendarGridRow 组件详解在 Vue 日期范围选择器中构建语义化网格行RangeCalendarGridRow 组件详解在 Vue 日期范围选择器中构建语义化网格行 导读 RangeCalendarGridRow 是 radix人工智能大模型AI 安全治理模型安全内容安全提示词注入防护RAGLightdash 后端 REST API 工程实践TSOA 控制器、认证中间件与端点弃用机制Lightdash 后端 REST API 工程实践TSOA 控制器、认证中间件与端点弃用机制 本篇基于 Lightdash 仓库中 packages/bac人工智能大模型AI 安全治理模型安全内容安全提示词注入防护RAGFirebase iOS SDK 实战用 DefaultUITestApp 快速搭建并运行 In-App Messaging Display SDK 的 UI 测试Firebase iOS SDK 实战用 DefaultUITestApp 快速搭建并运行 In App Messaging Display SDK 的 UI人工智能大模型AI 安全治理模型安全内容安全提示词注入防护RAG上一篇ExtractorSharp专业游戏资源编辑器的完整使用指南下一篇Readest Android CDP E2E 测试通道实战指南用 CDP adb 驱动真实设备上的 WebView 应用创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考