跳到主要内容
Supermarket
返回能力市场
Agent Pack
design
MIT

anthropic-anti-hallucinate-skills

Anti-hallucination guidelines for Claude Code — teach AI to say "I don't know" instead of guessing.

instantX-researchinstantX-research
71/ 100

公开评测 · 综合采用结论

核心能力可用,建议在受控范围内试用

查看评测依据 评测我的项目基于公开项目证据,非安全认证或安装推荐
11stars
1forks
最近更新 5个月前
评测生成时间(北京时间)
本报告引擎
v3.10.0
当前引擎
v3.16.0

本报告与当前引擎使用不同规则;原分数不会自动更新,不同版本的分数不宜直接对比。

重新评测此项目

进入后确认来源与额度,提交才会创建任务。

Evaluation report

综合采用结论

71
C
满分 100
值得试用低风险
决策摘要

核心能力可用,建议在受控范围内试用

78%
中置信度
75
文档
100
安全
72
质量
52
活跃
14
采用
  • 基础评测完成+25/25确定性评分与静态安全扫描已完成
  • README 有效证据+15/256,077 个去重后的有效字符
  • 独立证据来源+8/202 类非重复证据,重复文件不叠加
  • 仓库元数据+10/10已取得仓库状态与采用数据
  • 活跃记录+5/5已取得最近提交时间
  • AI 复核+15/15已完成结构化 AI 证据复核
How it works · 流程图

防幻觉指南应用流程

文档描述了从安装到应用再到评估的连续步骤,适合用流程图表示。

AI 提取 · 证据约束

左右滑动查看完整图示

防幻觉指南应用流程文档描述了从安装到应用再到评估的连续步骤,适合用流程图表示。加载到上下文根据需求调整检验行为安装指南应用原则定制规则评估效果
图示依据
  • • Install部分提供安装方法
  • • Principles部分列出应用原则
  • • Customization和How to Judge Effectiveness章节
五维表现
针对Claude Code幻觉问题提供实用指南,安装方式多样,示例和定制说明清晰。但缺少实际示例代码和详细参数说明,设计证据有限。
质量证据
  • README中'Why Hallucinations Happen'章节解释原理
  • Install部分提供CLAUDE.md和插件两种安装命令
  • Customization章节建议添加领域特定规则
  • How to Judge Effectiveness列出预期行为
  • File Structure展示文件结构,包含SKILL.md
采用建议
优势
  • 问题与用途描述
  • 有效 README
  • 可执行示例
  • 输出或结果说明
  • 未发现已知高风险模式
关注点
  • 缺少安装或接入步骤
  • 缺少输入、参数或工具说明
  • 缺少可执行示例或代码片段,仅文字说明
  • 未提供输入参数或工具的具体说明
  • 设计部分证据不足,如权限、错误处理等未提及
适合

希望减少Claude Code幻觉的开发者、需要全局或项目级防幻觉指南的团队、偏好CLAUDE.md或插件方式集成的用户、需要定制化防幻觉规则的项目

不建议直接用于

需要严格技术实现细节的开发者、非Claude Code环境或需要独立Skill的场景

也有自己的公开项目?先看完证据,再用当前规则生成独立报告。

评测我的项目 →
文档证据
75/100
问题与用途描述10 分
有效 README12 分
安装或接入步骤14 分
可执行示例16 分
输入、参数或工具说明11 分
输出或结果说明9 分
限制、权限或边界12 分
错误处理或排障8 分
许可证信息5 分
结构化章节3 分
安全证据
低风险
未发现已知高风险模式

静态扫描不是安全保证,生产接入前仍应人工复核权限和数据边界。

优先改进清单
  1. 01补充安装或接入步骤
  2. 02补充输入、参数或工具说明
方法、证据与局限展开
数据来源

GitHub Repository API

扫描范围

2 个文件 · 11,135 字符

评测引擎

v3.10.0 · AI 复核已启用(deepseek-chat)

局限
  • 静态评测不会安装或执行项目代码
  • 安全扫描基于高信号文件与已知模式,不能替代人工审计
  • 流行度只反映采用程度,不代表安全或工程质量

30 天热度趋势

README

Anti-Hallucinate Skills for Claude Code

"Hallucinations are hard to anticipate, hard to catch, and the wrong answer often looks exactly like it could be the right one." — Anthropic

A set of behavioral guidelines to reduce AI hallucinations in Claude Code, derived from Anthropic's official explanation of why AI models hallucinate and practical tactics to catch and prevent it.

Why Hallucinations Happen

AI models work by predicting the next word based on patterns learned from massive text datasets — similar to how your phone suggests the next word as you type, but at a much larger scale. When asked about obscure or niche topics, there isn't enough training data to draw from — so the model takes a guess. Combined with the training objective to be helpful, the model would rather produce a plausible-sounding answer than admit ignorance. Think of it like a well-read friend who takes pride in knowing everything — they'd sometimes rather say something confidently wrong than say "I don't know."

The key insight: being honest IS being helpful — they reinforce each other, not compete. Saying "I don't know" is not a failure — it's the more helpful response.

Anthropic regularly tests Claude with thousands of trick questions — obscure facts, niche topics, questions where the truthful answer is "I don't know" — and measures how often it correctly expresses uncertainty, whether it fabricates citations, and how often it hedges appropriately vs. states something false with confidence. Each version improves, but this remains an ongoing challenge for the entire AI field, not a solved problem. And as hallucinations become rarer, users check less — making the remaining errors more dangerous, not less.

Principles (for Claude)

  1. Honesty Over Helpfulness — Say "I don't know" instead of guessing
  2. Source Verification — Never fabricate citations, papers, or statistics
  3. Confidence Calibration — Hedge when uncertain, don't fake certainty
  4. High-Risk Awareness — Be extra careful with facts, dates, names, niche topics
  5. Self-Checking — Pause and ask: "Do I actually know this?"

Prompt Tactics (for Users)

These are strategies you can use in your prompts to reduce hallucinations:

Ask for Sources and Verify Them

Tell the AI to back up its claims with sources. If it already gave sources, ask it to check that those sources actually support what it's saying — not just that they exist.

Give Permission to Not Know

Tell the AI upfront: "It's okay if you don't know." This reduces the pressure to guess and makes honest responses more likely.

Probe Confidence

Ask the AI: "How confident are you? Is anything in your answer potentially wrong?" Often the AI knows it's uncertain but presented the answer confidently anyway.

Use a Fresh Chat for Verification

Start a new conversation and ask the AI to find errors in the previous answer and confirm that sources support the statements. A fresh context avoids confirmation bias.

Cross-Reference Critical Claims

For important work, never rely on AI alone. Be skeptical of specific numbers, dates, and citations — cross-check them against trusted primary sources.

Ask Follow-Up Questions

If something in the AI's response sounds off or feels too convenient, don't let it pass — ask follow-up questions to probe the claim. Hallucinations often unravel under targeted questioning.

Install

Option A: CLAUDE.md (Recommended)

Because anti-hallucination is a guardrail that should apply to every factual claim, CLAUDE.md is preferred — it's always loaded into context, so there's no risk of the skill failing to trigger at the moment you need it most.

A1: Global — one install, all projects (Recommended)

Append to your user-level ~/.claude/CLAUDE.md, which Claude Code loads for every project automatically:

mkdir -p ~/.claude
echo "" >> ~/.claude/CLAUDE.md
curl https://raw.githubusercontent.com/instantX-research/anthropic-anti-hallucinate-skills/main/CLAUDE.md >> ~/.claude/CLAUDE.md

A2: Per-project — share with your team via git

Install into the project root so teammates pulling the repo pick up the same guardrails automatically.

New project:

curl -o CLAUDE.md https://raw.githubusercontent.com/instantX-research/anthropic-anti-hallucinate-skills/main/CLAUDE.md

Existing project (append):

echo "" >> CLAUDE.md
curl https://raw.githubusercontent.com/instantX-research/anthropic-anti-hallucinate-skills/main/CLAUDE.md >> CLAUDE.md

Option B: Claude Code Plugin

From within Claude Code, first add the marketplace:

/plugin marketplace add instantX-research/anthropic-anti-hallucinate-skills

Then install the plugin:

/plugin install anti-hallucinate@anthropic-anti-hallucinate-skills

And reload plugins to apply:

/reload-plugins

This installs the guidelines as a Claude Code plugin, making the skill available across all your projects. Note that skills are loaded on-demand based on the skill's description — which means there's a chance Claude won't recognize the current context as high-risk and skip loading it.

File Structure

.
├── CLAUDE.md                              # Core guidelines (drop into any project)
├── EXAMPLES.md                            # Good vs bad examples for each principle
├── skills/
│   └── anti-hallucinate/
│       └── SKILL.md                       # Claude Code skill definition
├── LICENSE
└── README.md

Customization

These guidelines are a starting point. You can:

  • Add domain-specific rules — e.g., "Never guess medication dosages" for medical projects
  • Adjust strictness — tighten for research work, relax for creative brainstorming
  • Combine with other skills — these guidelines complement coding-style skills like andrej-karpathy-skills

How to Judge Effectiveness

After applying these guidelines, Claude should:

  • Say "I don't know" or "I'm not sure" more often (this is a feature, not a bug)
  • Provide fewer fabricated citations and statistics
  • Clearly distinguish facts from inferences
  • Proactively suggest verification steps for high-risk claims
  • Self-correct when challenged rather than doubling down

Acknowledgments

License

MIT