BHUSA-Anthropic-CyberSecurity-Skills
- 评测生成时间(北京时间)
- 本报告引擎
- v3.10.0
- 当前引擎
- v3.16.0
本报告与当前引擎使用不同规则;原分数不会自动更新,不同版本的分数不宜直接对比。
进入后确认来源与额度,提交才会创建任务。
综合采用结论
存在需要人工复核的风险或证据不足
- 基础评测完成+25/25确定性评分与静态安全扫描已完成
- README 有效证据+23/259,230 个去重后的有效字符
- 独立证据来源+8/202 类非重复证据,重复文件不叠加
- 仓库元数据+10/10已取得仓库状态与采用数据
- 活跃记录+5/5已取得最近提交时间
- AI 复核+15/15已完成结构化 AI 证据复核
Arsenal Lab场景执行流程
README描述了从准备到执行的连续步骤,包括安装、验证、运行场景,适合用流程图表示。
左右滑动查看完整图示
- • Day-before setup: 安装工具,运行setup-skills.sh,复制文件夹,运行verify.sh
- • The 30-minute repeated block: 运行一个场景,交替S1/S2
- • Advanced tier: 可选,使用generate-bigdata.py和bigquery.py
- README列出两个场景:TollBooth(DFIR)与OpenDoor(审计),共享Acme Rentals故事
- 提供setup-skills.sh、verify.sh等脚本,verify.sh预期9/9 PASS
- 高级层:generate-bigdata.py --events 1000000 --days 3生成大数据集,bigquery.py查询
- 安全声明:所有数据合成,无真实系统,使用RFC5737 IP与EXAMPLE密钥
- 已知缺口:无IMDS/SSRF专用技能,需绕行
- 有效 README
- 输入、参数或工具说明
- 限制、权限或边界
- 结构化章节
- 两个场景共享同一故事线,覆盖DFIR与审计,目标用户明确
- 发现提示词或系统信息提取意图
- 缺少问题与用途描述
- 缺少安装或接入步骤
- 缺少可执行示例
- 缺少错误处理或排障章节,遇到问题无指引
Black Hat Arsenal Lab现场演示、网络安全技能教学与实操、DFIR与云审计场景演练、展示AI代理在安全分析中的应用
生产环境真实安全事件响应、无现场网络或模型访问的离线环境、需要完整技能覆盖的通用安全分析
也有自己的公开项目?先看完证据,再用当前规则生成独立报告。
评测我的项目 →prompt-extractionsecurity/SECURITY.md:43high confidence`ANTHROPIC_BASE_URL` and exfiltrate your key *before the trust prompt*修复:移除提示词提取逻辑,并增加敏感上下文不可输出的边界说明。
- 01移除提示词提取逻辑,并增加敏感上下文不可输出的边界说明。
- 02补充问题与用途描述
- 03补充安装或接入步骤
- 04补充可执行示例
方法、证据与局限展开收起
GitHub Repository API
2 个文件 · 14,434 字符
v3.10.0 · AI 复核已启用(deepseek-chat)
- 静态评测不会安装或执行项目代码
- 安全扫描基于高信号文件与已知模式,不能替代人工审计
- 流行度只反映采用程度,不代表安全或工程质量
30 天热度趋势
README
Resources
RESOURCES.md — repo, frameworks, per-technique ATT&CK links, agent tooling, AWS hardening, and DFIR references.
Fielding questions at the booth
BOOTH-QA.md (presenter-only) has tight, honest answers for the off-scenario questions that actually bite:
affiliation, mapping validation, "is this just your company", safety, and "does it do ". Read it before the floor opens.
Running it on the floor (read ops/ first)
- Decide the posture —
ops/POSTURE.txt. Instructor-live+participant-static (simple, recommended) vs participant-live via gateway (richer, more to stand up). Pick before the show. - Day-before —
ops/preflight.shtells you which model path this network allows (cloud / local / static) and checks for stray commercial branding. There is no local-model path: this kit is API-only by design (see ops/POSTURE.txt). - Key safety — the real Anthropic key lives ONLY on the gateway laptop (
ops/gateway-up.sh+ops/gateway-litellm.yaml). Participant laptops useops/station.sh— a gateway token, never the key. - Model is pinned to a non-Fable-5 model to avoid the safeguard model-swap seen last run.
- Block-never-dies — if cloud AND local are down,
backstop/flagship-lsass.txtis the canned analysis. CLAUDE.mdauto-loads a defensive-analyst framing (cuts refusals + skill-misrouting).- Send
ops/arsenal-ops-email.txt— it unblocks all of the above (network + day-before slot).
Playable data the agent runs skills on
datasets/ holds six synthetic, analysis-only artifacts across domains — credential-dumping (Sysmon),
Kerberos 4769, NetFlow, domain-fronting proxy logs, a sandbox report, and OSINT indicators — each mapped to a real
loaded skill and threaded to the same attacker as the main scenarios. Prompts + answer key: SECURITY-OPS-COVERAGE.md.
Any security operation? Start here
SECURITY-OPS-COVERAGE.md routes any visitor — DFIR, hunting, detection, web, cloud, GRC, CVE triage — to a
copy-paste starter prompt + the skill families to use. Two rows are playable now (the scenarios); two more use
the analysis-only bonus logs in samples/; the rest are bring-your-own-artifact. Defensive/analysis only by design.
Arsenal Lab hands-on kit — Cloud + Network, driven by AI agent skills
Two self-contained scenarios for Black Hat USA 2026 Arsenal Lab (#53649). Promotes the
open-source Anthropic-Cybersecurity-Skills library. Not affiliated with Anthropic PBC. Apache-2.0.
The two scenarios (one shared universe: "Acme Rentals")
- Scenario 1 — TollBooth (reactive DFIR). Investigate a breach: find an SSRF -> EC2
metadata (IMDS) credential leak in
lab-tollbooth.pcap, pivot on the leaked AccessKeyId intocloudtrail/, and prove IAM enumeration + bulk S3 exfil by the same key from the same attacker IP. Map to MITRE ATT&CK + D3FEND. Skills: pcap/Wireshark + CloudTrail/S3-exfil. - Scenario 2 — OpenDoor (proactive audit). Before that breach: audit the cloud config in
opendoor/and find the 3 holes the attacker used — public S3 bucket, an IAM privilege-escalation path, and IMDSv1 enabled. Map to CIS. Skills: S3/IAM/CIS auditing + privesc assessment. They are the same story from two ends: OpenDoor is the audit that would have prevented TollBooth.
The 30-minute repeated block
<=5 min slides + <=5 min instructor demo + ~20 min hands-on.
The crowd rotates, so most attendees see ONE block. Run one scenario per block and alternate
S1 / S2 across the day, or let each participant pick. Both share the same start/reset/verify.
S1 is the stronger marquee demo; S2 is lighter-weight (pure jq on JSON) and a good fallback if
model access or the pcap tooling misbehaves.
What the participant does
Drives the AI agent (Claude Code, loaded with the skills); the agent runs the tools. Every step
also has a raw tshark/jq command on the cheat sheet, so anyone without a working agent still finishes.
Files
lab-tollbooth.pcap.............. S1 capture (SSRF + IMDS + leaked creds + benign noise)cloudtrail/*.json.............. S1 CloudTrail (12 malicious events in benign noise)opendoor/*.json............... S2 config artifacts (S3 policy, public-access-block, IAM policy, IMDS options)Arsenal-CheatSheet-Book.pdf... 6-page book, print 1 per laptop (pp.1-3 = S1, pp.4-6 = S2; each: instructions/hints/answers)start.sh...................... participant start + both missionsreset.sh...................... restore S1+S2 data + clear agent chat between attendees (<10s)verify.sh..................... instructor self-test (real tshark/jq; 9 checks; prints PASS/FAIL)setup-skills.sh............... day-before: clone repo, symlink the 10 needed skillsgenerate-bigdata.py.......... ADVANCED tier: build a large synthetic CloudTrail corpus at any scalebigquery.py.................. ADVANCED tier: query that corpus offline with DuckDB (pivot / anomaly)demo-one-shot.md............. INSTRUCTOR demo: one prompt that composes many skills (not the worksheet)
Day-before setup (build one Kali image, clone to all laptops)
apt-get install -y tshark jq(tcpdump, python3-scapy usually present on Kali).- Install Claude Code; pre-authenticate with the shared key (see model access).
./setup-skills.shto mount skills into~/.claude/skills.- Copy this folder to
~/tollbooth. Run./verify.sh— expect 9/9 PASS. - Print
Arsenal-CheatSheet-Book.pdf(color; answer pages are red) — one per laptop. - (Optional advanced tier)
pip install duckdband pre-build the big corpus (see below) if you want the haystack demo.
Advanced tier — folded into the scenarios (optional)
The big dataset is NOT a separate block; it extends the two scenarios:
- Scenario 1, Step 6 — the same attack hidden in ~1M CloudTrail events (scale).
- Scenario 2, Step 5 — cross-reference the CloudTrail logs to prove which audit finding was exploited (3), probed (2), or bypassed (1). Uses the existing logs, no new data.
The two scenarios above use small, tailored data so a beginner finishes in 20 min. For fidelity (and a stronger demo), the advanced tier points the SAME technique at a real haystack:
pip install duckdbthenpython3 generate-bigdata.py --events 1000000 --days 3builds ~1M realistic CloudTrail events, partitioned by hour, gzip'd. 400k = ~8.6 MB / ~24 s; 1M is a few tens of MB. The TollBooth needle (same leaked key + attacker IP) is embedded.python3 bigquery.py --buildindexes it once intoev.duckdb.python3 bigquery.py --key ASIAJ7A6EXAMPLEK3Y99-> the 12-event needle, found in seconds.python3 bigquery.py --anomaly-> finds it WITHOUT the key (external IP + IAM enum + GetObject burst). Why this scales on a laptop: the agent/skill must QUERY (DuckDB), never READ the corpus. Query results are tiny, so token cost and memory stay flat no matter how big the data is. Steer the agent to run DuckDB overbigcloud/, not to ingest files. Sweet spot 300k-500k; 1M+ only on good hardware. This tier is the payoff line: a human can't triage 1M events in 20 min; the agent writes one query.
One-shot setup (day-before, per laptop)
./setup.sh installs tooling (jq, tcpdump, tshark), clones + links the agent skills, resets the
data, and runs verify.sh. Idempotent. Flags: --all-skills (link the whole library), --gen-events N
(build the big corpus), --no-apt (offline / no sudo). RUN ON ONE LAPTOP FIRST, confirm 9/9, then the rest.
It cannot provision model access or test the safeguard — it reports those and leaves them to you.
Prerequisites on each laptop (day-before)
The agent path needs nothing extra, but the RAW FALLBACK commands and verify.sh need:
sudo apt-get update && sudo apt-get install -y jq tcpdump (tshark is usually already on Kali).
setup-skills.sh now does this. Run ./verify.sh first on every laptop - it preflights the tools.
Model access (open decision — pick before the show)
- Primary: Claude Code pre-logged-in with a shared, rate-limited key.
- Offline fallback: local Ollama model wired into Claude Code via
ANTHROPIC_BASE_URL(verify it emits tool calls, or it won't run commands). - Zero-model dry-run: read the answer page, run the printed commands. Always works.
start.shauto-detects the mode.
Safety / legal
All data synthetic and self-contained; no live system is touched. Fictional keys/IPs/company (RFC5737 doc IPs; AWS "EXAMPLE" key format). Answer pages carry the authorized-use banner.
Known repo gaps (state openly — a "contribute a skill" call to action)
No dedicated skill for: IMDS/SSRF-to-metadata (inside exploiting-server-side-request-forgery),
VPC Flow Log analysis (only NetFlow), or security-group/NACL audit (only inside
auditing-cloud-with-cis-benchmarks). Both scenarios route around these.
Presenter must still confirm before the show
- Exact skill directory names vs. the live repo
index.json(set changed 754 -> 817). - Exact D3FEND IDs (d3fend.mitre.org) and CIS numbers (your benchmark version).
- That
tshark/jqexist on the supplied Kali build and the pcap parses there (runverify.sh). - Model access on the laptops (the one true showstopper).
- Whether any online sandbox is allowed on the Arsenal network (only for optional deeper tracks) — confirm with Arsenal operations.