跳到主要内容
Supermarket
返回能力市场
MCP Server
data
MIT

geo-score

Can AI engines cite your site, and do they? Free 0–100 readiness score on an open GEO rubric, plus citation tracking via the OpenAI, Perplexity, Gemini and Claude APIs with your own keys. Zero dependencies, MCP server.

jianruntechjianruntech
92/ 100

公开评测 · 综合采用结论

证据充分,整体质量与安全表现优秀

查看评测依据 评测我的项目基于公开项目证据,非安全认证或安装推荐
618stars
12forks
最近更新 2天前
评测生成时间(北京时间)
本报告引擎
v3.16.0
当前引擎
v3.16.0

规则版本一致,但报告只反映生成时的证据,不代表项目代码和安全状态始终不变。

重新评测此项目

进入后确认来源与额度,提交才会创建任务。

Evaluation report

综合采用结论

92
A+
满分 100
值得推荐低风险
决策摘要

证据充分,整体质量与安全表现优秀

100%
高置信度
100
文档
100
安全
94
质量
100
活跃
42
采用
  • 基础评测完成+25/25确定性评分与静态安全扫描已完成
  • README 有效证据+25/2533,121 个去重后的有效字符
  • 独立证据来源+20/205 类非重复证据,重复文件不叠加
  • 仓库元数据+10/10已取得仓库状态与采用数据
  • 活跃记录+5/5已取得最近提交时间
  • AI 复核+15/15已完成结构化 AI 证据复核
How it works · 流程图

geo-score 三级测量流程

README与SKILL.md描述了从免费评分到API引用查询再到周度追踪的连续步骤,且明确引用结果不并入分数

AI 提取 · 证据约束

左右滑动查看完整图示

geo-score 三级测量流程README与SKILL.md描述了从免费评分到API引用查询再到周度追踪的连续步骤,且明确引用结果不并入分数21项检查打分可选升级固定问题集Readiness 0-100引用率另列不计入采样8个URL输入Level1 评分免费Level2 单问引用需密钥Level3 周度追踪需密钥报告与分数输出
图示依据
  • • README「Three levels, one tool」:Score→Ask→Watch 三级及各自所需条件
  • • SKILL.md「Reporting rules」:Never add a citation result to the readiness score,引用结果与分数并列报告
五维表现
解决AI引用可见性这一真实问题,三级能力(免费评分/API引用查询/周度追踪)与21项检查、公开rubric均有证据;最大缺口是Level 2/3需自备API密钥且需同时下载geo_watch.py,离线或无私钥用户只能用到Level 1。
质量证据
  • 通过: Agent Skills 格式校验 1/1 个通过
  • 通过: 1 个有效 Skill 含可核实的指令步骤或示例
  • README「Three levels, one tool」表:Level 1 `geo_score.py stripe.com` 无需密钥,Level 2 `--ask` 与 Level 3 `watch run` 需自有API密钥
  • SKILL.md「The rubric」:Readiness 100分=Reachable 15+Understandable 22+Content Citability 35+Brand Credibility 18+Answer Fit 10,另有+6 bonus不计入分母
  • SKILL.md「How to run an audit」:固定采样8个URL(首页、2产品页、2文档页、3近期内容页),必须跟随重定向、按状态码判断存在性、g.ssr不执行JS
  • README「Stability, privacy and install options」:无遥测,Level 1仅抓取指定站点及其标记/robots指向的URL与Wikidata/Wikipedia,密钥从不写入文件
  • README「Why did my score move 4 points?」:每次最多采样8页,实测波动±5,需用 `--urls-from last.json` 复现同一批页面
采用建议
优势
  • 问题与用途描述
  • 有效 README
  • 安装或接入步骤
  • 可执行示例
  • 通过: Agent Skills 格式校验 1/1 个通过
关注点
  • Level 2/3需用户自备OpenAI/Perplexity/Gemini/Anthropic/OpenRouter密钥,无密钥用户无法验证引用能力
  • Level 2/3需cli/geo_watch.py与geo_score.py并存,单文件curl方式仅覆盖Level 1
  • 评分存在采样波动(每次最多8页,±5分),跨版本v1.0与v1.1分数不可比
  • 明确不做修复(不生成robots.txt/JSON-LD/llms.txt模板),需要整改的用户须另寻工具
适合

想快速了解站点对AI检索爬虫可读性与可引用性的站长、已持有AI厂商API密钥、想实测并周度追踪引用率的团队、需要可审计、版本化评分规范并自行实现的研究者或工程团队、在CI中对站点AIV分数做回归门禁的开发者

不建议直接用于

希望工具直接生成修复模板或改写内容的用户、无任何AI厂商API密钥且需要引用实测的用户、需要单次查询即得出统计结论的用户

也有自己的公开项目?先看完证据,再用当前规则生成独立报告。

评测我的项目 →
文档证据
100/100
问题与用途描述10 分
有效 README12 分
安装或接入步骤14 分
可执行示例16 分
输入、参数或工具说明11 分
输出或结果说明9 分
限制、权限或边界12 分
错误处理或排障8 分
许可证信息5 分
结构化章节3 分
安全证据
低风险
未发现已知高风险模式

静态扫描不是安全保证,生产接入前仍应人工复核权限和数据边界。

方法、证据与局限展开
数据来源

GitHub Repository API

扫描范围

5 个文件 · 59,084 字符

评测引擎

v3.16.0 · AI 复核已启用(deepseek-flash)

局限
  • 静态评测不会安装或执行项目代码
  • 安全扫描基于高信号文件与已知模式,不能替代人工审计
  • 流行度只反映采用程度,不代表安全或工程质量

30 天热度趋势

README

English · 简体中文

geo-score — an open rubric for AI answer-engine visibility

geo-score

Can AI engines cite your site? A free 0–100 score in 20 seconds. Do they? Check through their APIs with your own keys.

License: MIT Rubric v1.1 No dependencies Claude Code Skill MCP server PRs welcome AIV readiness

Score (free, no key) → Ask one question (your API keys) → Watch a question set every week (your API keys). Three levels, one tool · the citation half is never added to the score.

curl -sL https://raw.githubusercontent.com/jianruntech/geo-score/v1.4.0/cli/geo_score.py \
  | python3 - stripe.com --brief

geo-score scoring stripe.com from the command line — 71 out of 100, band Solid

Recorded with geo-score 1.1.0 on 2026-09-09. A run with 1.2.0 on 2026-09-27 read 72.

One command. About twenty seconds. Every check, and what the next tier needs.

Full output — every check, its evidence, and what the next tier asks for
  AIV READINESS  https://stripe.com
──────────────────────────────────────────────────────────────────────────
  71 / 100   Solid
  12 points to Leading

  Reachable 11/15
   ◐ Crawlers allowed in robots.txt   ███████████░░░░░░░  3/5
   ✓ Reachable to retrieval agents    ██████████████████  5/5
   ◐ Main content server-rendered     ███████████░░░░░░░  3/5
  Understandable 15/22
   ◐ Sitemap discoverable and fresh   █████████░░░░░░░░░  2/4
   ✓ llms.txt present and structured  ██████████████████  5/5
   ✓ Organization + WebSite schema    ██████████████████  6/6
   ✗ BreadcrumbList on nested pages   ░░░░░░░░░░░░░░░░░░  0/3
   ◐ Page-type schema (Product, FAQ…) █████████░░░░░░░░░  2/4
  Content Citability 25/35
   ✓ Self-contained answer passages   ██████████████████  9/9
   ◐ Headings match how people ask    ████████░░░░░░░░░░  3/7
   ◐ Freshness signal present         █████████░░░░░░░░░  3/6
   ✓ Statistics carry a source        ██████████████████  7/7
   ◐ Named, verifiable authorship     █████████░░░░░░░░░  3/6
  Brand Credibility 8/10
   ⊘ Third-party listings             ··················   —
   ⊘ Independent mentions             ··················   —
   ✓ Knowledge-graph entity           ██████████████████  4/4
   ◐ sameAs links resolve             ████████████░░░░░░  2/3
   ◐ Video and multimodal presence    ████████████░░░░░░  2/3
  Answer Fit 2/4
   ◐ Content shaped for extraction    █████████░░░░░░░░░  2/4
   ⊘ Covers the questions people ask  ··················   —
   ⊘ Chinese engine readiness         ··················   —
  Biggest gaps
   +4   Headings match how people ask    about half do
   +3   Named, verifiable authorship     and the name links to a verifiable identity page
   +3   Freshness signal present         most pages do, and dateModified agrees with the visible date

  Scored 61 / 86 observable · 4 checks left the denominator · rubric v1.1
  Needs judgement: p3.listings, p3.mentions, p4.question-coverage, p4.cn-engines

  Full rubric and what each tier means:
  https://github.com/jianruntech/geo-score

Python 3.8+, standard library only, nothing to install. It reads public URLs and prints a score against a published, versioned rubric — not a black box.

See how 317 well-known sites score → · a quarter of them are unreadable to AI crawlers.

GEO means Generative Engine Optimization — getting cited by ChatGPT, Perplexity, Google AI Overviews, Gemini and Copilot. Nothing to do with geography or maps.


MCP server

geo-score-mcp is a local MCP server (stdio, Python standard library only). It gives an agent the same three levels as the CLI: score a site for free, ask the AI engines one question, and track a question set week over week.

Add geo-score to Cursor Install geo-score in VS Code Install geo-score in VS Code Insiders Add geo-score to LM Studio

The buttons and the lines below install geo-score 1.4.0 with uvx, which needs uv. They add no keys of their own, but a server started from a shell that exports a provider key (Claude Code passes its environment on) can use it. For a server that cannot spend, add --read-only after geo-score-mcp.

Claude Code

claude mcp add --scope user geo-score -- uvx --from git+https://github.com/jianruntech/geo-score@v1.4.0 geo-score-mcp

Standard config (no keys in the file) for any client that reads an mcpServers file:

{
  "mcpServers": {
    "geo-score": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/jianruntech/geo-score@v1.4.0", "geo-score-mcp"]
    }
  }
}

Claude Desktop does not read your shell's PATH: put the full path that which uvx prints in command.

ToolLevelCostWritesHintsArgs
score_site1FreeNoread-only, open-worldurl, urls, sample
ask2Paid (your keys)Nonon-destructive, open-worldquestion, engines, brand, domains
run3Paid (your keys); dry_run is freeA run file under .geo-score/watch/runs/, next to the confignon-destructive, open-worlddry_run, engines, limit, max_calls, budget_usd
list_runs3FreeNoread-only, idempotent—
report3FreeNoread-only, idempotentrun, format
diff3FreeNoread-only, idempotentfrom, to, format
status—FreeNoread-only, idempotent—

Resources: geo-score://rubric/v1.1.

Prompts: audit_site (url), check_citations (question, brand), weekly_watch, compare_runs (from, to), explain_check (check_id).

Bold arguments are required. Hints are the tools' MCP annotations: clients may use them to decide what to confirm, and they are hints, not guarantees. Only ask and run can spend money, and only with the provider keys you give the server.

Things to ask your agent:

PromptWhat it calls
Score https://acme.com for AI-search readiness and list the 3 checks with the most points to gain.score_site with url
Dry-run my tracked questions and tell me the planned calls and caps.run with dry_run: true (level 3, needs a config)
Compare my last two watch runs: which engines changed, and is any change outside the noise?diff (defaults to the latest two runs)

Keys and config. score_site needs no key. ask and run read the provider keys (OPENAI_API_KEY, PERPLEXITY_API_KEY, GEMINI_API_KEY, ANTHROPIC_API_KEY, OPENROUTER_API_KEY) from the server's environment, which is not your terminal's: a desktop app does not see what you exported in a shell. The level 3 tools also need a geo-score-watch.json. Pass it as -c /absolute/path/geo-score-watch.json, because a GUI client does not start the server in your project. From 1.4.0, --read-only gives a server that cannot spend anything.

Setup for each client (Claude Code, Claude Desktop, Codex, Cursor, VS Code, Gemini CLI, Devin Desktop, Zed), the Claude Code plugin, timeouts, security and troubleshooting: guide/mcp.md.

Three levels, one tool

LevelWhat it answersNeedsCommand
1 · ScoreCan AI engines reach, parse, trust and cite the site? 0–100 against the open rubricNothing: no key, no installgeo_score.py stripe.com
2 · AskRight now, for a question you care about, do they cite it?An API key for any of OpenAI, Perplexity, Gemini, Anthropic, OpenRoutergeo_score.py stripe.com --ask "best payments API for marketplaces"
3 · WatchHow often are you cited, against which competitors and sources, week over week?Your API keys and a fixed question listgeo_score.py watch run

Level 1 is the score. Levels 2 and 3 measure the outcome the rubric keeps out of the score on purpose (two scores, never one): they are reported next to the 100 and never summed into it. A site can score 90 and still lose every answer to a competitor with more third-party coverage, and the reverse also happens, which is why you want both.

Levels 2 and 3 need cli/geo_watch.py next to cli/geo_score.py: clone the repo or download both files. The one-line curl | python3 above runs level 1.

Why this is a different question from SEO

Classic SEO asks where do I rank. Answer engines don't rank — they retrieve passages, decide whether a source is worth quoting, and cite it. Different question, different failure modes: a site can sit at position 3 on Google and never be quoted, while a page nobody links to gets cited daily because its passages are clean.

Most of what determines this is mechanical and cheap to fix — a robots.txt line, a JSON-LD block, a date in a template, a paragraph rewritten so it stands on its own. The hard part is knowing which of them you are missing, and what each one is worth.

What it checks

21 tiered checks totalling 100 points, plus 4 bonus checks worth up to +6 outside the denominator. Full specification: rubric/v1.1.md · 简体中文

PillarPtsAsks
Reachable — gates15Can a retrieval crawler get the page at all? robots.txt, live reachability across 10 AI user-agents, server-rendered content
Understandable22Can it tell what the page and the company are? Organization + WebSite, llms.txt, sitemap, breadcrumbs, page-type schema
Content Citability35Is there anything here worth quoting? Self-contained answer passages, headings that match how people ask, sourced figures, real bylines, freshness
Brand Credibility18Why should an engine trust it? Knowledge-graph entity, third-party listings, sameAs that resolves, video presence
Answer Fit10Is the content shaped to be lifted into an answer?

Content Citability carries the most weight on purpose: answer engines retrieve passages, not domains. Passage shape beats domain authority more often than classic SEO intuition expects.

Every scored check is tiered — 2 to 4 tiers, each naming a count out of the 8 sampled pages, so two people scoring the same site agree on the arithmetic. Three checks are gates: score zero on crawler access, live reachability or server-rendered content and the result caps at 40, because until a crawler can reach the content nothing else you change has any effect.

Bands

0–3031–5051–6566–8283–100
Not startedEarlyGrowingSolidLeading

Band names describe a stage, not a verdict. External benchmarks put most business sites in the 30–55 range, so a score in the forties is ordinary, not alarming.

317 sites, scored in public

A quarter of them are unreadable to AI crawlers. 83 sites have a gate check at zero — an AI retrieval crawler cannot get the content, so it has nothing of theirs to quote. 23 block AI crawlers by name in robots.txt, which is an editorial choice and reported as such — amazon.com lands at 12 for exactly this reason. 47 serve a page whose body only exists after JavaScript runs. Their content is there, a browser sees it, and a crawler gets an empty shell. That group almost certainly did not choose it. A further 13 hand a crawler an outright error.

Median 56 (95% CI 52–58, stats.py). Range 11 to 99.

SiteScoreBand
pulumi.com99Leading
resend.com99Leading
lumalabs.ai94Leading
kayak.com92Leading
openrouter.ai90Leading
…
amazon.com12Not started
mercadolibre.com12Not started
keepa.com11Not started

The full table, by sector → · markdown · raw data · every site's full report · re-run it

Two more findings worth the click. Sites built for the Chinese market score 20 points lower than everyone else (95% CI 16–28; median 39 against 59; earlier samples measured with 1.1.0 put the gap between 16 and 23 points). The largest per-check differences are two heuristics, question-shaped headings and self-contained answer passages, which the CLI also read lower than a human on the one Chinese site in the hand-audit comparison, so part of the gap may be the tool reading Chinese pages conservatively. And a named, verifiable byline and an opening paragraph that stands on its own are among the three largest gaps on more than half the sites.

Every number here is reproducible with the command at the top of this page (the benchmark was run with 1.3.0 on 2026-09-27) — and we measured how reproducible. Running the whole benchmark twice with the same tool and comparing every site: 96% land within ±5, 45% land identically. Read one site's score as ±5 rather than as exact; medians are stable. The unstable part is the gate checks, where five sites flipped between runs because their bot protection answered a crawler differently. The band, the control experiment and the per-site pairs are in benchmark/REPRODUCIBILITY.md.

For five reference sites we also publish hand-scored audits covering all 21 checks, with the evidence behind each one: examples/audits/v1.1/.

Ways to run it

CLI — level 1 is one file with no dependencies, 20 seconds. Levels 2 and 3 add cli/geo_watch.py from the same release.

python3 cli/geo_score.py example.com            # human-readable
python3 cli/geo_score.py example.com --explain  # with the evidence behind every check
python3 cli/geo_score.py example.com --json     # conforms to schema/report.v2.json
python3 cli/geo_score.py example.com --compare competitor.com   # side by side
python3 cli/geo_score.py example.com --badge aiv-badge.svg      # embeddable SVG
python3 cli/geo_score.py example.com --share                    # one line to paste somewhere

GitHub Action (level 1) — score on every push, fail the build when it regresses. For level 3 on a schedule, see examples/ci/watch-weekly.yml.

- uses: jianruntech/geo-score@v1
  with:
    url: https://example.com
    fail-under: 40

Claude Code skill — the CLI measures what a static fetch can see. Four checks need off-site search or human judgement, and the skill does those too.

git clone https://github.com/jianruntech/geo-score ~/.claude/skills/geo-score
# then: /geo-score audit https://example.com

The CLI leaves those four checks out of the denominator rather than guessing. Re-scoring the five published hand audits on the same pages, it matched the auditor's tier on 79% of 86 check pairs (95% where it reads a rule, 64% where it approximates a judgement) and read 3.6 points lower on average, from 10 lower to 6 higher. n=5 is small: VALIDITY.md lists every disagreement.

MCP server (all three levels) — geo-score-mcp, or python3 cli/geo_score.py mcp from a clone, for Claude Code, Cursor and other agents: see MCP server.

Levels 2 and 3 · does AI actually cite you?

--ask and watch put the questions your buyers ask to ChatGPT, Perplexity, Gemini and Claude, through each provider's search-enabled API with your own keys, and record who the answers cite: you, your competitors, or the third-party pages (forums, review sites) the engines lean on instead. Keys are read from the environment; geo-score never writes them anywhere.

git clone https://github.com/jianruntech/geo-score && cd geo-score/cli
export OPENAI_API_KEY=… PERPLEXITY_API_KEY=…         # any subset of engines works
python3 geo_score.py acme.com --ask "best invoicing app for freelancers"   # level 2
python3 geo_score.py watch init --brand Acme --domain acme.com --competitor "Rival=rival.com"
# put the questions your buyers ask an AI assistant into queries.csv, then:
python3 geo_score.py watch run --dry-run            # the plan and the caps; no calls, no cost
python3 geo_score.py watch run                      # level 3: every question on every engine, saved
python3 geo_score.py watch diff                     # this run against the last, with a significance test
What a watch run prints (illustrative: made-up brands and canned answers, not a real measurement)
geo-score watch · Lumo · run 20260926T083000Z · api channel

6 questions × 2 engines × 2 = 24 planned · 23 answered · 1 failed · 0 skipped

Cited in 38% of answers to questions that do not name Lumo (6 of 16, 4 questions, 95% CI 12–62%) · mentioned in 38%
Questions that name Lumo (2, kept out of the headline): cited in 100% (7 of 7, 95% CI 65–100%) · mentioned in 100%
Your site was cited somewhere for 5 of 6 questions.

By engine
  engine          model             cited         95% CI   mentioned  avg rank
  chatgpt-api     gpt-6-luna        9/12 75%      42–100%  75%        1.0
  perplexity-api  sonar             4/11 36%      8–75%    36%        2.0       1 failed
  gemini-api      gemini-3.8-flash  not measured                                no_key: set GEMINI_API_KEY or GOOGLE_API_KEY

These rows, and the tables below, count every question, the 2 that name Lumo included.

Share of voice
              cited  95% CI  mentioned
  Lumo (you)  57%    26–86%  57%
  Pixa        52%    33–71%  100%

Sources the engines cite most (not yours)
  domain        answers  share  owner
  pixa.example  12       52%    Pixa
  reddit.com    11       48%

Questions where a competitor is cited and you are not (1)
  id   question                    cited instead
  q03  cheapest text to video app  Pixa

By question type
  type         cited     95% CI   mentioned
  alternative  2/4 50%   15–85%   50%
  list         4/8 50%   22–78%   50%
  pricing      3/7 43%   0–100%   43%
  vs           4/4 100%  51–100%  100%

What the engines searched for (from 12 answers that show it)
  times  search
  2      best free ai video generator
  2      ai video tool with no watermark
  2      cheapest text to video app
  2      lumo vs pixa
  2      pixa alternatives
  2      鹿末视频免费吗

Ledger
24 questions asked · 23 API requests · 14,620 in / 4,860 out tokens · 23 searches · $0.07 + 12 answers with no price (add prices to geo-score-watch.json)
Caps: at most 100 questions.

Read this before quoting the numbers
- API channel. Answers come from each provider's search-enabled API, which is not the consumer app. Compare runs with runs; never pool them with answers sampled by hand in the apps.
- The same question gets different answers from one ask to the next, and answers to one question move together. Rates carry a 95% interval: Wilson when each question was answered once, a bootstrap over questions when a question has several answers. diff pairs the questions both runs answered and calls a change a change only when an exact paired test (McNemar, or a sign-flip test when a question has several answers), Holm-corrected across engines, says so (p < 0.05).
- A question that names the brand invites an answer that cites it. Those questions are kept out of the headline rate and reported on a line of their own.
- Engines without a key, calls that failed and calls skipped by a cap count as not measured, never as zero.
- This measures citations. It does not predict traffic, rankings or revenue.

Level 2 asks each engine each question once and prints the result under the readiness report (and into the report's citation object with --json). One ask is an anecdote: use it to see what the engines say today, not to measure a rate.

Level 3 keeps a fixed question list, saves every answer under .geo-score/watch/runs/ (schema), and reports:

Meaning
citedThe answer links to a URL you own: one of your domains (subdomains included), or a url_prefixes entry such as your Amazon store or GitHub org
rankYour position among the distinct domains the answer cites. Rank 1 means you were the first source
mentionedThe answer names you, one of your aliases or your domain. Chinese, Japanese and Korean names match anywhere; all other names match whole words only
share of voiceThe same two rates for each competitor, over the same answers
sourcesThe third-party domains cited most often: the pages the engines trust in your category
gapsQuestions where a competitor is cited and you are not, in any answer
searchesThe searches the engine actually ran before answering, where the API exposes them (OpenAI, Gemini, Claude). This is the query fan-out, observed rather than guessed
ledgerQuestions asked, API requests, tokens, searches and cost. Every run keeps its own ledger

What it will not tell you.

  • It is not the ChatGPT app. Answers come from each provider's API with web search switched on. The consumer apps use other models, prompts and personalisation. Compare API runs with API runs; never pool them with answers sampled by hand in the apps.
  • One answer is an anecdote. Every rate carries a 95% interval (a bootstrap over questions when a question has several answers, since those answers move together), and diff calls something a change only when an exact paired test on the questions both runs answered, Holm-corrected across engines, says so (p < 0.05). Too few shared questions (under 6, or too few for the number of engines compared) is reported as too few to tell. The method: guide/watch-methodology.md.
  • Questions that name you are kept apart. "acme vs rival" or "is acme worth it" put your name in the engine's search and are cited almost every time. watch flags them when it runs, keeps them out of the headline rate, and reports them on their own line with their own count. An optional branded column in queries.csv corrects the flag by hand.
  • Not measured is not zero. An engine without a key, a failed call and a capped call are all reported as not measured, and none of them lowers your rate.
  • Citations are not traffic. Nothing here predicts visits, rankings or revenue.

Engines. Pin the model you mean in the config and keep it fixed between runs; diff flags a run where the model changed. OpenRouter covers hundreds of models with one key.

Engine id (default)ProviderKeyDefault modelWhat counts as citedSearches shownCost reported
chatgpt-apiOpenAI Responses API + web_searchOPENAI_API_KEYgpt-6-lunaurl_citation annotationsyesno, set prices
perplexity-apiPerplexity SonarPERPLEXITY_API_KEYsonarnumbered sources the answer usesnoyes
gemini-apiGemini Interactions API + google_searchGEMINI_API_KEY or GOOGLE_API_KEYgemini-3.8-flashurl_citation annotationsyesno, set prices
claude-apiAnthropic Messages + web_search toolANTHROPIC_API_KEYclaude-sonnet-5citations on the answer textyesno, set prices
any id you chooseOpenRouter + web pluginOPENROUTER_API_KEYopenai/gpt-6-lunaurl_citation annotationsnoyes

Caps and cost. max_calls is an exact cap on questions asked in a run, and a plan that exceeds it refuses to start. budget_usd is checked before every call against the spend so far plus the most expensive call seen on that engine, so it can be exceeded by at most one call per engine; with a budget set, engines whose cost cannot be known are left out unless you pass --allow-unpriced. Every run keeps a ledger of requests, tokens, searches and cost. Configuration, prices and all commands: cli/README.md.

From an agent. Levels 2 and 3 are also MCP tools, and an agent can only tighten your caps, never loosen them: see MCP server.

Every week. Citation rates move slowly and noisily: run on the same weekday with the same questions and models, and read diff, not single runs. examples/ci/watch-weekly.yml does it on a schedule with keys from repository secrets and commits each run, so the history lives in git. Run files contain your questions and the full answers: use a private repository if they are confidential.

Run it yourself, or have it run for you. Everything here is MIT; the tool has no paid edition. You pay your model providers directly and the ledger shows what each run used. What needs people rather than an API, Jianrun does as a service:

Run it yourself (free)Run by Jianrun
ChannelProvider APIs with searchAPIs and the consumer apps, sampled by hand each week
EnginesOpenAI, Perplexity, Gemini, Anthropic, anything on OpenRouterChatGPT with search, Perplexity, Gemini, Google AI Overviews, Copilot; Chinese engines when you sell into China
ReportText, Markdown, CSV, JSONA weekly report with the AIV dashboard: citation trend, per-engine rates, facts AI gets wrong about you, readiness history
When citations dropOut of scopeWe do the fixing
PriceYour API billAEO delivery system, from US$5,780 per 3 months, AIV dashboard included. See pricing

Why a rubric, not just a tool

A score you cannot audit is a number someone made up. So the specification is the product, and the tools are implementations of it:

  • Versioned. Every score reports the rubric version. 71 (v1.1) is a claim; 71 is not.
  • Tiered, with counts. Each tier names a page count out of 8, not "most".
  • Evidence-bound. Every check requires an observation someone else can reproduce.
  • Calibrated against public benchmarks, with the record published — including the four external sources the thresholds were checked against, and the eight specification ambiguities that real audits surfaced and v1.1 settled.
  • Machine-readable. rubric/v1.1.json with stable check ids, and schema/report.v2.json so results from different implementations are comparable.

Implement it in your own stack, disagree with a weight, open a rubric proposal. That is the main thing we want contributions on.

Scope — what this does not do

This is the part most tools leave out, so it's stated plainly.

AIV Score measures. It does not fix.

Not includedWhy
Fix templates — robots.txt, JSON-LD blocks, llms.txt boilerplateRemediation is where the actual work and judgement live. It is a separate, non-open project
Content rewriting — how to shape a passage so it gets quotedSame
Per-engine tactics — what to do differently for Perplexity vs GeminiSame
A remediation roadmapSame

Other honest limits:

  • It measures input-side readiness, not outcomes. A high readiness score means engines can cite you. Whether they do depends on competition, query intent and factors no external audit can observe. Citation performance is reported as a separate, unscored block and never folded into the 100 — see Two scores. Measure it with levels 2 and 3 above.
  • Brand Credibility and the named-author check need human judgement. "Is this a real identifiable person" and "is this mention independent" are not fully automatable. Treat those ~24 points as assisted, not automatic.
  • Tiers reduce disagreement, they do not remove it. Every tier names a count out of the 8 sampled pages, so two auditors agree on the arithmetic. They can still disagree on whether a given paragraph is a self-contained answer. The settled ambiguities are the ones we found; there will be more.
  • Heavily client-rendered sites score low, sometimes unfairly. If your content only appears after hydration, most checks will read the pre-hydration HTML — which is also roughly what a crawler sees, so the low score is usually right, but verify by hand.
  • Engine behaviour moves. The rubric is versioned for exactly this reason. A score from an older rubric version is not comparable to a current one.

Research behind the weights

The weights are opinionated but not invented. The two findings that most shaped them:

  • Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024 — citing sources, adding statistics and quoting experts raise visibility by up to 40% (measured as Position-Adjusted Word Count, not citation count). Notably, the paper found an authoritative tone produced no significant improvement — which is why this rubric scores structure and attribution, not voice.
  • llms.txt proposal, Answer.AI — the convention this rubric checks for in the Understandable pillar (p1.llms-txt).

Every check declares what its weight rests on (published research, a vendor's own documentation, our field observation, or a convention no engine has confirmed using), with sources and the date they were last verified: rubric/evidence-v1.1.md. Where the evidence is thin (p2.answer-passages, p2.question-intent, p1.llms-txt), the table says so. If you have evidence that a weight is wrong, open a rubric proposal — that is the main thing we want contributions on.

Related tools

Deliberately naming what this is not, so you can pick correctly:

ProjectWhat it doesRelationship
llms-txtThe llms.txt specification itselfAIV checks for compliance with it
yao-geo-skills21 categorized GEO skills, execution-orientedComplementary — they do production, this does measurement
GEOFlowFull GEO operations system for company sitesMuch larger scope; AGPL

If you need remediation and not just a score, those projects overlap with the part this repo deliberately excludes.

Stability, privacy and install options

  • Install: nothing to install for level 1 (the curl line above, pinned to a release). For a geo-score command with all three levels, pipx install git+https://github.com/jianruntech/geo-score@v1.4.0, or run it without installing: uvx --from git+https://github.com/jianruntech/geo-score@v1.4.0 geo-score example.com. Standard library only either way. Release downloads carry a SHA256SUMS file (how to check).
  • Stability: the rubric and the tool are versioned separately, and the report schema, check ids, CLI flags, exit codes, Action inputs and MCP tools are public contracts. What may change in which release: STABILITY.md.
  • No telemetry. geo-score sends nothing to us or anyone else. Level 1 fetches the site you name (following its redirects), the URLs its markup and robots.txt point to (sameAs profiles, logo, sitemap), and Wikidata and Wikipedia search; levels 2 and 3 call only the AI providers whose keys you set.
  • Security: text quoted from the audited site is fenced as data in every report, the MCP server's score_site connects only to public addresses, and keys never reach a file. Threat model and reporting: SECURITY.md.

Why did my score move 4 points? Each run samples up to 8 pages, and the sample changes; the measured spread is ±5 (REPRODUCIBILITY.md). To compare before and after a change, re-score the same pages with --urls-from last.json. Why does a check show —? It was not measured, so it left the denominator instead of scoring 0. More answers: troubleshooting.

Who maintains this

Built and maintained by Jianrun Tech (见润科技), Shenzhen — we run GEO and AI-adoption programs for cross-border commerce companies. The rubric came out of client work and out of optimizing our own products; publishing it is how we'd like AI visibility to be measured consistently, including by people who never become our clients.

Commercial use of this repository is unrestricted under MIT — including inside paid consulting work. You do not need our permission, and there is no separate commercial licence.

Contributing

The most valuable contribution is evidence about the weights. See CONTRIBUTING.md.

Citation

If you reference the rubric in research or a report, see CITATION.cff.

License

MIT