AI visibility scorecard rubric
AI Visibility Scorecard Rubric
The scoring rubric behind CiteKit's AI Visibility Scorecard, including dimension weights, 0-5 evidence standards, and follow-up review guidance.
Why the rubric is public
A score is useful only when users can see how it was produced. CiteKit publishes the AI Visibility Scorecard rubric so SaaS teams can apply the same evidence standard across baseline and follow-up reviews. The score should guide page work, not replace judgment or promise rankings, AI citations, AdSense approval, or advertising revenue.
- Use the rubric on one public URL at a time.
- Score only evidence that a normal visitor or crawler can verify.
- Keep notes for why each dimension received its score.
- Compare the same URL with the same rubric after meaningful page updates.
Dimension weights
The 100-point score weights technical access, content clarity, proof, and maintenance signals. The weights are intentionally practical rather than statistical. They reflect the order in which a SaaS page should usually be fixed: make it accessible, make product facts clear, make answers extractable, add proof, cover comparisons, and maintain the review process.
| Dimension | Weight | What it asks | Why it matters |
|---|---|---|---|
| Crawlability and technical access | 20 | Can public systems access and resolve the intended URL? | Blocked or ambiguous pages are weak citation candidates even if the copy is good |
| Entity and schema clarity | 18 | Are product identity, category, pricing, features, and schema facts visible and consistent? | Answer systems need clear entity facts before they can summarize accurately |
| Answer extraction | 18 | Can important sections be quoted or summarized without surrounding context? | Clear H2 answers reduce ambiguity for both buyers and answer engines |
| Citation proof | 18 | Does the page show public evidence behind important claims? | Screenshots, docs, changelogs, and third-party mentions make claims easier to verify |
| Comparison coverage | 14 | Does the site answer alternatives and competitor questions fairly? | Comparison queries often expose missing buyer criteria and unsupported claims |
| Maintenance and review process | 12 | Is there a dated owner, update process, correction path, and repeatable log? | Maintained pages are less likely to become stale, misleading, or ad-first |
0 to 5 scoring standard
Each dimension is scored from 0 to 5. Use whole numbers. A zero means the signal is absent, blocked, or contradicted by the page. A three means the signal exists but is incomplete or partly unverifiable. A five means the signal is visible, source-backed, consistent, and reviewed.
| Score | Evidence standard | Reviewer question |
|---|---|---|
| 0 | Missing, blocked, hidden, or contradicted by the live page | Would a user or crawler fail to find this signal at all? |
| 1 | Present only as a vague claim, hidden markup, or unsupported note | Is the claim visible but too weak to trust? |
| 2 | Partially present but missing clear proof, context, or consistency | Would a reviewer need to infer the important facts? |
| 3 | Usable baseline with gaps that limit citation or buyer confidence | Could the page work, but still leave obvious questions? |
| 4 | Strong evidence with minor gaps or maintenance risk | Would a user understand and verify the claim without much extra searching? |
| 5 | Complete, visible, source-backed, consistent, and recently reviewed | Could this section stand as the primary source for the claim? |
Dimension-specific examples
The rubric should be applied with concrete evidence. The examples below show how a reviewer can distinguish a weak score from a strong score without turning the tool into a black-box metric. Use these examples as a calibration aid before filling out the scorecard form.
| Dimension | Low score example | High score example |
|---|---|---|
| Crawlability | robots.txt allows most crawlers but the page has no canonical and is missing from the sitemap | The public page is indexable, canonical, listed in sitemap, and not blocked by crawler or WAF rules |
| Schema clarity | SoftwareApplication schema mentions integrations that do not appear in visible copy | Schema matches visible product facts, pricing model, FAQ answers, and organization details |
| Answer extraction | The first H2 opens with a slogan and no direct product definition | Each key H2 starts with a concise answer followed by specific proof and examples |
| Citation proof | The page claims Slack and GitHub support without screenshots or docs | Integration claims link to setup docs, screenshots, changelog entries, or public tutorials |
| Comparison coverage | The page says competitors are worse without criteria or sources | The page defines buyer criteria and explains when each competitor may be a better fit |
| Maintenance | No review date, owner, correction path, or prompt log exists | The page has a reviewed date, correction email, update log, and saved visibility test evidence |
How to use baseline and follow-up scores
The first score should be treated as a baseline. After the team fixes the weakest dimensions, rerun the scorecard on the same URL and compare both the total score and the dimension changes. A useful follow-up note explains what changed, which evidence was added, and whether AI visibility prompts cite better sources.
- Do not compare scores across unrelated pages unless the same reviewer and rubric were used.
- Do not refresh a score just because a date changed; score changes should reflect visible page or evidence changes.
- Attach the scorecard report to the saved prompt log so the team can compare readiness and actual citations.
- If a page remains below 60 after two review cycles, fix the page before publishing more long-tail content.
Use this with a tool
Turn this page into a concrete review by starting with the visibility checker, schema generator, or crawler checker.