Methodology
One question drives everything here: can search engines find the website, and can AI systems understand, trust, and recommend it? This page explains exactly how we score that, and, just as importantly, what we refuse to claim.
Scoring rules
- Undeterminable criteria are skipped, not zeroed. If a check cannot be established from public evidence (e.g. traceroute from Oman, backend geolocation behind a CDN), its weight is excluded and the section score is normalized over what could actually be measured. A site is never punished for the scanner's blind spots.
- Scores have separate models. The established 10-category, 100-point model remains the overall diagnostic view. AI Readability is now calculated independently from the 15 research criteria below, while Oman Localization Readiness keeps its own /100 methodology.
- Oman locality is binary. Evidence either points to Oman or it does not. A UAE or wider-GCC edge earns no Oman-locality points; regional proximity is reflected only in performance-related checks.
- One report per domain per day (Muscat time). Re-scanning the same domain on the same day serves the stored report (on-screen and PDF). Every scan issues a signed, timestamped PDF, and every criterion score is stored individually.
The overall category model (100 points)
Ten categories, weighted by real-world impact, provide the broad overall diagnostic model. SEO, category-level findings and sub-scores still use this model; the dedicated AI Readability and Oman scores are calculated separately.
Crawlability & Indexability
15 ptsHTTP status, robots.txt, noindex, canonicals, sitemap, static-HTML content, internal links and search/retrieval crawler access. Training-crawler consent is reported separately.
Site Architecture & Internal Linking
10 ptsHomepage positioning, service URL structure, per-service pages, breadcrumbs, cross-linking, footer completeness, orphan-page sampling.
Page Structure & AI Readability
15 ptsSingle H1, definitional opening (first 100–150 words), H2/H3 hierarchy, scannable formats, audience, offering, outcomes, CTA, fluff detection.
Content Usefulness & Answer Quality
15 ptsDirect answers, specificity signals, use cases, process, pricing, FAQs, limitations, keyword-stuffing check, plus five LLM-style extraction tests (heuristic in MVP).
Entity Clarity & Brand Understanding
10 ptsCompany name, brand consistency, market, team, service naming, contacts, sameAs, description consistency, superlative-claim check, category positioning.
Structured Data / Schema
10 ptsOrganization, LocalBusiness, Service, FAQPage, Article, BreadcrumbList; JSON-LD validity; schema-vs-visible-content match; enrichment properties.
Trust, Proof & Authority
10 ptsCase studies, quantified results, testimonials, demos/artifacts, team credibility, external references, dates, claim verifiability.
Agent / LLM Delivery Layer
7 ptsSame-URL Markdown negotiation, Vary: Accept, legacy .md-route redirects and optional llms.txt discovery.
Local, Multilingual & Market Signals
5 ptsService area, visible contacts, Arabic/English, hreflang, local proof (OMR, +968, Oman regulations, city references).
Performance, Accessibility & UX
3 ptsViewport, page-weight & TTFB proxies, semantic elements, labels/alt text, layout-stability hints. Lightweight MVP checks, not lab data.
Rating bands
- 90–100 Excellent
- 75–89 Good
- 60–74 Average
- 40–59 Weak
- 0–39 Poor
Entity, Structured Data, Trust and Agent-layer scores are their categories rescaled to 100. Rating bands are applied to each reported score, rather than blending AI Readability or Oman readiness into an unsupported all-purpose number.
Dedicated AI Readability (15 criteria, normalized /100)
The research rubric assigns 105 raw weight points across 15 criteria. Each applicable criterion is scored as pass, partial or fail. The displayed score is earned applicable weight ÷ total applicable weight × 100, so 105 is the research denominator before N/A exclusions, not a score above 100. Privacy/residency and linkability/authority are explicitly N/A when the scanner lacks measurable public evidence; their weights are then excluded rather than treated as zero.
1. Crawlers / robots access
weight 15Can search and retrieval crawlers reach the public pages intended for discovery?
2. llms.txt discovery file
weight 10If published, is this optional, emerging file valid Markdown and does it link to canonical regular URLs?
3. Same-URL Markdown negotiation
weight 15Does the canonical page return Markdown for Accept: text/markdown while continuing to return HTML for normal browser requests?
4. Legacy Markdown migration
weight 5Do previously used .md twins permanently redirect to regular canonical paths, with obsolete alternate/internal/llms.txt references removed? This is N/A when no twin evidence exists.
5. Structured data
weight 5Is relevant JSON-LD present, valid and consistent with visible content?
6. Canonicalisation
weight 10Do HTML pages identify the regular URL as canonical, without a competing active .md URL?
7. Vary: Accept
weight 5Do both negotiated representations send Vary: Accept so shared caches keep HTML and Markdown separate?
8. Bot-specific directives
weight 5Are search/retrieval controls assessed separately from model-training consent?
9. Rate limiting / anti-bot barriers
weight 5Can a modest, respectful crawl proceed without CAPTCHA, challenge pages or immediate throttling?
10. Content quality / answerability
weight 20Is the content specific, complete, well structured and able to support accurate answers and citations?
11. Token efficiency
weight 3Does the negotiated Markdown remove navigation, scripts and repeated boilerplate while preserving meaning and links?
12. Freshness
weight 2Are meaningful published or updated dates exposed where recency matters?
13. Language / accessibility
weight 3Are language, headings, semantics and image alternatives machine-readable and usable?
14. Privacy / residency signals
weight 1Are relevant public claims clear and supportable? Marked N/A when they cannot be measured from public evidence.
15. Linkability / authority
weight 1Can externally verifiable citations or authority signals be measured? Marked N/A when no suitable evidence source is available.
Canonical Markdown policy
The required target is one canonical URL with HTTP content negotiation. A normal request returns HTML; the same URL requested with Accept: text/markdown returns Content-Type: text/markdown. Both variants must include Vary: Accept.
Active .md twins are treated as legacy routes, not something to create or extend. Existing twins should permanently redirect with 301 or 308 to the regular canonical URL: /index.md → /, /section/index.md → /section, and /page.md → /page. Markdown is then served by negotiation at that destination.
llms.txt is an optional, emerging discovery convention, not a requirement, ranking guarantee or replacement for robots.txt. When present, it should link to regular canonical URLs—not legacy .md twins—so clients can request Markdown from the same URL. HTML or HTTP alternate links that still advertise a twin should be retired as part of the migration.
Search access is not training consent
The audit separates bots used for search or user-requested retrieval (for example OAI-SearchBot, ChatGPT-User, Claude-SearchBot and Claude-User) from bots used for model training (for example GPTBot and ClaudeBot). Blocking a training crawler is a consent choice and does not automatically mean the site is unavailable to AI-assisted search. Recommendations identify the affected purpose instead of telling every site to allow every AI bot.
Oman Localization Readiness (separate /100)
Checks whether a website is locally relevant, fast for Oman users, and whether there is evidence of Oman-hosted infrastructure. Locality checks (IP, ASN, DNS, API host) are binary: Oman or not Oman. A .om / .com.om domain is also credited: registration requires an Omani commercial presence, making it registry-verified local proof.
| Business & content localization | 20 |
| Technical hosting proximity | 25 |
| Asset/resource localization | 15 |
| Backend/API localization evidence | 15 |
| Data residency evidence | 10 |
| Performance for Oman users | 10 |
| Transparency/disclosure | 5 |
Latency interpretation bands (for a real Oman probe): 30–80ms strong local/GCC signal · 100–180ms possible UAE/GCC · 200–350ms likely farther region · 400ms+ poor regional localization. Latency is always evidence, not proof. The MVP measures TTFB from the scanner’s serverless region and labels it approximate; the architecture accepts future probes in Oman, UAE, Saudi Arabia, India, Europe and the USA.
Confidence labels
Every finding carries one of six labels. They are not decoration. They are the product.
Confirmed
Directly observed in a server response or file. E.g. “robots.txt blocks GPTBot”.
Strong evidence
Multiple consistent public signals. Very likely true, technically falsifiable.
Moderate evidence
A reasonable inference from public signals; heuristics may misread edge cases.
Weak evidence
A hint, not a conclusion. Often from limited crawl samples or single-region measurements.
Not externally verifiable
Cannot be established from outside at all: database location, backup regions, origin servers behind CDNs. We report claims, never facts, here.
Contradictory evidence
Public signals disagree with each other or with the site’s own claims.
What this scanner will never claim
Backend location. Modern websites often use CDNs and reverse proxies. Public tests may reveal the visible edge server, but not the private origin server, backend application server, database, logs, or backup location. When we detect Cloudflare, CloudFront, Akamai, Fastly, Vercel, Netlify, BunnyCDN or Azure Front Door, we say so: “The visible server appears to be a CDN/proxy. Origin backend location cannot be confirmed externally.”
Data residency. Database and data residency cannot usually be verified from outside. This scanner only reports public technical evidence and visible policy claims, always labeled “externally claimed, not independently verified.”
Rich results. Structured data may help search engines understand the page, but it does not guarantee rich results or ranking improvements.
llms.txt. This optional, emerging file is not a ranking factor we can guarantee, does not replace or modify robots.txt, and is never treated as a substitute for crawlable canonical pages or high-quality content.
Blocking training crawlers is not automatically wrong. GPTBot or ClaudeBot restrictions are reported as training-consent choices. Search/retrieval restrictions are reported separately because those can affect AI-assisted discovery, citation or user-requested access.
Scan mechanics
- Fetches the homepage, robots.txt and sitemaps, then crawls up to ~10 prioritized same-domain pages (about, contact, services, pricing, case studies, blog, privacy, terms…).
- Parses static HTML only, with no JavaScript execution. This mirrors how most AI crawlers see the site, and low static-HTML content is itself reported as a finding.
- Requests representative canonical pages as both HTML and Markdown, verifies Content-Type and Vary: Accept, and compares the negotiated bodies.
- Checks optional /llms.txt and /llms-full.txt files, then probes legacy .md patterns only to detect active twins and recommend their 301/308 migration to canonical negotiated URLs.
- Extracts JSON-LD (plus microdata types), asset hosts, API hosts from HTML, form actions and up to 5 same-site JS bundles.
- Resolves DNS, follows CNAME chains, fingerprints CDNs from headers, and geolocates the visible IP via a pluggable provider (free-tier accuracy; treated as evidence, not fact).
- Reports are stored in Supabase Database and private Storage so results are shareable by permalink.
Ready to see your numbers?
Run the free scan