Detection rules & scoring

Every check the scanner runs, grouped by category, and exactly how each one affects your score.

How a score is calculated

Each category starts at 100 points. Every rule that fires deducts:

penalty = rule weight × severity multiplier × coverage factor

The coverage factor scales the penalty by how many of the analysed pages were affected — one bad page barely dents the score, the same issue across most of the site costs much more. A critical technical issue also caps the whole Technical category at 59, regardless of other penalties. The global score is a weighted average of the category scores below.

Category weights

  • Technical24.8%
  • SEO20.7%
  • Catalogue16.6%
  • Conversion12.4%
  • Trust8.3%
  • GEO9.2%
  • Accessibility8%

Severity multipliers

  • Critical1.00
  • High0.65
  • Medium0.35
  • Low0.15
  • Info0.00

Technical 24.8% of the global score · 9 rules

  • Page returns a server error

    Criticalweight 60

    An internal page responded with a 5xx status code.

  • Slow server response time

    Mediumweight 20

    At least three pages took more than 800ms to respond.

  • Oversized HTML document

    Mediumweight 20

    The HTML document is larger than 1MB.

  • Mixed content: HTTP image on an HTTPS page

    Highweight 35

    An HTTPS page loads at least one <img> asset over plain HTTP, which browsers block or flag as insecure. Scoped to images — the only asset type the crawler currently inventories in page_assets; scripts, stylesheets and fonts are not yet checked, so this is partial mixed-content coverage, not exhaustive.

  • Missing viewport meta tag

    Highweight 35

    The page has no <meta name="viewport"> tag, so mobile browsers render it at desktop width, forcing users to pinch-zoom.

  • Broken internal link

    Highweight 35

    An internal link on a crawled page points to another crawled page that returned a 4xx/5xx status. Links to pages the crawler never reached are not flagged — their status is unknown, not broken.

  • Page requires multiple redirects

    Mediumweight 20

    The page took more than one redirect hop to reach its final destination, adding latency and eroding link equity with every extra hop.

  • Content overflows on mobile

    Highweight 35

    On a 390px-wide mobile viewport, the page's content is wider than the viewport itself, forcing horizontal scrolling — usually an unconstrained image, table, or fixed-width element.

  • JavaScript console error on page load

    Lowweight 10

    The page logged at least one JavaScript error to the browser console while loading — often a broken script, a failed third-party integration, or a bug that silently disables part of the page.

SEO 20.7% of the global score · 12 rules

  • Missing page title

    Highweight 40

    The page has no <title> element.

  • Duplicate page title

    Mediumweight 25

    Multiple indexable pages share the same title.

  • Title length outside recommended range

    Lowweight 10

    Title is shorter than 30 or longer than 60 characters.

  • Missing meta description

    Mediumweight 20

    The page has no meta description tag.

  • Thin page content

    Mediumweight 20

    The page has fewer than 100 words of visible text, which search engines and AI answer engines typically treat as too little content to rank or cite confidently.

  • Missing H1 heading

    Mediumweight 25

    The page has no H1 element.

  • Canonical tag mismatch

    Highweight 35

    The canonical tag points to a different URL than the page itself.

  • Multiple H1 headings

    Lowweight 10

    The page has more than one H1 element, diluting the primary heading signal for search engines. Complements MISSING_H1 rather than replacing it — a page can only be missing an H1 or have multiple H1s, never both at once.

  • Flat heading structure

    Lowweight 10

    The page has enough text to be a real content page and an H1, but no H2 elements underneath it — a flat wall of text with no sub-headings, which is harder for both search engines and AI systems to parse into sections.

  • No sitemap found

    Mediumweight 20

    Neither /sitemap.xml nor any sitemap referenced in robots.txt was reachable, making it harder for search engines to discover and prioritize the site's pages.

  • Sitemap URL has no internal link pointing to it

    Mediumweight 15

    A URL listed in the sitemap was never linked to from any crawled page, so shoppers (and crawlers that don't read the sitemap) have no way to navigate to it. The site's own homepage is never flagged even though nothing on the site typically links to it either.

  • Public product page is marked noindex

    Highweight 35

    The page's URL looks like a product page but its robots meta tag includes "noindex", hiding it from search results — often left over from staging/testing rather than intentional.

Catalogue 16.6% of the global score · 8 rules

  • Product page has no structured data

    Mediumweight 25

    The page's URL looks like a product page (e.g. a path containing "product"), but it has no Product JSON-LD, so search engines and AI systems can't reliably extract its name, price, or availability as structured facts.

  • Product page has no meaningful description

    Highweight 35

    The product page's meta description is under 80 characters (a proxy for "no real description content" — the crawler doesn't extract a dedicated product-description block, so this checks the meta description rather than the on-page copy itself).

  • Product page has no images

    Criticalweight 55

    The page's URL looks like a product page but the crawler found no images on it, which is close to disqualifying for an online shopper.

  • Product page has only one image

    Lowweight 10

    The product page has exactly one image. Shoppers buying without trying an item on typically expect multiple angles/detail shots before they trust a purchase.

  • Oversized product image

    Mediumweight 25

    The product page loads at least one image larger than 500KB, slowing page load — especially costly on mobile connections where shoppers are most likely to abandon a slow-loading product page.

  • Legacy image format on a sizeable image

    Lowweight 10

    The product page loads at least one JPEG/PNG image over 200KB with no modern-format (WebP/AVIF) alternative — those formats typically shrink the same image by 25-50% at equivalent visual quality.

  • Product page shows no price

    Highweight 35

    The page's URL looks like a product page but no currency-formatted amount (e.g. "$19.99", "19,99 €") was found anywhere in its visible text.

  • Product page shows no stock/availability signal

    Lowweight 10

    The page's URL looks like a product page but shows no in-stock/out-of-stock signal, either as visible text (in any supported language) or as structured Product offers.availability.

Conversion 12.4% of the global score · 2 rules

  • No cart page found

    Highweight 25

    None of the crawled pages or known internal links look like a shopping cart (e.g. a path containing "cart" or "basket"). Also checks links the crawler saw but didn't fetch (e.g. a cart URL disallowed by the site's own robots.txt), so this only false-positives on a slide-out/JS-only cart with no dedicated URL anywhere in the site's markup.

  • Product page has no add-to-cart control

    Criticalweight 55

    The page's URL looks like a product page but no button, link, or submit control with add-to-cart/buy-now text (in any supported language) was found — the single most direct conversion blocker a product page can have.

Trust 8.3% of the global score · 6 rules

  • No contact page found

    Mediumweight 20

    None of the crawled pages or known internal links look like a contact page (e.g. a path containing "contact"), so shoppers may struggle to reach the merchant.

  • No legal or privacy pages found

    Mediumweight 20

    None of the crawled pages or known internal links look like a privacy policy, terms of service, or other legal page.

  • No shipping information found

    Mediumweight 20

    None of the crawled pages or known internal links look like a shipping or delivery information page (e.g. a path containing "shipping" or "delivery").

  • No returns information found

    Lowweight 15

    None of the crawled pages or known internal links look like a returns or refund policy page (e.g. a path containing "returns" or "refund").

  • No cookie notice found

    Infoweight 10

    None of the crawled pages or known internal links look like a cookie notice or cookie policy page. Notice only — this does not assert or check legal cookie-consent compliance.

  • Invalid HTTPS certificate

    Criticalweight 60

    At least one page failed to load because of a TLS certificate problem (expired, self-signed, or not matching the hostname), which shows shoppers a security warning before they can even see the page.

GEO 9.2% of the global score · 8 rules

  • AI crawlers are blocked

    Criticalweight 70

    robots.txt disallows major AI answer-engine crawlers (GPTBot, ClaudeBot, PerplexityBot, and others) from reading the site, so the store cannot be cited by AI answer engines.

  • No llms.txt file found

    Lowweight 15

    llms.txt gives AI crawlers a clean, curated summary of the site's content.

  • No structured data found

    Mediumweight 30

    The page has no JSON-LD structured data, making it harder for AI systems to extract accurate facts about it.

  • No Organization structured data found

    Lowweight 20

    No crawled page publishes Organization (or a subtype like LocalBusiness) JSON-LD, so AI answer engines have no structured way to identify who runs the store.

  • Page content only exists after JavaScript runs

    Highweight 40

    The HTML this store sends is largely an empty shell — the real content (headings, copy, product details) is built in the browser by JavaScript. Search engines that render pages will usually still see it, but AI answer engines and most other crawlers do not execute JavaScript, so to them these pages look close to blank. It also means this report judges the pages it could render more accurately than the rest of the site.

  • Nonexistent pages return a 200 instead of a real error

    Mediumweight 20

    A request to a made-up, clearly nonexistent path on the site returned a successful (2xx) status instead of a 4xx error — a "soft 404". AI agents and crawlers rely on the HTTP status itself to tell a real page from a dead link; a soft 404 makes every broken link look like a valid page.

  • Organization structured data is incomplete

    Lowweight 10

    Organization (or a subtype like LocalBusiness) JSON-LD is present somewhere on the site, but it's missing address and/or contactPoint — the fields AI systems use to verify a business is real and reachable.

  • llms.txt has little to no real content

    Lowweight 10

    An /llms.txt file exists but is too short to give AI crawlers meaningful guidance about the site — closer to a placeholder than a real curated summary.

Accessibility 8% of the global score · 6 rules

  • Images missing alt text

    Lowweight 10

    The page has at least one <img> element with no alt attribute at all, hurting accessibility and image SEO. Images with an explicit empty alt="" (decorative) are not flagged.

  • Missing HTML lang attribute

    Mediumweight 25

    The page's <html> element has no lang attribute (or an empty one), so screen readers can't determine which language to use for pronunciation.

  • No main landmark

    Lowweight 15

    The page has no <main> element or role="main" landmark, so screen-reader users have no way to skip repeated header/navigation content and jump straight to the page's primary content.

  • Link with no accessible text

    Highweight 35

    At least one link on the page has no visible text and no aria-label, so a screen reader announces it as just "link" with no indication of where it goes or what it does.

  • Ambiguous link text

    Lowweight 15

    At least one link's text is a generic phrase like "click here" or "read more" with no surrounding context, which is meaningless to a screen-reader user navigating by a list of links out of context.

  • Form field missing a label

    Highweight 35

    At least one form field (input, select, or textarea) has no associated <label>, aria-label, or aria-labelledby, so a screen-reader user has no way to know what information the field expects.