Overview
The Auric Artisan URL Analyzer is a client-first website audit application that needs a signed-in
account to run. In production, server rendering and full-site scans come from the auth worker
(/auth/analyzer/scan and /auth/analyzer/crawl, metered against the plan and
tokens); an optional Node.js backend serves rendered snapshots, CORS-blocked pages, full-site scans,
paged scan results, and report export for local development. It combines accessibility checks, WCAG contrast testing, SEO metadata review, rendered DOM
comparison, framework detection, Core Web Vitals-style metrics, security and reliability signals,
media analysis, device readiness, color palette extraction, and developer fix generation.
This reference documents how the tool is wired, how data moves through the audit pipeline, which files own each concern, where the extension points live, and how to safely maintain or expand the analyzer from a code editor such as Cursor, VS Code, or any AI-assisted IDE.
1. File map
+
The analyzer UI is declared in tool/analyzer/index.html. It loads the site chrome,
cursor layer, scroll layer, loader, accessibility rest panel, SVG color-vision filters, the
analyzer application shell, a hidden preview iframe, and the analyzer modules.
- tool/analyzer/index.html: page shell, tabs, workflow buttons, scan controls, hidden iframe, metadata, and script order.
- js/tool/analyzer/analyzer.js: primary audit engine, UI renderers, scan
orchestration, scoring, local storage, exports, rules, and panel wiring. It loads
deep-analysis.js(enrichReportand the advanced modules),report-export.jsand the A11y+ engine on demand. - js/tool/analyzer/analyzer-schema.js: JSON-LD report schema, Open Graph tags, share scorecard SVG, score labels, and schema-enhanced report export.
- js/tool/analyzer/pagination.js: reusable pagination classes:
PaginationState,PaginationUI,PaginationController, andCategoryPagination. - js/tool/analyzer/analyzer-pagination-integration.js: drop-in pagination for Top 5 fixes, priority engine rows, and device result cards.
- css/analyzer.css: analyzer layout, cards, panels, issue rows, scan results, deep analysis, playground, report, and responsive behavior.
- css/components/pagination.css: pagination controls, category pagination, page stats, per-page selectors, and device card pagination styles.
- js/server.js: optional analyzer backend with fetch, snapshot, scan, report export, Playwright rendering, JSDOM parsing, rate limiting, and static file serving.
2. Runtime pipeline
+
The main control flow starts at startAnalysis(url). After normalizing the URL it asks
/auth/status; signed out, it sends the visitor to /auth/?next=… with the
URL carried, and the run starts on return. A run receives an incrementing
run ID and an AbortController, then updates the guided workflow to the Analyze stage.
Stale runs are ignored so a later analysis cannot be overwritten by an older request.
- Normalize the URL:
normalizeAnalyzerUrlaccepts pasted domains, absolute URLs, and current-page targets, then writes the normalized URL back to the input. - Same-page path: if the requested URL is the current page, the analyzer uses direct DOM access and captures browser timing from the active window.
- Iframe path: same-origin pages are loaded into the hidden iframe and audited directly.
- Rendered snapshot path: with a session and no configured backend, the page is
first rendered on the server through
/auth/analyzer/scan; with a backend, cross-origin pages try the backend/api/snapshotpath, which uses Playwright to capture rendered HTML, screenshot metadata, resources, failed requests, and browser audit data. - Static HTML path: if rendering fails, the analyzer fetches static HTML, injects it into a blob URL, and audits the parsed document.
- Backend analysis path: if iframe and fetch paths fail, the analyzer calls
/api/analyzeand enriches the returned server report on the client. - Cache path: rendered snapshots are stored in local storage for up to seven days with HTML and image size caps, then reused when live capture fails.
Network work goes through politeFetch, which applies per-host queues, adaptive
concurrency, Retry-After handling, 429/503 backoff, timeouts, and clear offline or CORS messages.
3. Report model and scores
+
The base report is built by analyzePage(doc, url, options). It starts with metadata,
accessibility, color, DOM, performance, links, issues, contrast pairs, and fixes. The enriched
report is produced by enrichReport(R, doc), which attaches all advanced modules and
computes final category scores.
- Core fields:
url,ts,meta,a11y,colors,dom,perf,links,issues,contrastPairs, andfixes. - Advanced fields:
metadataAudit,colorAudit,visualAudit,mediaAudit,deepAnalysis,performanceAudit,security,devices,reliability,seoAudit,wcagMapping,advancedSeo,aiInsights, andunifiedHealth. - Weighted score: accessibility 20%, contrast 16%, performance 13%, metadata 9%, responsive 8%, reliability 7%, security 7%, media 2%, visual 8%, and deep analysis 10%.
- Panel scores: SEO, A11y+, media, custom rules, and advanced panels also expose their own scores and findings, even when their exact panel score is not a direct top-level weight.
- Issue workflow: generated issue IDs normalize category, message, severity, confidence, cause chain, estimated score gain, status, and suggested fixes for UI rendering.
4. Backend APIs
+
Production runs use the auth worker routes /auth/analyzer/scan,
/auth/analyzer/crawl, /auth/analyzer/screenshots and
/auth/analyzer/limits. The optional local backend is js/server.js. Run it with npm run analyzer or
node js/server.js. It defaults to http://localhost:3001, or a custom
port from ANALYZER_PORT. The frontend reads auric_analyzer_backend from
local storage when a custom backend origin is needed.
- GET /api/health: lightweight readiness check used before backend calls.
- POST /api/analyze: accepts
{ url, rules }and returns a full single-page report built with server fetch, JSDOM, optional Playwright fallback, and shared audit helpers. - POST /api/snapshot: accepts
{ url }and returns rendered HTML, screenshot availability, resources, failed requests, browser audit timing, and visual data. - POST /api/fetch: fetches raw HTML or text for CORS-blocked client paths.
- POST /api/site-scan: starts a robots-aware multi-page scan; supports streamed NDJSON progress and compact result storage.
- GET /api/site-scan-page: returns 25 or 50 full page details at a time from a stored scan result.
- GET /api/site-scan-result: returns a stored scan summary by result ID.
- POST /api/export/report: exports analyzer reports as HTML or PDF when the backend is available.
Important backend limits are environment-configurable: body size, request timeout, analyzer concurrency, rate-limit window, scan result TTL, maximum scan pages, scan page size, scan concurrency, stored result count, and browser fallback concurrency.
5. Audit modules
+Each module focuses on one class of evidence, then adds findings to the common issue system. This keeps tab rendering flexible while preserving one report object for export and site-scan aggregation.
- Metadata: title, description, charset, viewport, language, canonical, robots, Open Graph, Twitter metadata, favicon, and site identity.
- Accessibility: heading order, H1 count, image alt text, link names, form labels, landmarks, skip link, positive tabindex, focus indication, table headers, autoplay media, nested interactivity, keyboard flow, screen-reader flow, audio accessibility, and WCAG mapping.
- Contrast and color: computed foreground/background pairs, WCAG grade, large text threshold, fix color generation, palette extraction, color mode detection, CVD-safe pair checks, and design lab integration.
- Performance: browser audit timing, Core Web Vitals-style LCP, CLS, and INP, resource waterfall, bundle analysis, dependency graph, network simulation profiles, and JS execution risk.
- Deep analyzer: raw HTML scoring, rendered DOM scoring, raw/rendered comparison, rendering mode classification, framework stack detection, route discovery, SEO generation, JavaScript intelligence, accessibility behavior, and developer fix plan.
- Media and visual: image/video/audio inventory, alt and caption checks, loading behavior, intrinsic dimensions, canvas sampling, scene detection, pixel palette, and visual risk notes.
- Security and reliability: HTTPS, mixed content, CSP, referrer policy, permissions policy, inline handlers, broken links, status errors, third-party scripts, and server stability signals.
- Responsive devices: mobile, tablet, foldable, laptop, desktop, wide desktop, and custom preview workflows using the responsive device profile list.
6. SEO and structured data
+The analyzer has both on-page SEO auditing and report-level structured data export. On-page SEO checks inspect title length, description length, canonical, robots directives, headings, schema, keyword signals, indexability, crawl quality, sitemap discovery, and raw versus rendered SEO changes.
- SEO panel: built by
buildSeoAuditand related advanced SEO helpers; includes metadata, headings, internal links, structured data, word count, and indexability. - Deep SEO panel: compares raw HTML with rendered DOM to identify content, schema, title, canonical, robots, and H1 fields injected or changed after rendering.
- Sitemap discovery: probes robots.txt, common sitemap paths, numbered sitemap patterns, sitemap indexes, and URL sets with capped traversal.
- Report schema:
analyzer-schema.jsgenerates a Schema.org graph with WebPage, audit event, Report, FAQPage, BreadcrumbList, and metrics data. - Share metadata: Open Graph and Twitter tags can be generated from the report, including a dynamic SVG scorecard image data URL.
- Export model:
exportReportWithSchema(report)wraps the report with@context, generated schema, export date, and analyzer version.
7. Full-site scan system
+
Full-site scans are backend-powered because they need robots handling, link discovery, sitemap
discovery, concurrency control, streamed progress, large result storage, and page-by-page retrieval.
In production they run through /auth/analyzer/crawl, which caps the pages per crawl
(24 on Cloudflare Browser Rendering); the local backend uses /api/site-scan.
The UI exposes robots, subdomains, max pages, page size, and worker controls in Advanced mode.
- Discovery: starts with the requested root URL, robots.txt sitemap entries, common sitemap paths, rendered links, static links, and scoped internal URLs.
- Scope: scanning stays on the same base host unless the subdomain option is enabled. Tracking parameters are normalized away for cleaner dedupe.
- Respect robots: enabled by default; the backend parser reads user-agent, allow, disallow, sitemap, and crawl-delay directives.
- Progress: streaming emits page-start, page-done, page-error, page-skip, status, result-ready, result, and done events.
- Paging: large scans store full pages in memory for a TTL and return compact
aggregate data first. Detailed page rows load through
/api/site-scan-pagein 25 or 50 item pages. - Sorting: site-scan rows can sort by score ascending, score descending, issue count, title, URL, or scan order.
8. Extension points
+The safest way to extend the analyzer is to add evidence to the report object, add findings through the shared issue pipeline, and render a panel from that normalized report data. Avoid one-off DOM state that cannot be exported.
- Add a new check: implement a function that returns checks, findings, score,
and optional fixes; call it from
enrichReport. - Add a tab: add a button and panel in
index.html, add it topanels, include it inMODE_TABS, then add a renderer ingetPanelRenderers(). - Add a custom rule template: extend
RULE_TEMPLATESwith selector, type, severity, message, fix, and enabled state. - Add report schema: extend
generateAnalyzerSchemawith new metrics orhasPartentries and keep export output serializable. - Add pagination: use
PaginationControllerfor single arrays orCategoryPaginationfor grouped categories. - Add backend evidence: compact large fields before returning JSON; avoid shipping raw multi-megabyte HTML, screenshots, or canvas payloads in full-site scan summaries.
9. Safety, privacy, and limits
+- Use the analyzer only on sites you own or are authorized to test.
- Keep robots.txt respect enabled for full-site scans unless you have clear permission to ignore it in a controlled environment.
- Use conservative worker counts. The UI caps user-requested backend site-scan concurrency to a safe range to avoid rate limits and unnecessary load.
- Do not persist raw HTML, screenshots, or scan artifacts longer than needed. The browser cache uses a seven-day snapshot TTL and trims oversized HTML and images.
- Review CSP setup before using live previews. Some pages intentionally block framing or cross origin access; the analyzer should report that constraint rather than bypass site policy.
- Backend deployment should set explicit allowed origins, request limits, body limits, and concurrency limits before public exposure.
10. Cursor and AI editor handoff
+
When editing this tool from Cursor or another AI-assisted editor, give the model the same file map
used in this guide. The most reliable workflow is to open tool/analyzer/index.html,
js/tool/analyzer/analyzer.js, analyzer-schema.js, and
js/server.js together, then make scoped changes around one audit module at a time.
- Best prompt shape: "Modify only the analyzer module for [audit area]. Preserve report schema, issue normalization, tab rendering, and export behavior."
- Ask for references: have the editor cite function names such as
enrichReport,buildSeoAudit,buildDeepAnalysis, orrenderReportbefore it edits. - Check UI modes: after a change, test both Simple and Advanced mode because
tabs are filtered by
MODE_TABS. - Check exports: any new report field should serialize to JSON and avoid circular references, DOM nodes, window objects, or unbounded HTML.
- Check generated docs: after adding or renaming documentation articles, run the discovery generator so RSS, sitemap, and the global feed stay in sync.
11. Troubleshooting
+- Backend server not running: start it with
npm run analyzer. On local hostnames the frontend defaults tohttp://localhost:3001; production hosts use a declared backend or none. - CORS or iframe failure: the analyzer should fall back to rendered snapshot, static fetch, backend analysis, or cached snapshot. If all fail, inspect the failure message for CORS, DNS, bot challenge, or TLS hints.
- Site scan is too slow: lower max pages, keep page size at 25, reduce workers, and keep subdomains disabled unless required.
- Large reports break export: compact raw HTML, screenshots, canvas data, and repeated resource arrays before shipping JSON.
- Pagination does not appear: confirm the pagination CSS is loaded, the target
section has enough items, and the integration file loads before
analyzer.js. - Structured data not injected: verify
window.AnalyzerSchemais available and the report object includes a validurl.