Queue-North-Website/scripts/prerender.js

264 lines
10 KiB
JavaScript
Raw Permalink Normal View History

feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
// Build-time prerenderer.
//
// Renders every route to static HTML so crawlers that do not execute JavaScript
// (Meta, LinkedIn, Slack, Bing) receive real content, correct per-route <title>,
// meta description, canonical, and Open Graph tags. Google also benefits: pages
// no longer sit in its deferred JS-rendering queue.
//
// Output layout (consumed by server/index.js):
// dist/index.html -> /
// dist/about/index.html -> /about
// dist/services/<slug>/index.html
// dist/404.html -> served with a real 404 status
//
// Run automatically as part of `npm run build`.
import { mkdirSync, readFileSync, writeFileSync } from 'fs'
import path from 'path'
import { fileURLToPath } from 'url'
feat(build): the copy is checked before a single page is built from it src/data is prose in a data structure, and nothing checked it. The long-form service pages make that dangerous in a specific way: their copy arrives as an owner-approved markdown sheet that MIXES DIRECTIONS TO THE WEBSITE MANAGER INTO THE COPY. "Do not promise that every number is always portable." "Keep this factual:" "Place an official 8x8 Work screenshot beside this section." Those lines look exactly like copy, and publishing one puts an internal instruction on a customer-facing page. scripts/lib/content.js decides whether the content layer is publishable, and prerender.js runs it before rendering anything, so every build enforces it: the pre-commit hook, npm run verify, and the Docker image build. It refuses a website-manager direction, an em dash, a U+FFFD, markdown or an HTML tag left in a string, an unknown block type, a section id that is not letter-first, unique and free of the layout's own ids, a section that does not open with its direct answer (unless it declares kind list or faq), a FAQ question with no answer, a link to a route or fragment that does not exist, an image whose src is missing from public/ or has no alt or no dimensions, and the missing benefits or idealFor list that the short layout maps without checking. A description over 160 characters is a note, not a failure: owner-approved copy is published as written. scripts/lib/routes.js is now the one route list. prerender.js built its own while src/routes.jsx built the router's, and nothing compared them: a route in one and not the other is never prerendered, so the server answers it with 404.html while the site's own navigation links to it. entry-server.jsx exports the router table so the build can compare the two. Proven by mutation, seventeen of them, each expecting exactly one finding and getting it: unknown block type, FAQ answer removed, answer moved below its list, duplicate id, digit-leading id, link to /services/contact-centre, #no-such-section, missing image file, image without dimensions, image without alt, em dash, a manager direction, markdown bold, U+FFFD, an HTML tag, missing h1, empty section. Both generated content modules pass unmutated. Against a real build: an em dash added to industries.js failed npm run build naming the field, and a /pricing route added to src/routes.jsx failed it naming the route. Closes #230. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:43:01 -05:00
import { render, routes as routerTable } from '../dist-ssr/entry-server.js'
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
import { services } from '../src/data/services.js'
import { industries } from '../src/data/industries.js'
fix(seo): the production sitemap had no dates, and the build context had secrets Three things, all in the path between this repository and the running image. Closes #225, #224 and #223. 1. THE PRODUCTION SITEMAP CARRIED NO LASTMOD AT ALL. Dates come from git history, and the image build cannot see git: .dockerignore excludes .git and node:alpine has no git binary. prerender.js read the failure into an empty catch commented "git unavailable or file untracked", so all 18 URLs came out undated while the build printed a success line. Local builds looked perfect, which is why nobody caught it. release.sh now computes the map where git exists, passes it as the SITEMAP_LASTMOD build arg, and then asks the built image whether its sitemap has dates, refusing to publish one that does not. prerender prints the count on every run, so "18 URLs, 0 dated" can never again read as success. The route-to-source map moved into scripts/lib/routes.js, where a service page now also counts its own content file, so editing one page's copy moves that page's date and no other. Proven: an image built with the arg carries 18 lastmod entries; a build with git deliberately unreadable and no arg reports "18 URLs, 0 carrying a lastmod" and warns. 2. THE DOCKER BUILD CONTEXT CARRIED CLIENT MATERIAL AND LIVE SECRETS. .drop/, zoho.md (the reCAPTCHA secret and the Zoho tokens), Levi.md and two 30 MB zips were all sent to the daemon on every build, along with four agent workspaces. The final image copies only built output, so none of it ever shipped, but one careless COPY would have changed that. Proven by listing the context from inside a throwaway image: before, all of it; after, none of it. 3. UNTRACKED FILES PASSED SILENTLY. docker build packs the working tree, so an untracked module the code imports produces an image that works and a tag that cannot rebuild it. release.sh now refuses while untracked files are present, and pre-commit's note counts them too. #223 also claimed post-commit hides a refused push. It does not: it printed "push was refused. The commit is safe locally and the branch is now ahead." during this batch. The issue was corrected on the tracker rather than acted on. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:58:37 -05:00
import { ROUTES as routes, lastModByRoute, routeDrift } from './lib/routes.js'
feat(build): the copy is checked before a single page is built from it src/data is prose in a data structure, and nothing checked it. The long-form service pages make that dangerous in a specific way: their copy arrives as an owner-approved markdown sheet that MIXES DIRECTIONS TO THE WEBSITE MANAGER INTO THE COPY. "Do not promise that every number is always portable." "Keep this factual:" "Place an official 8x8 Work screenshot beside this section." Those lines look exactly like copy, and publishing one puts an internal instruction on a customer-facing page. scripts/lib/content.js decides whether the content layer is publishable, and prerender.js runs it before rendering anything, so every build enforces it: the pre-commit hook, npm run verify, and the Docker image build. It refuses a website-manager direction, an em dash, a U+FFFD, markdown or an HTML tag left in a string, an unknown block type, a section id that is not letter-first, unique and free of the layout's own ids, a section that does not open with its direct answer (unless it declares kind list or faq), a FAQ question with no answer, a link to a route or fragment that does not exist, an image whose src is missing from public/ or has no alt or no dimensions, and the missing benefits or idealFor list that the short layout maps without checking. A description over 160 characters is a note, not a failure: owner-approved copy is published as written. scripts/lib/routes.js is now the one route list. prerender.js built its own while src/routes.jsx built the router's, and nothing compared them: a route in one and not the other is never prerendered, so the server answers it with 404.html while the site's own navigation links to it. entry-server.jsx exports the router table so the build can compare the two. Proven by mutation, seventeen of them, each expecting exactly one finding and getting it: unknown block type, FAQ answer removed, answer moved below its list, duplicate id, digit-leading id, link to /services/contact-centre, #no-such-section, missing image file, image without dimensions, image without alt, em dash, a manager direction, markdown bold, U+FFFD, an HTML tag, missing h1, empty section. Both generated content modules pass unmutated. Against a real build: an em dash added to industries.js failed npm run build naming the field, and a /pricing route added to src/routes.jsx failed it naming the route. Closes #230. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:43:01 -05:00
import { validateContent } from './lib/content.js'
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
const __dirname = path.dirname(fileURLToPath(import.meta.url))
const distDir = path.join(__dirname, '../dist')
feat(build): the copy is checked before a single page is built from it src/data is prose in a data structure, and nothing checked it. The long-form service pages make that dangerous in a specific way: their copy arrives as an owner-approved markdown sheet that MIXES DIRECTIONS TO THE WEBSITE MANAGER INTO THE COPY. "Do not promise that every number is always portable." "Keep this factual:" "Place an official 8x8 Work screenshot beside this section." Those lines look exactly like copy, and publishing one puts an internal instruction on a customer-facing page. scripts/lib/content.js decides whether the content layer is publishable, and prerender.js runs it before rendering anything, so every build enforces it: the pre-commit hook, npm run verify, and the Docker image build. It refuses a website-manager direction, an em dash, a U+FFFD, markdown or an HTML tag left in a string, an unknown block type, a section id that is not letter-first, unique and free of the layout's own ids, a section that does not open with its direct answer (unless it declares kind list or faq), a FAQ question with no answer, a link to a route or fragment that does not exist, an image whose src is missing from public/ or has no alt or no dimensions, and the missing benefits or idealFor list that the short layout maps without checking. A description over 160 characters is a note, not a failure: owner-approved copy is published as written. scripts/lib/routes.js is now the one route list. prerender.js built its own while src/routes.jsx built the router's, and nothing compared them: a route in one and not the other is never prerendered, so the server answers it with 404.html while the site's own navigation links to it. entry-server.jsx exports the router table so the build can compare the two. Proven by mutation, seventeen of them, each expecting exactly one finding and getting it: unknown block type, FAQ answer removed, answer moved below its list, duplicate id, digit-leading id, link to /services/contact-centre, #no-such-section, missing image file, image without dimensions, image without alt, em dash, a manager direction, markdown bold, U+FFFD, an HTML tag, missing h1, empty section. Both generated content modules pass unmutated. Against a real build: an em dash added to industries.js failed npm run build naming the field, and a /pricing route added to src/routes.jsx failed it naming the route. Closes #230. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:43:01 -05:00
// The route list and the router's own table both come from elsewhere now, so
// this file cannot disagree with either. See scripts/lib/routes.js.
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
// Tags the SEO component owns per-route. They are stripped from the template so
// Helmet's values replace them instead of duplicating them.
const TEMPLATE_TAGS_TO_STRIP = [
/<title>[\s\S]*?<\/title>\s*/,
/<meta name="description"[^>]*>\s*/,
/<meta property="og:[^"]*"[^>]*>\s*/g,
/<meta name="twitter:[^"]*"[^>]*>\s*/g,
/<!-- Open Graph fallback for crawlers that don't execute JavaScript -->\s*/,
/<!-- Twitter \/ X Card fallback -->\s*/,
]
// Metadata React renders inside the component tree. React 19 hoists these into
// <head> in the browser and in its streaming renderer, but renderToString leaves
// them inline, so the prerenderer performs the same hoist. Leaving them in <body>
// would put every title, canonical, and og: tag somewhere crawlers ignore.
fix(ui): React was throwing away the prerendered page on every route Suspected from the code while planning Batch 17, then confirmed in Chromium: every page logged React error #418, a hydration mismatch. React answers a mismatch by discarding the server DOM and re-rendering the page on the client. So the prerender ran, crawlers received it, and every visitor's browser threw it away and did the work again. Two causes, both ours: 1. prerender hoisted the JSON-LD scripts out of the body into <head>. React hoists only async scripts with a src, so on the client that script stays where its component renders it. The DOM and the client's first render therefore disagreed on every page that emits structured data. JSON-LD is valid anywhere in the document, so it now stays where React puts it. Title, meta and link tags are still hoisted, because React hoists those itself. 2. main.jsx rendered sonner's <Toaster> in the first client pass, and the server entry never rendered one, so the client expected a <section> the prerendered HTML did not have. It mounts after hydration instead, which costs nothing: a toast can only follow an interaction. Measured in a real browser, all 18 sitemap pages, before and after: hydration errors 18 to 0, other console and page errors 0. Each page keeps its server-rendered DOM (an h1 stamped before hydration survives), and still has exactly one head title and one canonical. Structured data is unchanged in substance: 27 JSON-LD blocks across 19 pages, and OAI-SearchBot still receives Service, BreadcrumbList and Organization on the contact-center page. Closes #226. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:37:33 -05:00
//
// JSON-LD is deliberately NOT in this list. React hoists only async scripts with
// a src, so on the client the ld+json script stays where its component renders
// it, in the body. Moving it to <head> here made the prerendered DOM disagree
// with the client's first render, and React threw out the whole prerendered page
// and re-rendered it (error #418, on every page). Structured data is valid
// anywhere in the document, so the honest fix is to leave it alone.
const HOISTABLE_TAGS = /<title[^>]*>[\s\S]*?<\/title>|<meta\b[^>]*?\/?>|<link\b[^>]*?\/?>/g
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
// An inline <svg> may carry its own <title>, and microdata rides in
// <meta itemprop> tags. Neither belongs in <head>: hoisting an SVG title gives
// the page two titles, which is the exact shape of defect this hoist exists to
// prevent. So SVG blocks are parked before the hoist and put back after.
const SVG_BLOCK = /<svg\b[\s\S]*?<\/svg>/gi
const SVG_TOKEN = 'svg'
// renderToString does not wait for Suspense: it emits the fallback and marks it.
// A page carrying one of these lost its content silently, and crawlers would
// receive the fallback as the page.
const SUSPENSE_MARKERS = ['<!--$!-->', '<!--$?-->']
// Substitute a marker that must appear exactly once, without regex replacement
// semantics. String.replace interprets `$&`, `$'` and `$$` inside the
// REPLACEMENT, and the replacement here is page copy, which is not ours to
// trust: one `$` before an escaped entity would inject markup into the page.
const replaceOnce = (text, marker, replacement, url) => {
const parts = text.split(marker)
if (parts.length !== 2) {
throw new Error(
`prerender: ${url}: expected exactly one ${marker} in the template, found ${parts.length - 1}.`,
)
}
return `${parts[0]}${replacement}${parts[1]}`
}
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
const buildPage = (template, url) => {
let html
try {
;({ html } = render(url))
} catch (error) {
// Without the route, the build fails with a stack trace and no clue which
// of the nineteen pages produced it.
throw new Error(`prerender: ${url} could not be rendered: ${error.message}`, { cause: error })
}
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
const svgs = []
const parked = html.replace(SVG_BLOCK, (svg) => `${SVG_TOKEN}${svgs.push(svg) - 1}${SVG_TOKEN}`)
const hoisted = []
const body = parked
.replace(HOISTABLE_TAGS, (tag) => {
if (/\bitemprop=/i.test(tag)) return tag
hoisted.push(tag)
return ''
})
.replace(new RegExp(`${SVG_TOKEN}(\\d+)${SVG_TOKEN}`, 'g'), (_, index) => svgs[Number(index)])
// Preload this route's own LCP image: the one React marked with a high fetch
// priority. Keying on `loading="eager"` matched the header logo, which is on
// every page, so every page preloaded the logo and no page preloaded its own
// hero. React writes the attribute camelCase in HTML, hence the /i.
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
const heroSrc = html
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
.match(/<img\b[^>]*\bfetchpriority="high"[^>]*>/i)?.[0]
?.match(/\bsrc="([^"]+)"/i)?.[1]
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
const head = []
if (heroSrc) {
head.push(`<link rel="preload" as="image" href="${heroSrc}" fetchpriority="high" />`)
}
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
// React emits its own preload for that hero; drop it so the hint above is not
// duplicated.
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
for (const tag of hoisted) {
if (/rel="preload"[^>]*as="image"/.test(tag)) continue
head.push(tag)
}
let page = template
for (const pattern of TEMPLATE_TAGS_TO_STRIP) {
page = page.replace(pattern, '')
}
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
page = replaceOnce(page, '</head>', ` ${head.join('\n ')}\n </head>`, url)
page = replaceOnce(page, '<div id="root"></div>', `<div id="root">${body}</div>`, url)
const marker = SUSPENSE_MARKERS.find((m) => page.includes(m))
if (marker) {
throw new Error(
`prerender: ${url} shipped a Suspense fallback (${marker}) instead of its content. ` +
'renderToString does not wait, so a lazy import or a Suspense boundary above this route ' +
'silently empties the page for every crawler.',
)
}
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
return page
}
const outputPathFor = (url) =>
url === '/'
? path.join(distDir, 'index.html')
: path.join(distDir, url, 'index.html')
const template = readFileSync(path.join(distDir, 'index.html'), 'utf8')
fix(build): the prerender preloaded the logo, hoisted any tag, and hid its errors Four faults in the one build step that decides what a crawler receives. Closes #233 and #232. 1. Every page preloaded /logo.png. The hero hint was keyed on the first `loading="eager"` image, and that is the header logo, on every page. So the image each page actually paints first was never preloaded, and React's own preload for it was being discarded as a duplicate. It now keys on the image React marked with a high fetch priority, matched case-insensitively, because React writes the attribute camelCase in HTML and a case-sensitive match would have quietly removed every preload instead. 2. The hoist moved ANY title, meta or link out of the body into <head>. An inline <svg><title> is a picture's label, and microdata rides in <meta itemprop>: both would have become page-level head tags the moment the long-form copy carried an icon with a title. SVG blocks are now parked before the hoist, and itemprop tags stay where they are. 3. `page.replace('</head>', body)` interprets `$&`, `$'` and `$$` INSIDE the replacement, and the replacement is page copy. React escapes & into &amp;, so any `$` immediately before an escaped character injected markup. Copy carries no `$` today; the next page with a price would have. Replacement is now a split and join, and each marker must appear exactly once. 4. A render error named no route, and a Suspense fallback shipped silently as an empty page. Both now fail the build and say which route. Also refuses to run over its own output: dist/index.html is both the template and the home page, so a second run without a rebuild gave every page two canonicals. Proven by mutation, each restored afterwards: a probe <svg><title> in the footer stays in the body on all 19 pages and no page gains a second title; a Suspense boundary fails the build naming the route; a second prerender run refuses; copy reading "Save $10 & more" reaches the page literally with one root div; and every page now preloads its own hero, with none preloading the logo. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:32:33 -05:00
// dist/index.html is both the template and the output for `/`, so running this
// script twice without a rebuild would treat a finished page as the template and
// give every page two canonicals and two of every head tag.
if (/rel="canonical"/.test(template)) {
throw new Error(
'prerender: dist/index.html already carries a canonical, so it is a rendered page rather than the ' +
'template. Run `vite build` before prerendering.',
)
}
feat(build): the copy is checked before a single page is built from it src/data is prose in a data structure, and nothing checked it. The long-form service pages make that dangerous in a specific way: their copy arrives as an owner-approved markdown sheet that MIXES DIRECTIONS TO THE WEBSITE MANAGER INTO THE COPY. "Do not promise that every number is always portable." "Keep this factual:" "Place an official 8x8 Work screenshot beside this section." Those lines look exactly like copy, and publishing one puts an internal instruction on a customer-facing page. scripts/lib/content.js decides whether the content layer is publishable, and prerender.js runs it before rendering anything, so every build enforces it: the pre-commit hook, npm run verify, and the Docker image build. It refuses a website-manager direction, an em dash, a U+FFFD, markdown or an HTML tag left in a string, an unknown block type, a section id that is not letter-first, unique and free of the layout's own ids, a section that does not open with its direct answer (unless it declares kind list or faq), a FAQ question with no answer, a link to a route or fragment that does not exist, an image whose src is missing from public/ or has no alt or no dimensions, and the missing benefits or idealFor list that the short layout maps without checking. A description over 160 characters is a note, not a failure: owner-approved copy is published as written. scripts/lib/routes.js is now the one route list. prerender.js built its own while src/routes.jsx built the router's, and nothing compared them: a route in one and not the other is never prerendered, so the server answers it with 404.html while the site's own navigation links to it. entry-server.jsx exports the router table so the build can compare the two. Proven by mutation, seventeen of them, each expecting exactly one finding and getting it: unknown block type, FAQ answer removed, answer moved below its list, duplicate id, digit-leading id, link to /services/contact-centre, #no-such-section, missing image file, image without dimensions, image without alt, em dash, a manager direction, markdown bold, U+FFFD, an HTML tag, missing h1, empty section. Both generated content modules pass unmutated. Against a real build: an em dash added to industries.js failed npm run build naming the field, and a /pricing route added to src/routes.jsx failed it naming the route. Closes #230. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:43:01 -05:00
// Content before pages. A page built from broken data is worse than no build:
// it looks finished. This is the check that keeps a website-manager direction
// out of the copy, and it runs on every build, including the image build.
const content = validateContent({ services, industries })
for (const warning of content.warnings) console.warn(`prerender: note: ${warning}`)
if (content.checked === 0) {
throw new Error('prerender: the content check examined nothing, which is not a pass. Did src/data fail to import?')
}
if (content.errors.length) {
console.error(`\nprerender: ${content.errors.length} content problem(s):`)
for (const error of content.errors) console.error(` ${error}`)
throw new Error('prerender: refusing to build pages from content that does not hold together.')
}
// The router and this script must agree on which pages exist. A route declared
// in src/routes.jsx and missing here is never written to dist/, and the server
// serves 404.html for it.
const drift = routeDrift(routerTable)
if (drift.length) {
throw new Error(
`prerender: ${drift.join(', ')} ${drift.length === 1 ? 'is a route' : 'are routes'} the router declares and this ` +
`build does not produce, so the server would answer ${drift.length === 1 ? 'it' : 'them'} with 404.html. ` +
'Add to STATIC_ROUTES in scripts/lib/routes.js.',
)
}
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
const written = []
for (const url of routes) {
const page = buildPage(template, url)
const outPath = outputPathFor(url)
mkdirSync(path.dirname(outPath), { recursive: true })
writeFileSync(outPath, page)
written.push([url, page.length])
}
// Dedicated 404 document. Rendering an unmatched path hits the catch-all route,
// which carries `noindex, follow` — now visible to crawlers in static HTML.
const notFoundPage = buildPage(template, '/__not_found__')
writeFileSync(path.join(distDir, '404.html'), notFoundPage)
written.push(['404.html', notFoundPage.length])
console.log(`\nPrerendered ${written.length} pages:`)
for (const [url, size] of written) {
console.log(` ${url.padEnd(42)} ${(size / 1024).toFixed(1)} KB`)
}
// --- sitemap.xml -------------------------------------------------------------
// Generated from the same route list that drives prerendering, so the sitemap can
// never drift out of sync with what the site actually serves.
const SITE_URL = 'https://queuenorth.com'
const PRIORITY = {
'/': '1.0',
'/services': '0.9',
'/contact': '0.9',
'/about': '0.8',
'/industries': '0.8',
'/support': '0.8',
'/privacy-policy': '0.3',
}
const CHANGEFREQ = { '/': 'weekly', '/privacy-policy': 'yearly' }
fix(seo): the production sitemap had no dates, and the build context had secrets Three things, all in the path between this repository and the running image. Closes #225, #224 and #223. 1. THE PRODUCTION SITEMAP CARRIED NO LASTMOD AT ALL. Dates come from git history, and the image build cannot see git: .dockerignore excludes .git and node:alpine has no git binary. prerender.js read the failure into an empty catch commented "git unavailable or file untracked", so all 18 URLs came out undated while the build printed a success line. Local builds looked perfect, which is why nobody caught it. release.sh now computes the map where git exists, passes it as the SITEMAP_LASTMOD build arg, and then asks the built image whether its sitemap has dates, refusing to publish one that does not. prerender prints the count on every run, so "18 URLs, 0 dated" can never again read as success. The route-to-source map moved into scripts/lib/routes.js, where a service page now also counts its own content file, so editing one page's copy moves that page's date and no other. Proven: an image built with the arg carries 18 lastmod entries; a build with git deliberately unreadable and no arg reports "18 URLs, 0 carrying a lastmod" and warns. 2. THE DOCKER BUILD CONTEXT CARRIED CLIENT MATERIAL AND LIVE SECRETS. .drop/, zoho.md (the reCAPTCHA secret and the Zoho tokens), Levi.md and two 30 MB zips were all sent to the daemon on every build, along with four agent workspaces. The final image copies only built output, so none of it ever shipped, but one careless COPY would have changed that. Proven by listing the context from inside a throwaway image: before, all of it; after, none of it. 3. UNTRACKED FILES PASSED SILENTLY. docker build packs the working tree, so an untracked module the code imports produces an image that works and a tag that cannot rebuild it. release.sh now refuses while untracked files are present, and pre-commit's note counts them too. #223 also claimed post-commit hides a refused push. It does not: it printed "push was refused. The commit is safe locally and the branch is now ahead." during this batch. The issue was corrected on the tracker rather than acted on. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:58:37 -05:00
// Dates come from git where git exists, and from the map release.sh injects
// where it does not. See scripts/lib/routes.js.
const lastmodByRoute = lastModByRoute()
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
const entries = routes.map((url) => {
const loc = url === '/' ? SITE_URL : `${SITE_URL}${url}`
fix(seo): the production sitemap had no dates, and the build context had secrets Three things, all in the path between this repository and the running image. Closes #225, #224 and #223. 1. THE PRODUCTION SITEMAP CARRIED NO LASTMOD AT ALL. Dates come from git history, and the image build cannot see git: .dockerignore excludes .git and node:alpine has no git binary. prerender.js read the failure into an empty catch commented "git unavailable or file untracked", so all 18 URLs came out undated while the build printed a success line. Local builds looked perfect, which is why nobody caught it. release.sh now computes the map where git exists, passes it as the SITEMAP_LASTMOD build arg, and then asks the built image whether its sitemap has dates, refusing to publish one that does not. prerender prints the count on every run, so "18 URLs, 0 dated" can never again read as success. The route-to-source map moved into scripts/lib/routes.js, where a service page now also counts its own content file, so editing one page's copy moves that page's date and no other. Proven: an image built with the arg carries 18 lastmod entries; a build with git deliberately unreadable and no arg reports "18 URLs, 0 carrying a lastmod" and warns. 2. THE DOCKER BUILD CONTEXT CARRIED CLIENT MATERIAL AND LIVE SECRETS. .drop/, zoho.md (the reCAPTCHA secret and the Zoho tokens), Levi.md and two 30 MB zips were all sent to the daemon on every build, along with four agent workspaces. The final image copies only built output, so none of it ever shipped, but one careless COPY would have changed that. Proven by listing the context from inside a throwaway image: before, all of it; after, none of it. 3. UNTRACKED FILES PASSED SILENTLY. docker build packs the working tree, so an untracked module the code imports produces an image that works and a tag that cannot rebuild it. release.sh now refuses while untracked files are present, and pre-commit's note counts them too. #223 also claimed post-commit hides a refused push. It does not: it printed "push was refused. The commit is safe locally and the branch is now ahead." during this batch. The issue was corrected on the tracker rather than acted on. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:58:37 -05:00
const lastmod = lastmodByRoute[url] ?? null
feat(seo): publish privacy policy, remove street address, prerender all routes (batch 0.9.3) Client directive (Levi Halford, 2026-08-01) ahead of Google/Meta lead forms. Privacy policy: - Publish approved policy verbatim at /privacy-policy (src/data/privacyPolicy.js is the single source of truth; 292/292 source lines verified present) - Privacy Policy link in the footer of every page - Effective/Last Updated 2026-07-31, [email protected] as mailto Remove St. Petersburg street address from every surface named in the brief: footer, contact page, schema markup, SEO metadata, Google Maps links. Collapse ProfessionalService + Organization schema into a single Organization with areaServed: United States; drop geo coordinates, priceRange, openingHours. Add the approved US-coverage sentence to About. No replacement address. Crawler visibility (the site previously served 0 bytes of body HTML without JS): - Prerender all 19 routes at build time via src/entry-server.jsx + scripts/prerender.js - Hoist title/meta/canonical/JSON-LD into <head>; renderToString does not do this and react-helmet-async's context is empty under React 19 - Serve prerendered HTML; return a real 404 for unknown paths instead of 200 - Hydrate instead of discarding the prerendered markup SEO/perf: - Titles <=60 and descriptions <=160 chars across all pages - Add BreadcrumbList to interior pages, WebSite to home - Generate sitemap.xml from the route list with git-derived lastmod - 301 duplicate URL forms (trailing slash, //, /index.html), preserving query - Immutable caching for content-hashed assets; no-cache for HTML - Split the 522 KB bundle into app/react-vendor/router/icons - loading/decoding/fetchpriority + per-route hero preload; drop unused asset Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-01 01:45:52 -05:00
return [
' <url>',
` <loc>${loc}</loc>`,
lastmod ? ` <lastmod>${lastmod}</lastmod>` : null,
` <changefreq>${CHANGEFREQ[url] || 'monthly'}</changefreq>`,
` <priority>${PRIORITY[url] || '0.7'}</priority>`,
' </url>',
]
.filter(Boolean)
.join('\n')
})
const sitemap = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
...entries,
'</urlset>',
'',
].join('\n')
writeFileSync(path.join(distDir, 'sitemap.xml'), sitemap)
fix(seo): the production sitemap had no dates, and the build context had secrets Three things, all in the path between this repository and the running image. Closes #225, #224 and #223. 1. THE PRODUCTION SITEMAP CARRIED NO LASTMOD AT ALL. Dates come from git history, and the image build cannot see git: .dockerignore excludes .git and node:alpine has no git binary. prerender.js read the failure into an empty catch commented "git unavailable or file untracked", so all 18 URLs came out undated while the build printed a success line. Local builds looked perfect, which is why nobody caught it. release.sh now computes the map where git exists, passes it as the SITEMAP_LASTMOD build arg, and then asks the built image whether its sitemap has dates, refusing to publish one that does not. prerender prints the count on every run, so "18 URLs, 0 dated" can never again read as success. The route-to-source map moved into scripts/lib/routes.js, where a service page now also counts its own content file, so editing one page's copy moves that page's date and no other. Proven: an image built with the arg carries 18 lastmod entries; a build with git deliberately unreadable and no arg reports "18 URLs, 0 carrying a lastmod" and warns. 2. THE DOCKER BUILD CONTEXT CARRIED CLIENT MATERIAL AND LIVE SECRETS. .drop/, zoho.md (the reCAPTCHA secret and the Zoho tokens), Levi.md and two 30 MB zips were all sent to the daemon on every build, along with four agent workspaces. The final image copies only built output, so none of it ever shipped, but one careless COPY would have changed that. Proven by listing the context from inside a throwaway image: before, all of it; after, none of it. 3. UNTRACKED FILES PASSED SILENTLY. docker build packs the working tree, so an untracked module the code imports produces an image that works and a tag that cannot rebuild it. release.sh now refuses while untracked files are present, and pre-commit's note counts them too. #223 also claimed post-commit hides a refused push. It does not: it printed "push was refused. The commit is safe locally and the branch is now ahead." during this batch. The issue was corrected on the tracker rather than acted on. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 04:58:37 -05:00
// Say how many pages carry a date, every time. The production sitemap carried
// none at all for months: the image build has no git, and the failure to read
// it was swallowed. Silence is what let that run.
const dated = routes.filter((url) => lastmodByRoute[url]).length
console.log(`\nGenerated sitemap.xml with ${routes.length} URLs, ${dated} carrying a lastmod`)
if (dated < routes.length) {
console.warn(
`prerender: ${routes.length - dated} URL(s) have no lastmod. git history is not readable here, which is ` +
'normal inside the image build. Pass SITEMAP_LASTMOD, as scripts/release.sh does.',
)
}