SEO & Static Assets Across Thin Clients
On this page 7 sections
QSCS aggressively caches and replicates dynamic responses, but search-engine
crawlers (Googlebot, Bingbot, DuckDuckBot, AI agents fetching llms.txt,
etc.) also expect a stable set of static URLs to exist on every public
endpoint they hit: /robots.txt, /sitemap.xml,
/favicon.ico, Open Graph images, apple-touch-icon.png,
PWA manifests, and any HTML pages you want crawled directly.
Because a QSCS domain is normally served by multiple thin clients
distributed across regions, a crawler hitting example.com in Tokyo and
again in London may reach different thin clients on consecutive requests.
If those nodes do not serve byte-identical static assets, you will see:
- Duplicate / inconsistent indexing, Googlebot sees a sitemap on one node and a 404 on another, and may de-prioritise the whole property.
- Favicon flicker, browsers fall back to a generic icon when one node returns the favicon and another returns 404.
- Broken
robots.txtdirectives, a crawler may obey one node's rules and ignore another's, blocking or exposing the wrong paths. - Validation failures for Search Console, Bing Webmaster Tools,
and the App Store domain-association files
(
/.well-known/apple-app-site-association,/.well-known/assetlinks.json).
1. The [Static Pages] section in qscs.conf
QSCS exposes a per-node static directory through one configuration key:
[Static Pages]
Path = /opt/qscs/pages
When Path is set and a request would otherwise miss the cache and
the origin, QSCS attempts to satisfy it from disk. The on-disk layout is
host-scoped, one subdirectory per public hostname the node
serves:
/opt/qscs/pages/
├── example.com/
│ ├── index.html
│ ├── about.html
│ ├── robots.txt
│ ├── sitemap.xml
│ ├── favicon.ico
│ ├── apple-touch-icon.png
│ ├── og-image.png
│ ├── llms.txt
│ └── .well-known/
│ ├── apple-app-site-association
│ └── assetlinks.json
└── docs.example.com/
├── index.html
├── robots.txt
└── favicon.ico
URL resolution rules:
GET /→<Path>/<Host>/index.htmlGET /about→<Path>/<Host>/about.html(extension auto-appended if the last segment has none)GET /sitemap.xml→<Path>/<Host>/sitemap.xmlGET /favicon.ico→<Path>/<Host>/favicon.ico- Path traversal (
..) and null bytes are rejected; the canonical resolved path must stay underPath.
Content-Type is inferred from the extension
(.html, .json, .css, .js) and
falls back to text/plain; charset=utf-8 for everything else. If you
need a specific binary type for assets like .ico, .png,
.webp, or .xml, you currently have two options:
- Put a CDN or thin reverse proxy in front (the Caddy/nginx config in the WASM page already covers MIME-type overrides), or
- Serve via the origin and let QSCS cache the response, the cache layer
preserves the upstream
Content-Typeverbatim.
2. The crawler-critical asset checklist
At minimum, for any public hostname, place the following files in
<Path>/<Host>/ on every thin client:
robots.txt- Must be byte-identical across nodes. If you generate it from a template, regenerate from the same source on every node, not per-node.
sitemap.xml(and any sub-sitemaps)- Reference absolute URLs (
https://example.com/...), never node-local hostnames. Add theSitemap:line torobots.txtso crawlers find it. favicon.ico,favicon-16.png,favicon-32.png,apple-touch-icon.png- Browsers and link-preview bots will probe several of these. A missing icon on one node manifests as flicker for users whose next request hits that node.
og-image.png/ Twitter card image- Referenced from
<meta property="og:image">; social crawlers (Slack, Discord, Twitter) cache these for days. An inconsistent image means link previews change unpredictably. llms.txt(optional)- Increasingly used by AI agents and search crawlers. Same rule: identical on every node.
.well-known/*- App-association files for iOS, Android App Links, Web Credentials, and ACME challenges. Any inconsistency breaks deep links or certificate issuance.
- Pre-rendered HTML for SEO-critical landing pages
- If your SPA is hydrated client-side (typical with the
WASM bootstrap),
drop a server-rendered
index.html,/pricing.html,/about.html, etc., so crawlers that do not execute JavaScript still see the canonical title, description, and Open Graph tags.
3. Keeping every thin client uniform
QSCS replicates dynamic responses through the encrypted control plane, but the static-pages directory is plain files on each node's filesystem. Pick one of the strategies below and apply it to every node that serves the public hostname.
3a. rsync from a single source of truth (recommended for small fleets)
Treat one host (or a build server) as the canonical source and push to every thin client over SSH on a timer:
# /usr/local/bin/qscs-pages-sync.sh
#!/bin/sh
set -eu
SRC=/srv/qscs-pages/
NODES="thin-eu1 thin-eu2 thin-us1 thin-us2 thin-ap1 thin-ap2"
for N in $NODES; do
rsync -az --delete --chown=qscs:qscs \
-e "ssh -i /etc/qscs/deploy.key" \
"$SRC" "deploy@$N:/opt/qscs/pages/"
done
Schedule with systemd or cron:
# /etc/systemd/system/qscs-pages-sync.timer
[Unit]
Description=Sync QSCS static pages to all thin clients
[Timer]
OnBootSec=2min
OnUnitActiveSec=5min
[Install]
WantedBy=timers.target
--delete so a file removed on the source is removed
everywhere. Without it, an old robots.txt can linger on one node
and quietly poison crawl behaviour for weeks.
3b. Git pull with a deploy hook
Store /opt/qscs/pages as a git repo and pull on every node from
a CI deploy hook or a systemd path-unit. This gives you an audit trail and an
obvious rollback path. The repo layout mirrors the on-disk structure
exactly (<repo>/<host>/...).
3c. Object store (S3 / R2 / GCS) + periodic pull
For globally distributed fleets where SSH-from-a-bastion is awkward, push to an object store from CI and let each node pull on a timer:
aws s3 sync s3://my-qscs-pages/ /opt/qscs/pages/ --delete
Pin a specific version tag in the bucket so partial rollouts can be inspected before "promotion".
3d. NFS / shared filesystem (only for co-located nodes)
If all your thin clients live in the same VPC/region, an NFS or EFS mount at
/opt/qscs/pages guarantees uniformity by definition. Do not
do this across regions, the latency penalty defeats the point of having
regional thin clients.
4. Favicons across multiple hostnames
If a single thin client serves more than one hostname (e.g.
example.com and blog.example.com), each host gets its
own subdirectory under Path and therefore its own favicon. Keep the
hostname-to-icon mapping explicit:
/opt/qscs/pages/example.com/favicon.ico # main brand
/opt/qscs/pages/example.com/apple-touch-icon.png
/opt/qscs/pages/blog.example.com/favicon.ico # blog sub-brand (or symlink)
/opt/qscs/pages/docs.example.com/favicon.ico # docs sub-brand (or symlink)
If every host should share the same icon, use a symlink, but make sure your
sync mechanism preserves it (rsync -a does, plain
cp -r may not).
Also reference the icon explicitly from each host's index.html:
<link rel="icon" href="/favicon.ico" sizes="any">
<link rel="icon" type="image/png" sizes="32x32" href="/favicon-32.png">
<link rel="apple-touch-icon" href="/apple-touch-icon.png">
5. Verifying uniformity
A simple shell check, run from any operator workstation, confirms that every node returns the same bytes for the canonical SEO files:
# Compare a hash of robots.txt / sitemap.xml / favicon.ico across nodes
for N in thin-eu1 thin-eu2 thin-us1 thin-us2 thin-ap1 thin-ap2; do
for F in robots.txt sitemap.xml favicon.ico; do
H=$(curl -sk --resolve example.com:443:$(getent hosts "$N" | awk '{print $1}') \
"https://example.com/$F" | sha256sum | cut -c1-12)
printf "%-12s %-14s %s\n" "$N" "$F" "$H"
done
done
Every column should show the same hash. Anything else means a node is out of sync, re-run your sync mechanism and investigate before the next crawl cycle.
6. A worked example
Suppose example.com is served by six thin clients
(thin-eu1/2, thin-us1/2, thin-ap1/2) and a
single canonical source on the build server at
/srv/qscs-pages/example.com/. Each thin client has:
# /etc/qscs/qscs.conf (excerpt)
[QSCS Daemon]
Port = 8080
Domain = example.com
WasmPath = /var/www/qscs-substrate.wasm
[Static Pages]
Path = /opt/qscs/pages
The build server's qscs-pages-sync.service runs every five
minutes and rsyncs /srv/qscs-pages/ →
/opt/qscs/pages/ on each node. Crawlers reach
https://example.com/robots.txt via whichever thin client is closest;
all six return identical bytes; Search Console reports a single canonical
sitemap and no inconsistency warnings.
NODES list (or its
region to the S3 puller), and the SEO surface area is automatically uniform
from the moment the node accepts its first request.
7. Common pitfalls
- Forgetting to set
[Static Pages] Path, the key is empty by default, which disables static serving entirely. Without it, even a correctly-populated directory is invisible. - Per-node dynamic generation of
robots.txtorsitemap.xml, if each node generates them on its own clock, they will differ. Generate once, distribute everywhere. - Stale
--delete-less syncs, old files accumulate. Always sync with delete semantics. - Mismatched MIME types for
favicon.ico, some operators put it behind a proxy that re-types it astext/html. Browsers reject it silently. - Hostname mismatch, the directory must be named after the Host header the client sends, not the node's own hostname.
- Symlinks broken by deploy tooling, copy with
rsync -aorcp -a, notcp -r.
See also: Configuration Reference, Browser Substrate & WASM Bootstrap, Recommended Topologies.