DocumentationBuild. Deploy. Operate.
DocsReference

SEO & Static Assets Across Thin Clients

On this page 7 sections

QSCS aggressively caches and replicates dynamic responses, but search-engine crawlers (Googlebot, Bingbot, DuckDuckBot, AI agents fetching llms.txt, etc.) also expect a stable set of static URLs to exist on every public endpoint they hit: /robots.txt, /sitemap.xml, /favicon.ico, Open Graph images, apple-touch-icon.png, PWA manifests, and any HTML pages you want crawled directly.

Because a QSCS domain is normally served by multiple thin clients distributed across regions, a crawler hitting example.com in Tokyo and again in London may reach different thin clients on consecutive requests. If those nodes do not serve byte-identical static assets, you will see:

  • Duplicate / inconsistent indexing, Googlebot sees a sitemap on one node and a 404 on another, and may de-prioritise the whole property.
  • Favicon flicker, browsers fall back to a generic icon when one node returns the favicon and another returns 404.
  • Broken robots.txt directives, a crawler may obey one node's rules and ignore another's, blocking or exposing the wrong paths.
  • Validation failures for Search Console, Bing Webmaster Tools, and the App Store domain-association files (/.well-known/apple-app-site-association, /.well-known/assetlinks.json).
The rule: every static asset a crawler or browser may request must exist, byte-for-byte identical, on every thin client that serves the public hostname. QSCS does not replicate the static-pages directory for you, that is an operations concern, addressed below.

1. The [Static Pages] section in qscs.conf

QSCS exposes a per-node static directory through one configuration key:

[Static Pages]
Path = /opt/qscs/pages

When Path is set and a request would otherwise miss the cache and the origin, QSCS attempts to satisfy it from disk. The on-disk layout is host-scoped, one subdirectory per public hostname the node serves:

/opt/qscs/pages/
├── example.com/
│   ├── index.html
│   ├── about.html
│   ├── robots.txt
│   ├── sitemap.xml
│   ├── favicon.ico
│   ├── apple-touch-icon.png
│   ├── og-image.png
│   ├── llms.txt
│   └── .well-known/
│       ├── apple-app-site-association
│       └── assetlinks.json
└── docs.example.com/
    ├── index.html
    ├── robots.txt
    └── favicon.ico

URL resolution rules:

  • GET / → <Path>/<Host>/index.html
  • GET /about → <Path>/<Host>/about.html (extension auto-appended if the last segment has none)
  • GET /sitemap.xml → <Path>/<Host>/sitemap.xml
  • GET /favicon.ico → <Path>/<Host>/favicon.ico
  • Path traversal (..) and null bytes are rejected; the canonical resolved path must stay under Path.

Content-Type is inferred from the extension (.html, .json, .css, .js) and falls back to text/plain; charset=utf-8 for everything else. If you need a specific binary type for assets like .ico, .png, .webp, or .xml, you currently have two options:

  • Put a CDN or thin reverse proxy in front (the Caddy/nginx config in the WASM page already covers MIME-type overrides), or
  • Serve via the origin and let QSCS cache the response, the cache layer preserves the upstream Content-Type verbatim.

2. The crawler-critical asset checklist

At minimum, for any public hostname, place the following files in <Path>/<Host>/ on every thin client:

robots.txt
Must be byte-identical across nodes. If you generate it from a template, regenerate from the same source on every node, not per-node.
sitemap.xml (and any sub-sitemaps)
Reference absolute URLs (https://example.com/...), never node-local hostnames. Add the Sitemap: line to robots.txt so crawlers find it.
favicon.ico, favicon-16.png, favicon-32.png, apple-touch-icon.png
Browsers and link-preview bots will probe several of these. A missing icon on one node manifests as flicker for users whose next request hits that node.
og-image.png / Twitter card image
Referenced from <meta property="og:image">; social crawlers (Slack, Discord, Twitter) cache these for days. An inconsistent image means link previews change unpredictably.
llms.txt (optional)
Increasingly used by AI agents and search crawlers. Same rule: identical on every node.
.well-known/*
App-association files for iOS, Android App Links, Web Credentials, and ACME challenges. Any inconsistency breaks deep links or certificate issuance.
Pre-rendered HTML for SEO-critical landing pages
If your SPA is hydrated client-side (typical with the WASM bootstrap), drop a server-rendered index.html, /pricing.html, /about.html, etc., so crawlers that do not execute JavaScript still see the canonical title, description, and Open Graph tags.

3. Keeping every thin client uniform

QSCS replicates dynamic responses through the encrypted control plane, but the static-pages directory is plain files on each node's filesystem. Pick one of the strategies below and apply it to every node that serves the public hostname.

Treat one host (or a build server) as the canonical source and push to every thin client over SSH on a timer:

# /usr/local/bin/qscs-pages-sync.sh
#!/bin/sh
set -eu
SRC=/srv/qscs-pages/
NODES="thin-eu1 thin-eu2 thin-us1 thin-us2 thin-ap1 thin-ap2"
for N in $NODES; do
    rsync -az --delete --chown=qscs:qscs \
        -e "ssh -i /etc/qscs/deploy.key" \
        "$SRC" "deploy@$N:/opt/qscs/pages/"
done

Schedule with systemd or cron:

# /etc/systemd/system/qscs-pages-sync.timer
[Unit]
Description=Sync QSCS static pages to all thin clients
[Timer]
OnBootSec=2min
OnUnitActiveSec=5min
[Install]
WantedBy=timers.target
Use --delete so a file removed on the source is removed everywhere. Without it, an old robots.txt can linger on one node and quietly poison crawl behaviour for weeks.

3b. Git pull with a deploy hook

Store /opt/qscs/pages as a git repo and pull on every node from a CI deploy hook or a systemd path-unit. This gives you an audit trail and an obvious rollback path. The repo layout mirrors the on-disk structure exactly (<repo>/<host>/...).

3c. Object store (S3 / R2 / GCS) + periodic pull

For globally distributed fleets where SSH-from-a-bastion is awkward, push to an object store from CI and let each node pull on a timer:

aws s3 sync s3://my-qscs-pages/ /opt/qscs/pages/ --delete

Pin a specific version tag in the bucket so partial rollouts can be inspected before "promotion".

3d. NFS / shared filesystem (only for co-located nodes)

If all your thin clients live in the same VPC/region, an NFS or EFS mount at /opt/qscs/pages guarantees uniformity by definition. Do not do this across regions, the latency penalty defeats the point of having regional thin clients.

4. Favicons across multiple hostnames

If a single thin client serves more than one hostname (e.g. example.com and blog.example.com), each host gets its own subdirectory under Path and therefore its own favicon. Keep the hostname-to-icon mapping explicit:

/opt/qscs/pages/example.com/favicon.ico        # main brand
/opt/qscs/pages/example.com/apple-touch-icon.png
/opt/qscs/pages/blog.example.com/favicon.ico   # blog sub-brand (or symlink)
/opt/qscs/pages/docs.example.com/favicon.ico   # docs sub-brand (or symlink)

If every host should share the same icon, use a symlink, but make sure your sync mechanism preserves it (rsync -a does, plain cp -r may not).

Also reference the icon explicitly from each host's index.html:

<link rel="icon" href="/favicon.ico" sizes="any">
<link rel="icon" type="image/png" sizes="32x32" href="/favicon-32.png">
<link rel="apple-touch-icon" href="/apple-touch-icon.png">

5. Verifying uniformity

A simple shell check, run from any operator workstation, confirms that every node returns the same bytes for the canonical SEO files:

# Compare a hash of robots.txt / sitemap.xml / favicon.ico across nodes
for N in thin-eu1 thin-eu2 thin-us1 thin-us2 thin-ap1 thin-ap2; do
    for F in robots.txt sitemap.xml favicon.ico; do
        H=$(curl -sk --resolve example.com:443:$(getent hosts "$N" | awk '{print $1}') \
              "https://example.com/$F" | sha256sum | cut -c1-12)
        printf "%-12s %-14s %s\n" "$N" "$F" "$H"
    done
done

Every column should show the same hash. Anything else means a node is out of sync, re-run your sync mechanism and investigate before the next crawl cycle.

6. A worked example

Suppose example.com is served by six thin clients (thin-eu1/2, thin-us1/2, thin-ap1/2) and a single canonical source on the build server at /srv/qscs-pages/example.com/. Each thin client has:

# /etc/qscs/qscs.conf  (excerpt)
[QSCS Daemon]
Port      = 8080
Domain    = example.com
WasmPath  = /var/www/qscs-substrate.wasm

[Static Pages]
Path = /opt/qscs/pages

The build server's qscs-pages-sync.service runs every five minutes and rsyncs /srv/qscs-pages/ → /opt/qscs/pages/ on each node. Crawlers reach https://example.com/robots.txt via whichever thin client is closest; all six return identical bytes; Search Console reports a single canonical sitemap and no inconsistency warnings.

Once the sync is in place, adding a new region is purely additive: spin up a new thin client, append its hostname to the NODES list (or its region to the S3 puller), and the SEO surface area is automatically uniform from the moment the node accepts its first request.

7. Common pitfalls

  • Forgetting to set [Static Pages] Path, the key is empty by default, which disables static serving entirely. Without it, even a correctly-populated directory is invisible.
  • Per-node dynamic generation of robots.txt or sitemap.xml, if each node generates them on its own clock, they will differ. Generate once, distribute everywhere.
  • Stale --delete-less syncs, old files accumulate. Always sync with delete semantics.
  • Mismatched MIME types for favicon.ico, some operators put it behind a proxy that re-types it as text/html. Browsers reject it silently.
  • Hostname mismatch, the directory must be named after the Host header the client sends, not the node's own hostname.
  • Symlinks broken by deploy tooling, copy with rsync -a or cp -a, not cp -r.

See also: Configuration Reference, Browser Substrate & WASM Bootstrap, Recommended Topologies.

Need a hand with your deployment?Contact support ↗Back to top ↑