Files
thptqg2017/docs/deployment-guide.md
T
tiennm99 3fd137a233 refactor(web): rebuild the frontend on SvelteKit, TypeScript and Tailwind
The app was a single-entry React SPA: one index.html that the assembler
copied to every dataset path, with a hand-rolled router resolving the
dataset from window.location and the title patched in at runtime because
one file had to serve every route.

SvelteKit prerenders a real page per route instead. The entry generator
in routes/[dataset]/+page.ts reads the same datasets.json the assembler
does, so the set of pages and the set of databases cannot drift apart,
and each dataset page ships its own <title> and description.

Everything framework-free moved across unchanged in behaviour and became
typed: the admission blocks, subject list, query classifier and SQL
presets. `Student` now mirrors the 22-column table, so a mistyped column
name fails the build rather than rendering blank.

Tailwind replaces the stylesheet. Theme tokens are CSS variables, which
keeps dark mode a single block of overrides rather than a `dark:` variant
on every class. The score tiers stay hand-written CSS: the class is
chosen at runtime from a score, and no utility generator can see that.

Adds the tests the frontend never had, over the three modules where a
silent wrong answer is possible — most importantly that toAscii here
folds exactly as ToAscii does in the parser, which is what makes
accent-insensitive search find anything.

The assembler stops copying index.html per dataset and checks that the
build prerendered each one instead. The database presence, raw-artifact
and idempotence guards are untouched.

Verified: 25 tests, ESLint and svelte-check clean, and a full
`assemble site` producing /thptqg/, /thptqg/2016/, /thptqg/2017/ and
404.html with absolute asset URLs. Not verified in a browser — this
machine has none.
2026-08-14 11:38:29 +07:00

4.2 KiB
Raw Blame History

Deployment Guide

Deploys to GitHub Pages via .github/workflows/deploy-pages.yml. Every push to main rebuilds both datasets and redeploys the whole site.

One-time setup: Settings → Pages → Source: GitHub Actions.

What the workflow does

  1. Checkout, Go toolchain, Node 24, npm ci in web/
  2. Parser, crawler and assembler test suites, web lint, govulncheck over all three modules
  3. go -C assembler run ./cmd/assemble — the whole pipeline: compile the parser, build and verify each database, compress it into .build/public/db/, run the web build, assemble _site/
  4. actions/upload-pages-artifact + actions/deploy-pages

The database build dominates the runtime: roughly 348 MB of Excel is parsed on every deploy.

Resulting URLs

https://<user>.github.io/thptqg/
https://<user>.github.io/thptqg/2016/
https://<user>.github.io/thptqg/2017/

/thptqg/2017/old/ and /thptqg/2017/old2/ were the pre-flattening URLs for the two removed 2017 archives. They are no longer served; like any unknown path they now render the hub via 404.html.

Local reproduction

(cd web && npm ci)
go -C assembler run ./cmd/assemble
npx serve _site

To rebuild a single dataset, or only the site:

go -C assembler run ./cmd/assemble db 2017
go -C assembler run ./cmd/assemble site

Base path

svelte.config.js sets paths.base: "/thptqg". If you fork under a different repo name, update it to match — assets are referenced absolutely, so a mismatch shows up as a blank page with 404s on /_app/....

Adding a dataset

  1. Put the Excel files in data/<id>/
  2. Add parser/configs/<id>.yml with the parse rules — sheet mode, column indices, SBD validation, header tokens, blank-row stripping. No SQL: the schema is canonical and lives in parser/internal/schema/schema.go
  3. Add an entry to datasets.json — id, expectedRows, dbSizeMb
  4. Add its presentation to CONTENT in web/src/datasets.js

Nothing else. The assembler and the router both read the registry, and the frontend adapts to whichever columns the dataset populates. The last two steps check each other, so forgetting either fails rather than half-working.

Why no uncompressed database can ship

The assembler deletes the source once compression succeeds, so the raw file does not survive the build, and it then fails the job if any .db, .db-journal, .db-wal or .db-shm reached the output.

Both guards exist because the previous pipeline wrote a 100+ MB uncompressed database into the source tree and relied on an rm step to keep it out of the artifact — one missing line away from publishing it.

Notes

  • The gzipped database is not cacheable across deploys. Every rebuild produces a different .db.gz, because SQLite does not lay pages out deterministically. First-visit users pay the full download; later visits hit browser cache until the next deploy.
  • GitHub Pages caps individual files at 100 MB. The largest gzipped database is about 48 MB. Uncompressed they run 135–234 MB and would not fit — which is why the browser decompresses via DecompressionStream.
  • No server-side compression is assumed. The app fetches the .gz bytes directly rather than relying on Content-Encoding: gzip; Pages does not reliably compress arbitrary paths on the fly.
  • Total artifact is about 93 MB, well inside the 1 GB site limit.

Rollback

Deploys are stateless snapshots. Revert the commit on main and push; the next run rebuilds the older state. There is no data to migrate.

Troubleshooting

Symptom Typical cause
Blank page, 404 on assets paths.base in svelte.config.js does not match the repo name
Failed to fetch database: 404 Dataset id in datasets.json does not match the file in db/
A route 404s The site step did not run, or the id is missing from datasets.json
WASM fails to load sql.js.org unreachable — self-host sql-wasm.wasm and update SQL_WASM_URL in lib/sqlite.svelte.ts
Deploy fails on assembly An uncompressed database artefact reached the output; the error names the files
Missing rows after a data update Unknown Excel header — check the per-file row counts the parser prints