The app was a single-entry React SPA: one index.html that the assembler copied to every dataset path, with a hand-rolled router resolving the dataset from window.location and the title patched in at runtime because one file had to serve every route. SvelteKit prerenders a real page per route instead. The entry generator in routes/[dataset]/+page.ts reads the same datasets.json the assembler does, so the set of pages and the set of databases cannot drift apart, and each dataset page ships its own <title> and description. Everything framework-free moved across unchanged in behaviour and became typed: the admission blocks, subject list, query classifier and SQL presets. `Student` now mirrors the 22-column table, so a mistyped column name fails the build rather than rendering blank. Tailwind replaces the stylesheet. Theme tokens are CSS variables, which keeps dark mode a single block of overrides rather than a `dark:` variant on every class. The score tiers stay hand-written CSS: the class is chosen at runtime from a score, and no utility generator can see that. Adds the tests the frontend never had, over the three modules where a silent wrong answer is possible — most importantly that toAscii here folds exactly as ToAscii does in the parser, which is what makes accent-insensitive search find anything. The assembler stops copying index.html per dataset and checks that the build prerendered each one instead. The database presence, raw-artifact and idempotence guards are untouched. Verified: 25 tests, ESLint and svelte-check clean, and a full `assemble site` producing /thptqg/, /thptqg/2016/, /thptqg/2017/ and 404.html with absolute asset URLs. Not verified in a browser — this machine has none.
3.9 KiB
thptqg
Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high
school graduation exam. Client-side SQL (sql.js) over a SQLite database built
from the published .xls/.xlsx score files by the Go parser module. Where
those files come from: data pipeline.
Live at tiennm99.github.io/thptqg.
| Dataset | Exam | Candidates | Site |
|---|---|---|---|
2016 |
2016 | 877,460 | /2016/ |
2017 |
2017 | 861,068 | /2017/ |
Two earlier 2017 publications (2017-old, 2017-old2) were kept for a while
because they disagreed with the current one. They have been removed; they remain
in git history.
Layout
The repository is one directory per pipeline stage, plus the two stores they pass between them.
crawler/ Go — re-fetches the source spreadsheets → data/
parser/ Go — Excel to SQLite data/ → .db
assembler/ Go — verifies, compresses, builds, assembles .db + web/ → _site/
web/ npm — the frontend, one SvelteKit app for every dataset
data/<id>/ raw Excel files, one directory per dataset
datasets.json the registry: which datasets exist, and their expected size
docs/ architecture, data pipeline, deployment
Each stage runs on its own and hands its output to the next through the stores.
web/ is the only npm project; the three stages are independent Go modules.
datasets.json is the contract between them. It is JSON because Go and the web
app both read it and neither needs a dependency to do so; presentation stays in
web/src/lib/datasets.ts, keyed by id, which fails loudly if the two disagree.
The dataset id is one identifier end to end:
data/2017/ → parser/configs/2017.yml → db/2017.db.gz → /thptqg/2017/
Build
(cd web && npm ci)
go -C assembler run ./cmd/assemble # databases, then the site, into _site/
npx serve _site
That one command compiles the parser, builds and verifies each database against
its registry row count, compresses it, builds the web app and assembles _site —
refusing to continue if a database is short, an artifact looks truncated, or one
is missing altogether. Sub-steps when iterating:
go -C assembler run ./cmd/assemble db 2017 # one database
go -C assembler run ./cmd/assemble site # web build and _site only
go -C assembler run ./cmd/assemble verify A B # compare two sets of databases
(cd web && npm run dev) # the app against staged databases
The source spreadsheets are committed, so a crawl is only needed to refresh them:
go -C crawler run ./cmd/crawl 2016
go -C crawler run ./cmd/crawl 2017
Each reads the download links out of the article that published the dataset, so no link list is kept in the repository. Crawling is idempotent — files already present are skipped — and is never part of the build.
Pushing to main runs the same steps in
.github/workflows/deploy-pages.yml and publishes to GitHub Pages.
Adding a dataset
- Put the Excel files in
data/<id>/ - Add
parser/configs/<id>.yml— sheet mode, column indices, validation guards. No SQL; the schema is canonical. - Add an entry to
datasets.jsonwith its expected row count and size - Add the matching presentation to
CONTENTinweb/src/lib/datasets.ts
Everything else follows: the assembler, the router and the hub all read the registry, and the UI adapts to whichever columns the dataset fills. Steps 3 and 4 check each other, so forgetting either one fails rather than half-working.
Docs
See docs/ — overview,
architecture,
data pipeline,
deployment.