Files
thptqg/README.md
T
tiennm99 dbc23c25c5 feat: read the databases over HTTP range requests
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.

That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:

  so_bao_danh = ?              SEARCH via PK          ~20 KB
  ho_ten_ascii LIKE '%x%'      SCAN                   127 MB
  ho_ten_ascii LIKE 'x%'       SCAN                   127 MB
  COUNT(*)                     covering index scan     20 MB
  ORDER BY toan DESC LIMIT 10  SCAN + temp b-tree     127 MB

Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.

name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.

idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).

2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.

The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.

Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
2026-08-14 12:42:48 +07:00

3.9 KiB

thptqg

Tra cứu điểm thi THPT Quốc gia — exam-score lookup for Vietnam's national high school graduation exam. Client-side SQL over a SQLite database read in place by HTTP range request, built from the published .xls/.xlsx score files by the Go parser module. Where those files come from: data pipeline.

Live at tiennm99.github.io/thptqg.

Dataset Exam Candidates Site
2016 2016 877,460 /2016/
2017 2017 861,068 /2017/

Two earlier 2017 publications (2017-old, 2017-old2) were kept for a while because they disagreed with the current one. They have been removed; they remain in git history.

Layout

The repository is one directory per pipeline stage, plus the two stores they pass between them.

crawler/      Go   — re-fetches the source spreadsheets      → data/
parser/       Go   — Excel to SQLite                          data/ → .db
assembler/    Go   — verifies, compresses, builds, assembles  .db + web/ → _site/
web/          npm  — the frontend, one SvelteKit app for every dataset
data/<id>/         raw Excel files, one directory per dataset
datasets.json      the registry: which datasets exist, and their expected size
docs/              architecture, data pipeline, deployment

Each stage runs on its own and hands its output to the next through the stores. web/ is the only npm project; the three stages are independent Go modules.

datasets.json is the contract between them. It is JSON because Go and the web app both read it and neither needs a dependency to do so; presentation stays in web/src/lib/datasets.ts, keyed by id, which fails loudly if the two disagree.

The dataset id is one identifier end to end:

data/2017/ → parser/configs/2017.yml → db/2017.sqlite3 → /thptqg/2017/

Build

(cd web && npm ci)
go -C assembler run ./cmd/assemble        # databases, then the site, into _site/
npx serve _site

That one command compiles the parser, builds and verifies each database against its registry row count, compresses it, builds the web app and assembles _site — refusing to continue if a database is short, an artifact looks truncated, or one is missing altogether. Sub-steps when iterating:

go -C assembler run ./cmd/assemble db 2017   # one database
go -C assembler run ./cmd/assemble site      # web build and _site only
go -C assembler run ./cmd/assemble verify A B  # compare two sets of databases
(cd web && npm run dev)                      # the app against staged databases

The source spreadsheets are committed, so a crawl is only needed to refresh them:

go -C crawler run ./cmd/crawl 2016
go -C crawler run ./cmd/crawl 2017

Each reads the download links out of the article that published the dataset, so no link list is kept in the repository. Crawling is idempotent — files already present are skipped — and is never part of the build.

Pushing to main runs the same steps in .github/workflows/deploy-pages.yml and publishes to GitHub Pages.

Adding a dataset

  1. Put the Excel files in data/<id>/
  2. Add parser/configs/<id>.yml — sheet mode, column indices, validation guards. No SQL; the schema is canonical.
  3. Add an entry to datasets.json with its expected row count and size
  4. Add the matching presentation to CONTENT in web/src/lib/datasets.ts

Everything else follows: the assembler, the router and the hub all read the registry, and the UI adapts to whichever columns the dataset fills. Steps 3 and 4 check each other, so forgetting either one fails rather than half-working.

Docs

See docs/ — overview, architecture, data pipeline, deployment.