Files
thptqg/docs/project-overview.md
T
tiennm99 dbc23c25c5 feat: read the databases over HTTP range requests
The browser downloaded 45 MB of gzipped SQLite before it could answer
anything. Now sql.js-httpvfs asks for the pages a query touches and the
databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip
stream is not a byte range of a database.

That only works if every query the site issues is index-driven, and
measured against the real 2016 file, most were not:

  so_bao_danh = ?              SEARCH via PK          ~20 KB
  ho_ten_ascii LIKE '%x%'      SCAN                   127 MB
  ho_ten_ascii LIKE 'x%'       SCAN                   127 MB
  COUNT(*)                     covering index scan     20 MB
  ORDER BY toan DESC LIMIT 10  SCAN + temp b-tree     127 MB

Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE
index; a range comparison does use the index. So the schema changed to
suit the access pattern rather than the search changing to suit the
schema.

name_word holds one row per word of each name, WITHOUT ROWID so the
table is the index, carrying ho_ten_ascii so a multi-word query is
resolved inside a single b-tree. name_word_freq says which word of a
query is rarest — the vocabulary is 4,397 words across 2.87M entries, so
"buu loc" seeks on 287 entries rather than walking the 300,000 that
"thi" would. Searching by any word of a name survives, at a few hundred
KB a query.

idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either.
Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL
presets off a full scan. The footer's candidate count now comes from
datasets.json instead of COUNT(*).

2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site
is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged.

The SQL tab is the one place a user can still write a query that reads
the whole table, so it asks before it opens, runs under a byte budget
that stops a runaway query, and shows what each query actually fetched.

Verified: row counts through the assembler guards, every app query
index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206
with a correct Content-Range. Not verified in a browser — this machine
has none — and the library refuses to open a file the host compresses,
so the deployed response headers need a look.
2026-08-14 12:42:48 +07:00

67 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Project Overview
## Goal
A public lookup tool for Vietnam's National High School Graduation Exam scores,
running entirely in the browser and hosted for free on GitHub Pages. Covers the
2016 and 2017 exams — 1.7 million candidates across two datasets.
## Scope
- Lookup by exam ID or full name, with Vietnamese diacritics handled
- Read-only SQL queries against a single `student` table
- Admission-block (khối thi) totals computed per candidate
- Static datasets — both exams are long over and the data is frozen
## Target users
- Former candidates checking their scores
- Education researchers and data journalists running aggregate statistics
- Developers exploring SQL against a real-world dataset
## Constraints
- **Zero backend.** The database (238–289 MB per dataset) stays on the server
and the browser reads the pages a query touches over HTTP range requests.
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
into thinking edits persist. The file is fetched, never written.
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
hangs.
- **Vietnamese-first UI.** App labels and data are Vietnamese; documentation is
English.
## Datasets
| id | Exam | Candidates | Notes |
| --- | --- | --- | --- |
| `2016` | 2016 | 877,460 | 119 files, four column layouts |
| `2017` | 2017 | 861,068 | current generation of three publications |
Two further 2017 datasets (`2017-old`, `2017-old2`) were kept alongside these
because the three publications disagreed. They have been removed; git history
still has them.
Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN,
which is still live. 2016 comes from an aggregator article on
`dtnt.bacninh.edu.vn`, also still online. Both datasets have been crawled
successfully, so either can be rebuilt from source — see
[data-pipeline](./data-pipeline.md#sources) for the full article URLs.
## History
Each year began as a standalone repository (`thptqg2016`, `thptqg2017`), merged
here with full history. They initially kept separate frontends and separate
copies of the same Rust parser, synchronised by hand. That duplication was
removed: there is now one frontend, one parser, and one canonical schema, with
per-dataset differences confined to one small config file and one registry
entry each.
The unification also fixed a latent data-loss bug — neither year's parser
config listed the complete set of subjects, so 1,691 candidates were missing
their foreign-language score. See `data-pipeline.md`.
## Status
Stable, data frozen. Work is limited to UX polish and keeping the pipeline
maintainable.