mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-05 00:13:44 +00:00
The browser downloaded 45 MB of gzipped SQLite before it could answer anything. Now sql.js-httpvfs asks for the pages a query touches and the databases ship uncompressed as <id>.sqlite3 — a byte range of a gzip stream is not a byte range of a database. That only works if every query the site issues is index-driven, and measured against the real 2016 file, most were not: so_bao_danh = ? SEARCH via PK ~20 KB ho_ten_ascii LIKE '%x%' SCAN 127 MB ho_ten_ascii LIKE 'x%' SCAN 127 MB COUNT(*) covering index scan 20 MB ORDER BY toan DESC LIMIT 10 SCAN + temp b-tree 127 MB Prefix LIKE scans because SQLite's LIKE optimisation needs a NOCASE index; a range comparison does use the index. So the schema changed to suit the access pattern rather than the search changing to suit the schema. name_word holds one row per word of each name, WITHOUT ROWID so the table is the index, carrying ho_ten_ascii so a multi-word query is resolved inside a single b-tree. name_word_freq says which word of a query is rarest — the vocabulary is 4,397 words across 2.87M entries, so "buu loc" seeks on 287 entries rather than walking the 300,000 that "thi" would. Searching by any word of a name survives, at a few hundred KB a query. idx_ho_ten and idx_ho_ten_ascii are gone: no plan could use either. Partial indexes on toan, khtn and khxh cost 12 MB and keep the SQL presets off a full scan. The footer's candidate count now comes from datasets.json instead of COUNT(*). 2016 grows 223.5 MB to 288.6 MB, 2017 162.7 MB to 237.7 MB, and the site is 528 MB against the 1 GB GitHub Pages limit. Row counts are unchanged. The SQL tab is the one place a user can still write a query that reads the whole table, so it asks before it opens, runs under a byte budget that stops a runaway query, and shows what each query actually fetched. Verified: row counts through the assembler guards, every app query index-driven under EXPLAIN QUERY PLAN, and GitHub Pages returning 206 with a correct Content-Range. Not verified in a browser — this machine has none — and the library refuses to open a file the host compresses, so the deployed response headers need a look.
67 lines
2.6 KiB
Markdown
67 lines
2.6 KiB
Markdown
# Project Overview
|
||
|
||
## Goal
|
||
|
||
A public lookup tool for Vietnam's National High School Graduation Exam scores,
|
||
running entirely in the browser and hosted for free on GitHub Pages. Covers the
|
||
2016 and 2017 exams — 1.7 million candidates across two datasets.
|
||
|
||
## Scope
|
||
|
||
- Lookup by exam ID or full name, with Vietnamese diacritics handled
|
||
- Read-only SQL queries against a single `student` table
|
||
- Admission-block (khối thi) totals computed per candidate
|
||
- Static datasets — both exams are long over and the data is frozen
|
||
|
||
## Target users
|
||
|
||
- Former candidates checking their scores
|
||
- Education researchers and data journalists running aggregate statistics
|
||
- Developers exploring SQL against a real-world dataset
|
||
|
||
## Constraints
|
||
|
||
- **Zero backend.** The database (238–289 MB per dataset) stays on the server
|
||
and the browser reads the pages a query touches over HTTP range requests.
|
||
- **Read-only.** `INSERT`/`UPDATE`/`DELETE` are rejected, so nobody is misled
|
||
into thinking edits persist. The file is fetched, never written.
|
||
- **Row caps.** 100 rows for lookups, 1000 for custom SQL, to prevent browser
|
||
hangs.
|
||
- **Vietnamese-first UI.** App labels and data are Vietnamese; documentation is
|
||
English.
|
||
|
||
## Datasets
|
||
|
||
| id | Exam | Candidates | Notes |
|
||
| --- | --- | --- | --- |
|
||
| `2016` | 2016 | 877,460 | 119 files, four column layouts |
|
||
| `2017` | 2017 | 861,068 | current generation of three publications |
|
||
|
||
Two further 2017 datasets (`2017-old`, `2017-old2`) were kept alongside these
|
||
because the three publications disagreed. They have been removed; git history
|
||
still has them.
|
||
|
||
Both datasets now have a crawler source. 2017 comes from the baotintuc.vn CDN,
|
||
which is still live. 2016 comes from an aggregator article on
|
||
`dtnt.bacninh.edu.vn`, also still online. Both datasets have been crawled
|
||
successfully, so either can be rebuilt from source — see
|
||
[data-pipeline](./data-pipeline.md#sources) for the full article URLs.
|
||
|
||
## History
|
||
|
||
Each year began as a standalone repository (`thptqg2016`, `thptqg2017`), merged
|
||
here with full history. They initially kept separate frontends and separate
|
||
copies of the same Rust parser, synchronised by hand. That duplication was
|
||
removed: there is now one frontend, one parser, and one canonical schema, with
|
||
per-dataset differences confined to one small config file and one registry
|
||
entry each.
|
||
|
||
The unification also fixed a latent data-loss bug — neither year's parser
|
||
config listed the complete set of subjects, so 1,691 candidates were missing
|
||
their foreign-language score. See `data-pipeline.md`.
|
||
|
||
## Status
|
||
|
||
Stable, data frozen. Work is limited to UX polish and keeping the pipeline
|
||
maintainable.
|