Files
thptqg/2017/docs/data-pipeline.md
T
tiennm99 8ca6bfa7a8 docs: expand README, add architecture / deployment / data-pipeline docs
- README: live-site URLs, features, scripts, project layout, quickstart
- docs/system-architecture.md: data flow, 3-variant deploy mechanism,
  schema, parse-quirk matrix, score-tier model, admission-block model,
  frontend key behaviors (deep-link, share, keyboard shortcuts)
- docs/data-pipeline.md: source Excel shape, score regex, why three
  parsers, overflow-sheet gotcha, audit tooling, refresh flow
- docs/deployment-guide.md: CI workflow, adding a new variant, rollback,
  GH Pages file-size note
- docs/README.md: index for the docs dir
2026-04-14 23:25:51 +07:00

89 lines
3.2 KiB
Markdown

# Data Pipeline
From raw Excel files to a compressed SQLite file the browser can load.
## Canonical source
63 Excel files from baotintuc.vn CDN. Re-download anytime:
```bash
node scripts/crawl-baotintuc.js
```
Idempotent — skips files already present. Saves to `data/<ascii-kebab-province>.xls`.
Source article URL: `https://baotintuc.vn/tuyen-sinh/tra-cuu-diem-thi-thpt-2017-cua-63-tinh-thanh-pho-tren-baotintucvn-20170706073512672.htm`
## Source Excel shape
Every file is a single sheet (for `data-old/`) or up to 3 sheets (for `data/` .xls) with columns:
| Col | Name | Content |
|---|---|---|
| 0 | HO_TEN | full name in Vietnamese |
| 1 | NGAY_SINH | `dd/mm/yyyy` |
| 2 | SOBAODANH | 8-digit string, first 2 digits = province |
| 3 | DIEM_THI | concatenated Vietnamese score string, e.g. `"Toán: 6.80 Ngữ văn: 5.25 … Tiếng Anh: 5.80"` |
Some files lack a header row. `isHeaderRow()` in `build-lib.js` detects and skips.
## Score text parsing
`SCORE_PATTERNS` in `build-lib.js` defines one regex per subject:
```js
{ toan: /Toán:\s*(\d+(?:\.\d+)?)/, ngu_van: /Ngữ văn:…/, … }
```
Subjects supported: Toán, Ngữ văn, Vật lí, Hóa học, Sinh học, KHTN, Lịch sử, Địa lí, GDCD, KHXH, Tiếng Anh, Tiếng Pháp, Tiếng Nga, Tiếng Trung.
Missing-in-string → `NULL` in DB (student didn't take that subject).
## Why three parsers
One shared lib + three thin parsers, not one parser with flags. Each dataset's quirks stay visible in its own file:
| File | Source dir | Sheet strategy | Extra guard |
|---|---|---|---|
| `build-database.js` | `data/` | iterate ALL sheets (.xls row overflow) | — |
| `build-database-old.js` | `data-old/` | sheet 0 only | reject rows where `so_bao_danh` isn't digits-only (rejects an `HO_TEN` → `SOBAODANH` header leak present in this export) |
| `build-database-old2.js` | `data-old2/` | iterate ALL sheets (HCM overflow) | skip fully blank rows before counting |
## Overflow-sheet gotcha
The old `.xls` binary format caps at 65,536 rows per sheet. Hà Nội and HCM exceed that in the baotintuc export, so Excel auto-splits the overflow into `Sheet2`. **Only reading `Sheet1` silently drops 13,720 students** (Hà Nội +7,275, HCM +6,445).
The audit `node scripts/audit-row-counts.js` catches this by comparing source-row-count (across all sheets) to DB-row-count.
## Audits
Run any time after a rebuild:
```bash
node scripts/audit-row-counts.js # source rows vs DB rows (zero-loss check)
node scripts/check-duplicates.js # md5 + row-content duplicates in data/
node scripts/diff-datasets.js # compare two DBs (needs backup/ dir)
```
Expected baseline:
| DB | Source rows | Skipped | DB rows |
|---|---|---|---|
| main | 861,131 | 63 empty rows | **861,068** |
| old | 847,349 | 1 header-leak | **847,348** |
| old2 | 679,764 | 0 | **679,764** |
## Refreshing the data
```bash
rm data/*.xls # clear current snapshot
node scripts/crawl-baotintuc.js # re-download from CDN
npm run build:db # rebuild
node scripts/audit-row-counts.js # verify
gzip -kf -9 public/thptqg2017.db # compress for shipping
```
## Adding a new source dataset
See `docs/deployment-guide.md` → "Adding a new variant".