mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-05 08:14:55 +00:00
Both sources carried their download links as a hardcoded array, which is not a crawl: the lists could drift from what the articles actually published, and nothing would say so. A source now names the article and how to name what it finds there, and internal/article reads the links out of that page at run time. 2016 takes its filenames straight from the URL. 2017 cannot — the CDN names are inconsistent (Angiang.xls, 1BaRiaVungTau.xls, 23HaiPhong.xls) — so it derives them from the province in the link text, transliterated to ASCII the same way go-parser builds ho_ten_ascii. Filenames stay load-bearing: go-parser sorts inputs bytewise and inserts last-wins, so they decide which row survives a duplicate exam number. Saved copies of both articles are committed as fixtures, and a test asserts that reading them and applying each naming rule reproduces data/<id> exactly, in both directions. Resolve also rejects a page that yields the wrong number of links or two links that would write the same file, since either silently costs the dataset files that only the row-count guard would notice afterwards. Verified against the live 2017 article: a from-scratch crawl of all 63 files leaves the committed data unchanged.