mirror of
https://github.com/tiennm99/thptqg.git
synced 2026-10-06 04:15:00 +00:00
Both sources carried their download links as a hardcoded array, which is not a crawl: the lists could drift from what the articles actually published, and nothing would say so. A source now names the article and how to name what it finds there, and internal/article reads the links out of that page at run time. 2016 takes its filenames straight from the URL. 2017 cannot — the CDN names are inconsistent (Angiang.xls, 1BaRiaVungTau.xls, 23HaiPhong.xls) — so it derives them from the province in the link text, transliterated to ASCII the same way go-parser builds ho_ten_ascii. Filenames stay load-bearing: go-parser sorts inputs bytewise and inserts last-wins, so they decide which row survives a duplicate exam number. Saved copies of both articles are committed as fixtures, and a test asserts that reading them and applying each naming rule reproduces data/<id> exactly, in both directions. Resolve also rejects a page that yields the wrong number of links or two links that would write the same file, since either silently costs the dataset files that only the row-count guard would notice afterwards. Verified against the live 2017 article: a from-scratch crawl of all 63 files leaves the committed data unchanged.
5 lines
308 B
Plaintext
5 lines
308 B
Plaintext
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
|
|
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
|
|
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
|
|
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
|