reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
Replace the pinned 2018 undertheseanlp wordlist with kaikki.org's current
wiktextract export of the Vietnamese Wiktionary, read with --kaikki. The
same authors and website, eight years fresher: 34,813 words instead of
26,845, with every graph metric up and bot game length unchanged.
The export is fetched fresh for every build and is not pinned, by the
owner's decision: kaikki keeps no dated snapshots, so a checksum would
break weekly. The builder therefore hashes the file as it streams it and
records source_sha256, source_rows and source_fetched_at in meta; the
fetch downloads to a .part name and renames on success; a truncated or
non-JSON body fails the build, and the --min-words floor rises to 30,000.
DICT_SHA256 and verify-dict are gone; the Makefile, Dockerfile and the
builder's URL constant are held in agreement by a test.
Current Wiktionary text is CC BY-SA 4.0, so the data licence returns to
4.0: data/LICENSE is restored, and NOTICE, ATTRIBUTION, README, the image
docs and the in-game footer credit Wiktionary tiếng Việt's contributors
and wiktextract/kaikki.org. builder_version becomes 3 for the changed
meta contract.
Replace the 179 MB minhqnd SQLite aggregate with the 4.8 MB
undertheseanlp/dictionary JSONL, pinned by commit and SHA-256, reading
only rows tagged "wiktionary". The two other wordlists in that file are
never read: hongocduc is GPL and would force a relicense, tudientv is an
unlicensed derivative of a commercial dictionary.
build-dictionary gains --merged and --sources (names validated, default
wiktionary), a shared finish() tail, and meta rows for the source commit
and the sources kept and excluded. The SQLite --in path, its schema
auto-detection and their tests are removed. Fixture builds now record
that they carry no upstream data instead of inheriting a licence string.
The --min-words floor moves from 40,000 to 20,000; the corpus is 26,845
words, down from 48,216, all of the loss being words absent from the
2018 Wiktionary scrape. Capitalization is not a filter.
The data licence follows the source text: CC BY-SA 3.0 Unported, which
is what vi.wiktionary.org carried in 2018. LICENSE, NOTICE, ATTRIBUTION,
the README, the image docs, the builder's meta string and the in-game
footer all name Wiktionary tiếng Việt's contributors as the authors and
undertheseanlp as the intermediary. The Makefile/Dockerfile pin test now
also checks the commit the builder stamps into the database.
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.
Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.
Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).
Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.
Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.
Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.