Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game wordlist: 48,216 Vietnamese words of two or more syllables, indexed by first and last syllable with an out-degree table for dead-end detection. Source schema is auto-detected rather than hardcoded, since it is someone else's release artifact; explicit flags override it and are validated against the real tables, because SQLite silently reads an unknown double-quoted column as a string literal. Accept a word only if every syllable fits Vietnamese phonotactics. An alphabet check is not enough: "credit card" and "come out" use only letters Vietnamese has, and the multilingual source tags them as Vietnamese. Onset matching backtracks so the gi digraph does not swallow the nucleus of common words like "gi", "gin" and "gi" (rust). Record accepted spelling variants in an alias table rather than solving tone placement at runtime. Tone shifting applies only to open oa/oe/uy syllables, since "hoan" and "hoai" have a single correct spelling, and "qu" is a consonant onset. The i/y alternation uses an onset allowlist plus explicit pairs, because it is lexical rather than productive. Variants are generated as a cross product over syllables so a word with two variable syllables still offers the fully modern spelling. Build to a temporary file and rename only after commit, so a failed run cannot leave an empty database where a good one was, then re-open the result and verify its invariants on disk. Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact from the Apache-2.0 code, loaded at runtime and never embedded. Ships NOTICE, data/LICENSE and an attribution file recording every change.
4.0 KiB
Dictionary Data Attribution
The Vietnamese dictionary data used by this game is not original work of this project. It is derived from a third-party dataset licensed under CC BY-SA 4.0, and this file records the attribution and the modifications required by that license.
Source
| Field | Value |
|---|---|
| Project | minhqnd/dictionary |
| Asset | dictionary.db, release v2.0.0 (~179 MB) |
| Author | minhqnd |
| Data license | CC BY-SA 4.0 — full text in LICENSE |
| Code license (not used here) | MIT |
That project is itself an aggregation. Its own upstream sources include Wiktionary (CC BY-SA) and vntk/dictionary, among other Vietnamese dictionary projects. Those upstream attributions carry through this file.
Modifications made by this project
server/cmd/build-dictionary transforms the upstream dictionary.db into data/noitu.db.
The derived database is a modified version of the source data. Changes:
- Language filter — kept only entries with
lang_code = 'vi'; all other languages (of 1,500+ language pairs in the source) were dropped. - Length filter — kept only words of 2 or more space-separated syllables, as required by the nối từ game rules. Single-syllable entries were dropped.
- Content filter — dropped entries containing digits or punctuation, entries using letters Vietnamese does not have (f, j, w, z), and entries whose syllables do not fit Vietnamese phonotactics (a closed inventory of onsets, nuclei and codas). This removes loanwords and foreign phrases such as "credit card" and "world cup" that the multilingual source tags as Vietnamese. Diacritic-free Vietnamese words ("con cua") are kept.
- Normalization — all words Unicode NFC-normalized, lowercased, and whitespace-collapsed.
- Spelling aliases — added an
aliasestable mapping alternative Vietnamese spellings to canonical entries. Two kinds: competing tone placement in open oa/oe/uy syllables (hoà→hòa,thuý→thúy), and i/y alternation in Sino-Vietnamese syllables (quí→quý,lí→lý). The majority are the i/y kind. These aliases are generated by this project and are not present upstream. - Added columns and tables —
firstandlastsyllable columns, asyllablescount, an index onfirst, asyllablesout-degree table, and ametatable recording provenance. All added for game lookups. - Deduplication — the source lists a word once per sense; entries were deduplicated by normalized form. A generated spelling variant that is itself a real word, or that more than one word would claim, is discarded rather than recorded as an alias.
- Dropped fields — all definitions, translations, pronunciations, examples, part-of-speech tags, and relations from the source were discarded. The derived database contains only word forms, not meanings.
Share-alike obligation
CC BY-SA 4.0 is a share-alike license. The derived database data/noitu.db, and any
distribution of it, remains licensed under CC BY-SA 4.0 — including when it is shipped
inside a container image or any other packaged build of this project.
This obligation applies to the data only. The source code of this project is licensed
separately under Apache-2.0 (see the repository root LICENSE and NOTICE). The derived
database is loaded at runtime from a file and is never compiled or linked into the binary,
keeping the two licensing regimes on separate artifacts.
How to reproduce the derived data
make fetch-dict # downloads the ~179 MB upstream dictionary.db into data/
make dict # derives data/noitu.db from it
Neither file is committed to version control; both are build artifacts.