Files
noitu/data/ATTRIBUTION.md
T
tiennm99 4b4afea679 feat(dictionary): build Vietnamese wordlist from upstream dictionary
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.

Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.

Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).

Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.

Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.

Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.
2026-09-04 16:25:43 +07:00

4.0 KiB

Dictionary Data Attribution

The Vietnamese dictionary data used by this game is not original work of this project. It is derived from a third-party dataset licensed under CC BY-SA 4.0, and this file records the attribution and the modifications required by that license.

Source

Field Value
Project minhqnd/dictionary
Asset dictionary.db, release v2.0.0 (~179 MB)
Author minhqnd
Data license CC BY-SA 4.0 — full text in LICENSE
Code license (not used here) MIT

That project is itself an aggregation. Its own upstream sources include Wiktionary (CC BY-SA) and vntk/dictionary, among other Vietnamese dictionary projects. Those upstream attributions carry through this file.

Modifications made by this project

server/cmd/build-dictionary transforms the upstream dictionary.db into data/noitu.db. The derived database is a modified version of the source data. Changes:

  1. Language filter — kept only entries with lang_code = 'vi'; all other languages (of 1,500+ language pairs in the source) were dropped.
  2. Length filter — kept only words of 2 or more space-separated syllables, as required by the nối từ game rules. Single-syllable entries were dropped.
  3. Content filter — dropped entries containing digits or punctuation, entries using letters Vietnamese does not have (f, j, w, z), and entries whose syllables do not fit Vietnamese phonotactics (a closed inventory of onsets, nuclei and codas). This removes loanwords and foreign phrases such as "credit card" and "world cup" that the multilingual source tags as Vietnamese. Diacritic-free Vietnamese words ("con cua") are kept.
  4. Normalization — all words Unicode NFC-normalized, lowercased, and whitespace-collapsed.
  5. Spelling aliases — added an aliases table mapping alternative Vietnamese spellings to canonical entries. Two kinds: competing tone placement in open oa/oe/uy syllables (hoà → hòa, thuý → thúy), and i/y alternation in Sino-Vietnamese syllables (quí → quý, lí → lý). The majority are the i/y kind. These aliases are generated by this project and are not present upstream.
  6. Added columns and tables — first and last syllable columns, a syllables count, an index on first, a syllables out-degree table, and a meta table recording provenance. All added for game lookups.
  7. Deduplication — the source lists a word once per sense; entries were deduplicated by normalized form. A generated spelling variant that is itself a real word, or that more than one word would claim, is discarded rather than recorded as an alias.
  8. Dropped fields — all definitions, translations, pronunciations, examples, part-of-speech tags, and relations from the source were discarded. The derived database contains only word forms, not meanings.

Share-alike obligation

CC BY-SA 4.0 is a share-alike license. The derived database data/noitu.db, and any distribution of it, remains licensed under CC BY-SA 4.0 — including when it is shipped inside a container image or any other packaged build of this project.

This obligation applies to the data only. The source code of this project is licensed separately under Apache-2.0 (see the repository root LICENSE and NOTICE). The derived database is loaded at runtime from a file and is never compiled or linked into the binary, keeping the two licensing regimes on separate artifacts.

How to reproduce the derived data

make fetch-dict   # downloads the ~179 MB upstream dictionary.db into data/
make dict         # derives data/noitu.db from it

Neither file is committed to version control; both are build artifacts.