Files
noitu/NOTICE
T
tiennm99 4b4afea679 feat(dictionary): build Vietnamese wordlist from upstream dictionary
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.

Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.

Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).

Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.

Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.

Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.
2026-09-04 16:25:43 +07:00

38 lines
1.7 KiB
Plaintext

noitu — Vietnamese "nối từ" word-chain game
Copyright 2026 the noitu authors
This product is distributed under two licenses, applying to different artifacts.
------------------------------------------------------------------------------
1. SOURCE CODE — Apache License 2.0
------------------------------------------------------------------------------
All source code in this repository is licensed under the Apache License,
Version 2.0. See the LICENSE file at the repository root.
This covers everything under server/, tools/, web/, and proto/.
------------------------------------------------------------------------------
2. DICTIONARY DATA — Creative Commons Attribution-ShareAlike 4.0 International
------------------------------------------------------------------------------
The Vietnamese dictionary database is NOT covered by the Apache License. It is
derived from a third-party dataset and remains under CC BY-SA 4.0:
Affected artifacts: data/noitu.db (and data/dictionary.db, its source)
License text: data/LICENSE
Attribution and
list of changes: data/ATTRIBUTION.md
Upstream source: https://github.com/minhqnd/dictionary (release v2.0.0)
CC BY-SA 4.0 is a share-alike license. Any distribution of the derived database
-- including inside a container image or other packaged build -- must carry the
same license, this attribution, and the record of modifications.
The two regimes are kept on separate artifacts deliberately: the database is
loaded at runtime from a file and is never embedded, compiled, or linked into
the Go binary.
Neither database file is committed to version control. Both are build artifacts
produced by `make fetch-dict` and `make dict`.