mirror of
https://github.com/tiennm99/noitu.git
synced 2026-10-05 14:14:53 +00:00
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game wordlist: 48,216 Vietnamese words of two or more syllables, indexed by first and last syllable with an out-degree table for dead-end detection. Source schema is auto-detected rather than hardcoded, since it is someone else's release artifact; explicit flags override it and are validated against the real tables, because SQLite silently reads an unknown double-quoted column as a string literal. Accept a word only if every syllable fits Vietnamese phonotactics. An alphabet check is not enough: "credit card" and "come out" use only letters Vietnamese has, and the multilingual source tags them as Vietnamese. Onset matching backtracks so the gi digraph does not swallow the nucleus of common words like "gi", "gin" and "gi" (rust). Record accepted spelling variants in an alias table rather than solving tone placement at runtime. Tone shifting applies only to open oa/oe/uy syllables, since "hoan" and "hoai" have a single correct spelling, and "qu" is a consonant onset. The i/y alternation uses an onset allowlist plus explicit pairs, because it is lexical rather than productive. Variants are generated as a cross product over syllables so a word with two variable syllables still offers the fully modern spelling. Build to a temporary file and rename only after commit, so a failed run cannot leave an empty database where a good one was, then re-open the result and verify its invariants on disk. Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact from the Apache-2.0 code, loaded at runtime and never embedded. Ships NOTICE, data/LICENSE and an attribution file recording every change.
38 lines
1.7 KiB
Plaintext
38 lines
1.7 KiB
Plaintext
noitu — Vietnamese "nối từ" word-chain game
|
|
Copyright 2026 the noitu authors
|
|
|
|
This product is distributed under two licenses, applying to different artifacts.
|
|
|
|
------------------------------------------------------------------------------
|
|
1. SOURCE CODE — Apache License 2.0
|
|
------------------------------------------------------------------------------
|
|
|
|
All source code in this repository is licensed under the Apache License,
|
|
Version 2.0. See the LICENSE file at the repository root.
|
|
|
|
This covers everything under server/, tools/, web/, and proto/.
|
|
|
|
------------------------------------------------------------------------------
|
|
2. DICTIONARY DATA — Creative Commons Attribution-ShareAlike 4.0 International
|
|
------------------------------------------------------------------------------
|
|
|
|
The Vietnamese dictionary database is NOT covered by the Apache License. It is
|
|
derived from a third-party dataset and remains under CC BY-SA 4.0:
|
|
|
|
Affected artifacts: data/noitu.db (and data/dictionary.db, its source)
|
|
License text: data/LICENSE
|
|
Attribution and
|
|
list of changes: data/ATTRIBUTION.md
|
|
Upstream source: https://github.com/minhqnd/dictionary (release v2.0.0)
|
|
|
|
CC BY-SA 4.0 is a share-alike license. Any distribution of the derived database
|
|
-- including inside a container image or other packaged build -- must carry the
|
|
same license, this attribution, and the record of modifications.
|
|
|
|
The two regimes are kept on separate artifacts deliberately: the database is
|
|
loaded at runtime from a file and is never embedded, compiled, or linked into
|
|
the Go binary.
|
|
|
|
Neither database file is committed to version control. Both are build artifacts
|
|
produced by `make fetch-dict` and `make dict`.
|