expvar counters for connections, rooms, games, submissions by rejection
reason, eliminations, chat, joins and bot moves, served on a separate
debug address so they never sit on the public mux, plus one structured
word_rejected log line per refused word carrying the normalized word and
its link. GET /readyz flips to 503 while draining; SIGTERM stops new
rooms, waits up to NOITU_DRAIN_TIMEOUT for live games, then shuts down.
GET /version and the startup log carry the build's git describe.
Fuzz targets for the frame decoder, the text sanitizer and Vietnamese
normalization; the last one found that composing before lowercasing
could leave a non-NFC result, now recomposed after lowering.
CI runs on dev as well as main, gates gofmt and golangci-lint, tracks the
buf major instead of an exact pin, and dependabot watches every
ecosystem. The lint findings that had been hidden by the default
per-issue cap are fixed.
PlayedWord.player_id was declared and read by the client but never set by
the server, so a room of three or four never showed who played each word.
The chain byline now comes from the room, with a producer-side test.
The process also gains the ceilings it was missing: a cap on live rooms and
on open sockets, a per-connection frame-rate limit so a payload-less frame
is no longer free, and an opt-in trusted-proxy list so the join limiter can
tell players apart behind the documented reverse proxy instead of putting
them in one bucket. The typed word is sanitized before the engine stores it,
since every seat is shown it; a room exiting on its idle clock releases the
sessions still bound to it; the dictionary builder escapes its SQLite path
like the store does and renames over the old database instead of deleting
it first.
bot.BoardFor and hub.roomCount were reachable from tests only and now live beside them; session.serve drops an always-true nil check (readLoop never returns nil); invisible characters in test literals become escapes; comments no longer cite plan phases; `.dockerignore` keeps the dump out of the build context. No behaviour change; vet, deadcode, staticcheck, go test -race, svelte-check and vitest green.
reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
Drop what nothing uses any more: the --max-syllables flag with its
sourceSpec field, meta row and reject reason; the store's clearRejection;
the --gap CSS token; three orphaned i18n strings; two exports with no
importer; the preview npm script; the proto.yml "baseline exists" step
that has been unconditionally true since the schema landed on main; and
two root-level .gitignore entries for paths Playwright never writes.
Correct text that outlived its subject: NOTICE named a tools/ directory
that never existed, the README still pointed at the first plan and
carried a migration note for retired sources, and a few comments still
said "release" for an upstream that is a weekly export. The pin test
header now says three copies and also holds README and ATTRIBUTION to
the same URL; the Make-targets table lists clean and help.
Wire fixtures regenerated so client_hello uses the ProtocolVersion
constant and server_game_started carries the 30s turn limit.
builder_version becomes 4 for the dropped meta row.
Replace the pinned 2018 undertheseanlp wordlist with kaikki.org's current
wiktextract export of the Vietnamese Wiktionary, read with --kaikki. The
same authors and website, eight years fresher: 34,813 words instead of
26,845, with every graph metric up and bot game length unchanged.
The export is fetched fresh for every build and is not pinned, by the
owner's decision: kaikki keeps no dated snapshots, so a checksum would
break weekly. The builder therefore hashes the file as it streams it and
records source_sha256, source_rows and source_fetched_at in meta; the
fetch downloads to a .part name and renames on success; a truncated or
non-JSON body fails the build, and the --min-words floor rises to 30,000.
DICT_SHA256 and verify-dict are gone; the Makefile, Dockerfile and the
builder's URL constant are held in agreement by a test.
Current Wiktionary text is CC BY-SA 4.0, so the data licence returns to
4.0: data/LICENSE is restored, and NOTICE, ATTRIBUTION, README, the image
docs and the in-game footer credit Wiktionary tiếng Việt's contributors
and wiktextract/kaikki.org. builder_version becomes 3 for the changed
meta contract.
Replace the 179 MB minhqnd SQLite aggregate with the 4.8 MB
undertheseanlp/dictionary JSONL, pinned by commit and SHA-256, reading
only rows tagged "wiktionary". The two other wordlists in that file are
never read: hongocduc is GPL and would force a relicense, tudientv is an
unlicensed derivative of a commercial dictionary.
build-dictionary gains --merged and --sources (names validated, default
wiktionary), a shared finish() tail, and meta rows for the source commit
and the sources kept and excluded. The SQLite --in path, its schema
auto-detection and their tests are removed. Fixture builds now record
that they carry no upstream data instead of inheriting a licence string.
The --min-words floor moves from 40,000 to 20,000; the corpus is 26,845
words, down from 48,216, all of the loss being words absent from the
2018 Wiktionary scrape. Capitalization is not a filter.
The data licence follows the source text: CC BY-SA 3.0 Unported, which
is what vi.wiktionary.org carried in 2018. LICENSE, NOTICE, ATTRIBUTION,
the README, the image docs, the builder's meta string and the in-game
footer all name Wiktionary tiếng Việt's contributors as the authors and
undertheseanlp as the intermediary. The Makefile/Dockerfile pin test now
also checks the commit the builder stamps into the database.
build-dictionary gains a --words mode that reads a plain list instead of the
upstream database. Everything after that — filtering, alias generation, writing
and verification — is the code the real build uses, so a fixture cannot drift
into being shaped differently from what the server loads.
testdata/fixture-words.txt is hand-written rather than extracted, so no test
artifact carries the upstream release's licence. Its graph is built around one
hub syllable: it is the only one with enough continuations for the server to
open a game on, so every game starts on a word ending in it and a scripted test
always knows the first answer.
One goroutine owns each room and its engine. The room goroutine starts before
anyone is seated and seating is itself a message, so reading run() is a complete
proof of the concurrency contract rather than a convention to uphold. The bot
searches a frozen copy of the board instead of the live engine.
Reads carry no deadline; liveness is ping-based, because a read timeout cannot
distinguish a healthy player idling in the lobby from a dead socket.
The hub no longer binds a joiner to a seat before the room decides whether to
seat them. Anyone holding a room code could previously resign or play on a
seated player's behalf, and the room code is the only credential online 1v1 has.
cmd/noitu-server serves the API and, when NOITU_WEB_DIR is set, the built
frontend, with unknown paths falling back to index.html for client routes. All
configuration is environment-only and every variable has a working default.
Reduce the 179 MB minhqnd/dictionary SQLite release to a ~3 MB game
wordlist: 48,216 Vietnamese words of two or more syllables, indexed by
first and last syllable with an out-degree table for dead-end detection.
Source schema is auto-detected rather than hardcoded, since it is someone
else's release artifact; explicit flags override it and are validated
against the real tables, because SQLite silently reads an unknown
double-quoted column as a string literal.
Accept a word only if every syllable fits Vietnamese phonotactics. An
alphabet check is not enough: "credit card" and "come out" use only
letters Vietnamese has, and the multilingual source tags them as
Vietnamese. Onset matching backtracks so the gi digraph does not swallow
the nucleus of common words like "gi", "gin" and "gi" (rust).
Record accepted spelling variants in an alias table rather than solving
tone placement at runtime. Tone shifting applies only to open oa/oe/uy
syllables, since "hoan" and "hoai" have a single correct spelling, and
"qu" is a consonant onset. The i/y alternation uses an onset allowlist
plus explicit pairs, because it is lexical rather than productive.
Variants are generated as a cross product over syllables so a word with
two variable syllables still offers the fully modern spelling.
Build to a temporary file and rename only after commit, so a failed run
cannot leave an empty database where a good one was, then re-open the
result and verify its invariants on disk.
Licensing: the derived data is CC BY-SA 4.0 and stays a separate artifact
from the Apache-2.0 code, loaded at runtime and never embedded. Ships
NOTICE, data/LICENSE and an attribution file recording every change.