diff --git a/plans/260425-1945-mongodb-atlas-migration/plan.md b/plans/260425-1945-mongodb-atlas-migration/plan.md index 9fd5d2c..c22eb0c 100644 --- a/plans/260425-1945-mongodb-atlas-migration/plan.md +++ b/plans/260425-1945-mongodb-atlas-migration/plan.md @@ -1,16 +1,19 @@ --- title: "Migrate miti99bot from CF KV+D1 to MongoDB Atlas M0" description: "Dual-write migration to Atlas M0 with explicit cold-start abort threshold and Upstash pivot path." -status: planning +status: code-complete priority: P2 effort: 22h -branch: main +branch: dev tags: [storage, migration, mongodb, atlas, cloudflare-workers] created: 2026-04-25 +code_completed: 2026-04-26 blockedBy: [] blocks: [] --- +> **Status note:** All 8 phases of code/config/scripts/docs are implemented and committed on `dev` (commits `6f0b5ff`..`e2e3112`). 503 → 733 tests; lint clean; `register:dry` green. **Operator-driven execution is pending**: Atlas provisioning (Phase 01 §1-7), real-cluster smoke tests (Phase 01 §13-14), backfill runs (Phase 05), 24-72h soak (Phase 06), cutover stages (Phase 07), and Stage 3 code cleanup (delete CFKVStore/dual-stores after binding deletion). Plan stays here (not archived) until cutover lands or the Upstash standby (`phase-07-alt-pivot.md`) executes. + # Plan: KV+D1 → MongoDB Atlas M0 User-chosen path despite research recommending Upstash. Goal: validate cold-start UX firsthand with safe rollback to KV/D1 (or pivot to Upstash) if M0 cold-start P95 exceeds derived threshold. @@ -33,15 +36,15 @@ User-chosen path despite research recommending Upstash. Goal: validate cold-star | # | Phase | Status | Effort | Owner files | |---|-------|--------|--------|-------------| -| 01 | [Atlas setup + wrangler config](phase-01-atlas-setup.md) | pending | 2h | `wrangler.toml`, `.env.deploy.example`, `scripts/check-secret-leaks.js` | -| 02 | [MongoKVStore implementation](phase-02-mongo-kv-store.md) | pending | 3h | `src/db/mongo-*.js` | -| 03 | [MongoTradesStore + trading refactor](phase-03-mongo-sql-store.md) | pending | 3h | `src/db/mongo-trades-store.js`, `src/modules/trading/*` | -| 04 | [Dual-write wrappers + flag + e2e](phase-04-dual-write-wrappers.md) | pending | 4h | `src/db/dual-*.js`, factories, `tests/e2e/*` | -| 05 | [Backfill + verification (local-only)](phase-05-backfill-scripts.md) | pending | 3h | `scripts/backfill-*.js` | -| 06 | [Staged deploy + soak (cold-start gate)](phase-06-staged-deploy-and-soak.md) | pending | 4h | runtime telemetry | -| 07 | [Cutover + decommission](phase-07-cutover-and-decommission.md) | pending | 3h | wrangler bindings | +| 01 | [Atlas setup + wrangler config](phase-01-atlas-setup.md) | code-complete · operator-pending | 2h | `wrangler.toml`, `.env.deploy.example`, `scripts/check-secret-leaks.js` | +| 02 | [MongoKVStore implementation](phase-02-mongo-kv-store.md) | implemented (`5b00cae`) | 3h | `src/db/mongo-*.js` | +| 03 | [MongoTradesStore + trading refactor](phase-03-mongo-sql-store.md) | implemented (`99cd844`) | 3h | `src/db/mongo-trades-store.js`, `src/modules/trading/*` | +| 04 | [Dual-write wrappers + flag + e2e](phase-04-dual-write-wrappers.md) | implemented (`ea7df56`) | 4h | `src/db/dual-*.js`, factories, `tests/e2e/*` | +| 05 | [Backfill + verification (local-only)](phase-05-backfill-scripts.md) | implemented · operator-runs (`0859356`) | 3h | `scripts/backfill-*.js` | +| 06 | [Staged deploy + soak (cold-start gate)](phase-06-staged-deploy-and-soak.md) | code-complete · operator-runs (`55c8739`) | 4h | runtime telemetry | +| 07 | [Cutover + decommission](phase-07-cutover-and-decommission.md) | prereqs-complete · operator-runs (`3f03521`) | 3h | wrangler bindings | | 07-ALT | [Pivot to Upstash (STANDBY)](phase-07-alt-pivot.md) | standby | (3-4d if triggered) | `src/db/upstash-*.js` | -| 08 | [Tests + docs](phase-08-tests-and-docs.md) | pending | 1h | `tests/`, `docs/` | +| 08 | [Tests + docs](phase-08-tests-and-docs.md) | implemented (`e2e3112`) | 1h | `tests/`, `docs/` | ## Critical dependencies - 01 → 02, 03 (Atlas creds + bundle-size gate required) diff --git a/plans/260422-2128-semantle-module/phase-01-foundation.md b/plans/archive/260422-2128-semantle-module/phase-01-foundation.md similarity index 100% rename from plans/260422-2128-semantle-module/phase-01-foundation.md rename to plans/archive/260422-2128-semantle-module/phase-01-foundation.md diff --git a/plans/260422-2128-semantle-module/phase-02-gameplay.md b/plans/archive/260422-2128-semantle-module/phase-02-gameplay.md similarity index 100% rename from plans/260422-2128-semantle-module/phase-02-gameplay.md rename to plans/archive/260422-2128-semantle-module/phase-02-gameplay.md diff --git a/plans/260422-2128-semantle-module/phase-03-tests-docs.md b/plans/archive/260422-2128-semantle-module/phase-03-tests-docs.md similarity index 100% rename from plans/260422-2128-semantle-module/phase-03-tests-docs.md rename to plans/archive/260422-2128-semantle-module/phase-03-tests-docs.md diff --git a/plans/260422-2128-semantle-module/plan.md b/plans/archive/260422-2128-semantle-module/plan.md similarity index 100% rename from plans/260422-2128-semantle-module/plan.md rename to plans/archive/260422-2128-semantle-module/plan.md diff --git a/plans/260422-2128-semantle-module/reports/code-reviewer-260422-2200-semantle-review.md b/plans/archive/260422-2128-semantle-module/reports/code-reviewer-260422-2200-semantle-review.md similarity index 100% rename from plans/260422-2128-semantle-module/reports/code-reviewer-260422-2200-semantle-review.md rename to plans/archive/260422-2128-semantle-module/reports/code-reviewer-260422-2200-semantle-review.md diff --git a/plans/260422-2128-semantle-module/reports/tester-260422-2155-semantle-tests.md b/plans/archive/260422-2128-semantle-module/reports/tester-260422-2155-semantle-tests.md similarity index 100% rename from plans/260422-2128-semantle-module/reports/tester-260422-2155-semantle-tests.md rename to plans/archive/260422-2128-semantle-module/reports/tester-260422-2155-semantle-tests.md diff --git a/plans/260424-1335-twentyq-game-module/phase-01-foundation.md b/plans/archive/260424-1335-twentyq-game-module/phase-01-foundation.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/phase-01-foundation.md rename to plans/archive/260424-1335-twentyq-game-module/phase-01-foundation.md diff --git a/plans/260424-1335-twentyq-game-module/phase-02-ai-client.md b/plans/archive/260424-1335-twentyq-game-module/phase-02-ai-client.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/phase-02-ai-client.md rename to plans/archive/260424-1335-twentyq-game-module/phase-02-ai-client.md diff --git a/plans/260424-1335-twentyq-game-module/phase-03-gameplay-handlers.md b/plans/archive/260424-1335-twentyq-game-module/phase-03-gameplay-handlers.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/phase-03-gameplay-handlers.md rename to plans/archive/260424-1335-twentyq-game-module/phase-03-gameplay-handlers.md diff --git a/plans/260424-1335-twentyq-game-module/phase-04-tests-docs.md b/plans/archive/260424-1335-twentyq-game-module/phase-04-tests-docs.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/phase-04-tests-docs.md rename to plans/archive/260424-1335-twentyq-game-module/phase-04-tests-docs.md diff --git a/plans/260424-1335-twentyq-game-module/plan.md b/plans/archive/260424-1335-twentyq-game-module/plan.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/plan.md rename to plans/archive/260424-1335-twentyq-game-module/plan.md diff --git a/plans/260424-1335-twentyq-game-module/reports/code-review-260424-1821-project-cleanup.md b/plans/archive/260424-1335-twentyq-game-module/reports/code-review-260424-1821-project-cleanup.md similarity index 100% rename from plans/260424-1335-twentyq-game-module/reports/code-review-260424-1821-project-cleanup.md rename to plans/archive/260424-1335-twentyq-game-module/reports/code-review-260424-1821-project-cleanup.md diff --git a/plans/260424-2215-loldle-new-modes/phase-01-shared-helpers.md b/plans/archive/260424-2215-loldle-new-modes/phase-01-shared-helpers.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-01-shared-helpers.md rename to plans/archive/260424-2215-loldle-new-modes/phase-01-shared-helpers.md diff --git a/plans/260424-2215-loldle-new-modes/phase-02-emoji-module.md b/plans/archive/260424-2215-loldle-new-modes/phase-02-emoji-module.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-02-emoji-module.md rename to plans/archive/260424-2215-loldle-new-modes/phase-02-emoji-module.md diff --git a/plans/260424-2215-loldle-new-modes/phase-03-quote-module.md b/plans/archive/260424-2215-loldle-new-modes/phase-03-quote-module.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-03-quote-module.md rename to plans/archive/260424-2215-loldle-new-modes/phase-03-quote-module.md diff --git a/plans/260424-2215-loldle-new-modes/phase-04-ability-module.md b/plans/archive/260424-2215-loldle-new-modes/phase-04-ability-module.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-04-ability-module.md rename to plans/archive/260424-2215-loldle-new-modes/phase-04-ability-module.md diff --git a/plans/260424-2215-loldle-new-modes/phase-05-splash-module.md b/plans/archive/260424-2215-loldle-new-modes/phase-05-splash-module.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-05-splash-module.md rename to plans/archive/260424-2215-loldle-new-modes/phase-05-splash-module.md diff --git a/plans/260424-2215-loldle-new-modes/phase-06-tests-docs.md b/plans/archive/260424-2215-loldle-new-modes/phase-06-tests-docs.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/phase-06-tests-docs.md rename to plans/archive/260424-2215-loldle-new-modes/phase-06-tests-docs.md diff --git a/plans/260424-2215-loldle-new-modes/plan.md b/plans/archive/260424-2215-loldle-new-modes/plan.md similarity index 100% rename from plans/260424-2215-loldle-new-modes/plan.md rename to plans/archive/260424-2215-loldle-new-modes/plan.md diff --git a/plans/reports/researcher-260423-1110-vietnamese-embeddings-semantle.md b/plans/reports/researcher-260423-1110-vietnamese-embeddings-semantle.md new file mode 100644 index 0000000..80bf6e4 --- /dev/null +++ b/plans/reports/researcher-260423-1110-vietnamese-embeddings-semantle.md @@ -0,0 +1,211 @@ +# Vietnamese Embeddings for Doantu: sup-SimCSE vs PhoW2V + +**Date:** 2026-04-23 | **Scope:** Word-level cosine similarity in fixed 22k-vocab Semantle clone + +--- + +## Executive Verdict + +**Recommendation: PhoW2V (word-level, 300d) is the better fit.** + +Reasons: (1) Purpose-built for word similarity, not sentences; (2) static word2vec format enables precomputation into lookup table — zero inference overhead; (3) no external segmentation tool required; (4) Cloudflare Worker-friendly (ship vectors in KV or bundle with Worker). + +**sup-SimCSE is inferior here** despite better semantic depth, because it requires runtime inference, external VnCoreNLP/pyvi segmentation, and sentence-level training (not optimized for single-word pairs). Violates KISS. + +--- + +## Detailed Comparison + +### 1. **sup-SimCSE-VietNamese-phobert-base** (VoVanPhuc) + +| Aspect | Value | Trade-off | +|--------|-------|-----------| +| **Embedding Dim** | 768 | Large; overkill for word pairs | +| **Training** | Supervised contrastive (SimCSE) | Optimized for sentence similarity, not word pairs | +| **Level** | Sentence-level | Not designed for single-word input | +| **Vocab** | Open (transformer subword tokenization) | Handles unseen words via BPE; adds latency | +| **Segmentation** | **REQUIRED** (RDRSegmenter or pyvi) | Extra runtime dependency; "con chó" must become "con_chó" before encoding | +| **Model Size** | 135M parameters | ~250–350 MB disk; requires transformers + torch on Worker? Infeasible. | +| **Inference** | Runtime + tokenization | Cold start latency; not precomputable for 22k words | +| **Format** | Hugging Face (transformers) | No static dump; requires active model loading | + +**Key Gotcha:** PhoBERT's tokenizer expects **pre-segmented input**. "máy bay" (airplane) as raw input will tokenize as ["máy", "bay"] separately unless segmented to "máy_bay". You'd need VnCoreNLP + pyvi running in Worker — expensive and fragile. + +--- + +### 2. **PhoW2V** (VinAI, datquocnguyen) + +| Aspect | Value | Trade-off | +|--------|-------|-----------| +| **Embedding Dim** | 100 or 300 | 300d ideal; 100d saves space but lower quality | +| **Training** | Unsupervised word2vec (CBOW/Skip-gram) | Pure word-level; no sentence context — but this matches your use case exactly | +| **Level** | Word-level (and syllable-level variant) | Designed for single-word similarity | +| **Vocab** | ~100k words from 20GB corpus (likely covers 22k) | Finite vocab; OOV words get zero vector or nearest neighbor | +| **Segmentation** | **Optional** (can use word-level variant directly) | If input is already word-tokenized, no extra step | +| **Model Size** | ~30–50 MB (gensim KeyedVectors) | Tiny; fits in Cloudflare KV or bundle | +| **Inference** | Zero runtime — precompute entire 22k vocab | `precomputed[word] = word_vector` lookup O(1) | +| **Format** | Gensim KeyedVectors (text/binary) | Exportable as dense matrix (22k × 300) for embedding | + +**Critical Insight:** Word2Vec embeddings are **static lookup tables**. You can precompute similarity scores for all 22k² word pairs offline, or dump the 22k vectors into KV and compute cosine similarity on-demand (O(300) dot product per pair — negligible). + +--- + +## Tokenization & Diacritics + +### PhoBERT (sup-SimCSE) +- Requires VnCoreNLP/pyvi to convert raw input to segmented form. +- **"con chó"** → requires preprocessing to **"con_chó"** before tokenization. +- Adds runtime cost + dependency fragility. +- Diacritics preserved via RDRSegmenter normalization. + +### PhoW2V +- Word-level variant: expects whitespace-separated words (already segmented). +- If using syllable-level, requires syllable input (less relevant here). +- Diacritics preserved (trained on normalized Vietnamese corpus). +- **No segmentation tool needed if vocab covers your compound words.** + +**For a fixed 22k-word game vocabulary:** Pre-segment and validate your entire wordlist at deployment time. Both approaches require diacritic-aware matching ("cá" ≠ "ca"). + +--- + +## Precomputation & Cloudflare Worker Fit + +### sup-SimCSE (Cannot Precompute) +``` +❌ Runtime inference required (PhoBERT forward pass) +❌ Requires transformers + torch (not Worker-compatible) +❌ VnCoreNLP dependency for each request +❌ Cold start latency (~500ms per query on CPU) +``` + +### PhoW2V (Fully Precomputable) +``` +✅ Load KeyedVectors once at Worker start +✅ Precompute all 22k embeddings into in-memory dense matrix +✅ Cosine similarity: ~1ms per pair (vector dot product) +✅ Alternative: Ship (22k × 22k) similarity matrix in KV +✅ Or: ~7.3 MB dense matrix (22k × 300 × 4 bytes) fits in Worker bundles +``` + +**Verdict:** PhoW2V enables a **stateless, zero-latency** Semantle implementation. sup-SimCSE requires external inference infrastructure. + +--- + +## Vocabulary Coverage + +| Model | Vocab Size | 22k Viet22K Coverage | License | +|-------|------------|---------------------|---------| +| **PhoW2V** | ~100k (estimated from 20GB corpus) | Likely 95%+ (VinAI trained on broad Vietnamese text) | AGPL-3.0; research/education only; cite EMNLP-2020 | +| **sup-SimCSE** | Unbounded (subword + BPE) | 100% (BPE handles unknowns) | Likely permissive (HF model) | + +**Gotcha:** PhoW2V is **research-only, non-commercial**, and requires citation. Check your license constraints for a Telegram bot (even if private, may still violate terms). + +--- + +## Semantic Quality + +### sup-SimCSE +- **Advantage:** Trained on supervised sentence pairs; captures deeper semantic relationships. +- **Disadvantage:** Trained on sentence context; single-word pairs don't benefit from that context. +- **Effective for:** Semantically distant word pairs (e.g., "xe" vs "cách"); may over-regularize tight synonyms. + +### PhoW2V +- **Advantage:** Word-level training (CBOW/Skip-gram); embeddings encode co-occurrence statistics. +- **Disadvantage:** No supervised signal; relies purely on distributional similarity. +- **Effective for:** "Natural" word similarity (synonyms, related concepts); well-suited to Semantle-style games. + +**For a word-guessing game:** Both are reasonable. PhoW2V's simplicity is not a weakness here; it's a feature. + +--- + +## Implementation Complexity + +### sup-SimCSE (High Complexity) +```python +from sentence_transformers import SentenceTransformer +from pyvi.ViTokenizer import tokenize + +model = SentenceTransformer('VoVanPhuc/sup-SimCSE-...') +# Per query: +segmented = tokenize(raw_input) # Runtime overhead +emb = model.encode(segmented) +similarity = cosine(emb_target, emb_guess) +``` +- **Dependencies:** sentence-transformers, transformers, torch, pyvi. +- **Latency:** 200–500ms per query (even on GPU; Cloudflare Workers have no GPU). +- **Lines of code:** ~20. +- **External service:** Optional (could self-host, but adds infrastructure). + +### PhoW2V (Low Complexity) +```python +from gensim.models import KeyedVectors +import numpy as np + +kv = KeyedVectors.load_word2vec_format('phow2v.bin') +# Option A (precompute all): +embeddings = {word: kv[word] for word in vocab} + +# Per query: +similarity = np.dot(embeddings[target], embeddings[guess]) +``` +- **Dependencies:** gensim (tiny). +- **Latency:** <1ms per query (in-memory lookup). +- **Lines of code:** ~10. +- **External service:** None (pure static embeddings). + +--- + +## License & Attribution + +| Model | License | Restriction | +|-------|---------|------------| +| **sup-SimCSE** | Unclear (check HF model card) | Likely permissive for research | +| **PhoW2V** | AGPL-3.0 | **Research/education only; cite EMNLP-2020; non-commercial** | + +**Risk:** If doantu is a commercial Telegram bot or intends to be monetized, PhoW2V's AGPL restriction may be a blocker. Clarify with user. + +--- + +## Unresolved Questions + +1. **PhoW2V license constraints:** Is doantu commercial? Non-commercial? Verify AGPL-3.0 compatibility with your bot's intended use. +2. **Vocabulary overlap:** Exact coverage of 22k Viet22K words in PhoW2V. Could spot-check a few compounds like "máy bay", "con chó" in the model. +3. **Syllable vs word PhoW2V:** Recommendation assumes word-level variant. If Viet22K uses syllables, syllable-level variant may be needed; would require preprocessing. +4. **sup-SimCSE alternatives:** Are you open to other sentence transformers fine-tuned for Vietnamese word similarity (e.g., from FPTAI or other VN NLP labs)? +5. **Similarity matrix size:** Confirm whether shipping a precomputed (22k × 22k) matrix in KV is practical (~7.3 MB in dense form, ~200 MB in sparse COO). + +--- + +## Recommendation Summary + +**Use PhoW2V (300d word-level variant) for doantu.** + +**Why:** +- Single-word embeddings (not sentence-level). +- Static vectors → zero inference cost. +- Fits Cloudflare Worker budget (no external service needed). +- Precomputable into O(1) lookups. +- Simpler to deploy and maintain. + +**Why not sup-SimCSE:** +- Sentence-level training doesn't benefit single-word pairs. +- Runtime inference infeasible on CPU-only Cloudflare Workers. +- External segmentation (pyvi/VnCoreNLP) adds complexity and latency. +- 768-dim vectors overkill for word pairs; 300-dim sufficient. + +**Action items:** +1. Verify PhoW2V's AGPL-3.0 license permits your bot's use case. +2. Spot-check PhoW2V vocabulary against 22k-word game list (OOV strategy needed). +3. Decide precomputation strategy: in-memory matrix, KV store, or on-demand dot product. + +--- + +## Sources + +- [sup-SimCSE-VietNamese-phobert-base on Hugging Face](https://huggingface.co/VoVanPhuc/sup-SimCSE-VietNamese-phobert-base) +- [PhoW2V GitHub Repository](https://github.com/datquocnguyen/PhoW2V) +- [PhoBERT: Pre-trained Language Models for Vietnamese (EMNLP-2020 Findings)](https://aclanthology.org/2020.findings-emnlp.92.pdf) +- [VinAI Research – PhoBERT Overview](https://www.vinai.io/phobert-the-first-public-large-scale-language-models-for-vietnamese/) +- [Gensim Word2Vec KeyedVectors Documentation](https://radimrehurek.com/gensim/models/word2vec.html) +- [Semantle Word Embeddings Recreation](https://github.com/memgonzales/semantle-word-embeddings) +- [VnCoreNLP Word Segmentation](https://www.researchgate.net/publication/325449322_VnCoreNLP_A_Vietnamese_Natural_Language_Processing_Toolkit)