diff --git a/.github/workflows/deploy-pages.yml b/.github/workflows/deploy-pages.yml index 4dc3f31..b121217 100644 --- a/.github/workflows/deploy-pages.yml +++ b/.github/workflows/deploy-pages.yml @@ -121,3 +121,30 @@ jobs: steps: - id: deployment uses: actions/deploy-pages@v5 + + - uses: actions/checkout@v7 + + # The site reads the databases a page at a time over range requests, so + # what matters is not that the host leaves the file alone in general — it + # gzips the un-ranged response, and browsers work around that by sending + # Accept-Encoding: identity whenever a request carries a Range header — + # but that a ranged read returns raw database bytes. Checking headers is + # what missed this before: a bare `curl -sI` advertises no encoding and + # so passes whatever the host does. Check the bytes instead. + - name: Verify ranged reads return database bytes + env: + PAGE_URL: ${{ steps.deployment.outputs.page_url }} + run: | + set -euo pipefail + for id in $(jq -r '.datasets[].id' datasets.json); do + url="${PAGE_URL%/}/db/${id}.sqlite3" + if ! magic=$(curl -sf -r 0-14 -H 'Accept-Encoding: identity;q=1, *;q=0' "$url"); then + echo "::error::$url is not fetchable" + exit 1 + fi + if [ "$magic" != "SQLite format 3" ]; then + echo "::error::$url did not return database bytes over a range request" + exit 1 + fi + echo "$url: SQLite format 3" + done diff --git a/CLAUDE.md b/CLAUDE.md index e640750..fd837ef 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -61,8 +61,15 @@ hashes every real input file. That is the point of it; do not skip it. assembler refuses to publish an artifact that falls below a ratio of it. - **The databases ship uncompressed, as `.sqlite3`.** The browser reads byte ranges of them, and a range of a gzip stream is not a range of the - database. The host must not apply `Content-Encoding` either — check with - `curl -sI` after a deploy. + database. +- **The file length comes from a range request, not from the host's HEAD.** + GitHub Pages gzips `application/octet-stream`, so a HEAD reports the + compressed size and `sql.js-httpvfs` refuses to open the file. Ranged reads + are unaffected — browsers must send `Accept-Encoding: identity` whenever a + request carries a `Range` header — so `web/src/lib/db-probe.js` reads the + header over a range and passes `fileLength`. Verify the way a browser asks: + `curl -sI -r 0-99 -H 'Accept-Encoding: identity' …`, never a bare `curl -sI`, + which advertises no encoding and hides the problem. - **Every query the site runs must be index-driven.** Over range requests an unindexed query fetches the whole table. Hence no index on `ho_ten` (nothing can use one), `name_word` for name search, partial indexes for the score diff --git a/docs/deployment-guide.md b/docs/deployment-guide.md index b34a9dc..8d01024 100644 --- a/docs/deployment-guide.md +++ b/docs/deployment-guide.md @@ -87,10 +87,23 @@ artifact — one missing line away from publishing it. - **Total artifact is about 552 MB**, inside the 1 GB site limit but with less headroom than before: a third dataset of this size would not fit. The fallback is `sql.js-httpvfs`'s chunked mode, which splits a database into parts. -- **The server must not compress the databases.** Ranges of a compressed body - address the wrong bytes, and the library refuses to open a file whose HEAD - carries a `Content-Encoding`. `.sqlite3` is an unknown type to Pages, so it is - served as `application/octet-stream` and left alone — verify after a deploy. +- **Pages does compress the databases, and that is survivable.** `.sqlite3` is + unknown to Pages, so it is served as `application/octet-stream`, which is + marked compressible in `mime-db` and gzipped: a plain request returns + `Content-Encoding: gzip` and the compressed length. Ranged reads are not + affected, because the Fetch standard makes browsers send + `Accept-Encoding: identity` on any request carrying a `Range` header. Only + the length probe breaks, and the site supplies `fileLength` itself instead of + trusting HEAD — see `web/src/lib/db-probe.js`. +- **Verify the way a browser asks.** A bare `curl -sI` advertises no encoding + and so reports success whatever the host does; it is what let this reach + production. Check ranged reads instead, and check the bytes, not the headers: + + ```bash + curl -s -r 0-15 -H 'Accept-Encoding: identity;q=1, *;q=0' \ + https://.github.io/thptqg/db/2016.sqlite3 | head -c 16 + # must print: SQLite format 3 + ``` ## Rollback @@ -104,7 +117,8 @@ run rebuilds the older state. There is no data to migrate. | Blank page, 404 on assets | `paths.base` in `svelte.config.js` does not match the repo name | | `Failed to fetch database: 404` | Dataset id in `datasets.json` does not match the file in `db/` | | A route 404s | The site step did not run, or the id is missing from `datasets.json` | -| Database fails to open | The host compressed it. `curl -sI …/db/.sqlite3` must show no `content-encoding`; ranges of a compressed body are unusable | +| `Length of the file not known` | The host gzipped the un-ranged response, so HEAD reports the compressed size. The site supplies `fileLength` from a range probe; if this returns, that probe failed | +| Database fails to open | A ranged read did not return raw database bytes. The range check above must print `SQLite format 3` | | Every query is slow or huge | It is not using an index. `EXPLAIN QUERY PLAN` it: a `SCAN` means the browser is fetching the whole table | | Deploy fails on assembly | An uncompressed database artefact reached the output; the error names the files | | Missing rows after a data update | Unknown Excel header — check the per-file row counts the parser prints | diff --git a/docs/system-architecture.md b/docs/system-architecture.md index d0a027c..5eab684 100644 --- a/docs/system-architecture.md +++ b/docs/system-architecture.md @@ -204,10 +204,15 @@ total descending. - **Unindexed queries are expensive.** The SQL tab can express a query that walks the table, which over range requests means fetching 100+ MB. A byte budget stops one before it gets that far, and the tab warns before it opens. -- **`Content-Encoding` breaks everything.** If the host ever compresses - `.sqlite3` on the wire, ranges address compressed bytes and - `sql.js-httpvfs` refuses to open the file. Verify after a deploy: - `curl -sI …/db/2016.sqlite3` must show no `content-encoding`. +- **`Content-Encoding` on a ranged response would break everything.** A range + of a compressed body addresses the wrong bytes. In practice browsers prevent + it: the Fetch standard requires `Accept-Encoding: identity` on any request + carrying a `Range` header. GitHub Pages *does* gzip the un-ranged response — + `application/octet-stream` is compressible in `mime-db` — which is why the + file length is probed with a range request and passed as `fileLength` rather + than left to the library's HEAD. `db-probe.js` checks the returned bytes + start with the SQLite magic, so a host that ever compresses a ranged response + fails loudly instead of returning nonsense. - **`sql.js-httpvfs` is unmaintained** (0.8.12, September 2022) and ships its own SQLite WASM. `sqlite-wasm-http`, on the official build, is the fallback. - **Hosted size.** 552 MB for both datasets against the 1 GB GitHub Pages diff --git a/web/src/lib/db-probe.js b/web/src/lib/db-probe.js new file mode 100644 index 0000000..2312c54 --- /dev/null +++ b/web/src/lib/db-probe.js @@ -0,0 +1,88 @@ +/** + * How long the database is, asked in the one way that survives a CDN. + * + * `sql.js-httpvfs` sizes a file with a HEAD request. That request carries no + * Range header, so the browser advertises gzip, and GitHub Pages answers with + * `Content-Encoding: gzip` and the length of the *compressed* body — 66 MB for + * a 302 MB database. The library rightly refuses to believe it and gives up + * with "Length of the file not known. It must either be supplied in the config + * or given by the HTTP server." + * + * Range requests do not have that problem: the Fetch standard requires + * `Accept-Encoding: identity` on any request carrying a Range header, so the + * page reads that do the actual work always come back uncompressed. Asking for + * the first hundred bytes therefore yields both a trustworthy total, from + * Content-Range, and the file header itself to check it against. + */ + +/** A SQLite file opens with these characters and then a NUL byte. */ +const MAGIC = "SQLite format 3"; + +/** Enough for the whole SQLite header. */ +const PROBE_BYTES = 100; + +/** Page size lives at offset 16, big-endian; the value 1 encodes 65536. */ +const PAGE_SIZE_OFFSET = 16; + +/** Total size of the representation, from `bytes -/`. */ +export function parseTotalBytes(contentRange) { + const match = /\/\s*(\d+)\s*$/.exec(contentRange ?? ""); + if (!match) { + throw new Error( + `the server did not say how large the database is (Content-Range: ${contentRange ?? "absent"})`, + ); + } + return Number(match[1]); +} + +/** Page size the file was written with, from its header. */ +export function readPageSize(header) { + const raw = (header[PAGE_SIZE_OFFSET] << 8) | header[PAGE_SIZE_OFFSET + 1]; + return raw === 1 ? 65536 : raw; +} + +/** True when these bytes begin a SQLite database. */ +export function looksLikeSqlite(header) { + const text = Array.from(MAGIC).every((ch, i) => header[i] === ch.charCodeAt(0)); + return text && header[MAGIC.length] === 0; +} + +/** + * Read the file header over a range request and return the database's length. + * + * Doubles as the check that the host is serving raw database bytes: a body that + * does not start with the SQLite magic means something rewrote it in transit — + * compression being the way that happens — and every later page read would be + * reading the wrong bytes. + */ +export async function probeDatabase(url, expectedPageSize, fetchImpl = fetch) { + const response = await fetchImpl(url, { headers: { Range: `bytes=0-${PROBE_BYTES - 1}` } }); + + if (response.status !== 206) { + throw new Error( + `${url}: expected 206 for a range request, got ${response.status}. ` + + "The host must serve byte ranges of the database.", + ); + } + + const total = parseTotalBytes(response.headers.get("Content-Range")); + const header = new Uint8Array(await response.arrayBuffer()); + + if (!looksLikeSqlite(header)) { + throw new Error( + `${url}: the first bytes are not a SQLite header, so the host is not ` + + "serving the database as stored — check for Content-Encoding on ranged responses.", + ); + } + + const pageSize = readPageSize(header); + if (expectedPageSize && pageSize !== expectedPageSize) { + // Not fatal: it still reads correctly, just at more requests per page. + console.warn( + `[httpvfs] ${url} has page size ${pageSize}, but requests are ${expectedPageSize} bytes. ` + + "Every page read now spans more than one request.", + ); + } + + return total; +} diff --git a/web/src/lib/db-probe.test.js b/web/src/lib/db-probe.test.js new file mode 100644 index 0000000..698425a --- /dev/null +++ b/web/src/lib/db-probe.test.js @@ -0,0 +1,90 @@ +import { describe, expect, it, vi } from "vitest"; +import { looksLikeSqlite, parseTotalBytes, probeDatabase, readPageSize } from "./db-probe.js"; + +/** The first bytes of a real database: magic, NUL, then the page size at 16. */ +function header(pageSize = 1024, magic = "SQLite format 3") { + const bytes = new Uint8Array(100); + for (let i = 0; i < magic.length; i += 1) bytes[i] = magic.charCodeAt(i); + bytes[16] = pageSize >> 8; + bytes[17] = pageSize & 0xff; + return bytes; +} + +function respond(bytes, { status = 206, contentRange = `bytes 0-99/${317096960}` } = {}) { + return { + status, + headers: { get: (name) => (name.toLowerCase() === "content-range" ? contentRange : null) }, + arrayBuffer: async () => bytes.buffer, + }; +} + +describe("parseTotalBytes", () => { + it("takes the total from a Content-Range", () => { + expect(parseTotalBytes("bytes 0-99/317096960")).toBe(317096960); + }); + + it("refuses an unknown total", () => { + expect(() => parseTotalBytes("bytes 0-99/*")).toThrow(/how large/); + expect(() => parseTotalBytes(null)).toThrow(/absent/); + }); +}); + +describe("readPageSize", () => { + it("reads the two big-endian bytes at offset 16", () => { + expect(readPageSize(header(1024))).toBe(1024); + expect(readPageSize(header(4096))).toBe(4096); + }); + + it("treats 1 as 65536, as the file format does", () => { + expect(readPageSize(header(1))).toBe(65536); + }); +}); + +describe("looksLikeSqlite", () => { + it("accepts a real header", () => { + expect(looksLikeSqlite(header())).toBe(true); + }); + + it("rejects a gzip stream, which is what a compressing host returns", () => { + const gzip = new Uint8Array([0x1f, 0x8b, 0x08, 0x00, 0x00, 0x00, 0x00, 0x00]); + expect(looksLikeSqlite(gzip)).toBe(false); + }); + + it("requires the NUL that terminates the magic", () => { + expect(looksLikeSqlite(header(1024, "SQLite format 3x"))).toBe(false); + }); +}); + +describe("probeDatabase", () => { + it("asks for the header by range and returns the total length", async () => { + const fetchImpl = vi.fn(async () => respond(header())); + const total = await probeDatabase("/db/2016.sqlite3", 1024, fetchImpl); + + expect(total).toBe(317096960); + expect(fetchImpl).toHaveBeenCalledWith("/db/2016.sqlite3", { + headers: { Range: "bytes=0-99" }, + }); + }); + + it("fails when the host ignores the range", async () => { + const fetchImpl = async () => respond(header(), { status: 200 }); + await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).rejects.toThrow(/expected 206/); + }); + + it("fails when the bytes are not a database", async () => { + const gzip = new Uint8Array([0x1f, 0x8b, 0x08]); + const fetchImpl = async () => respond(gzip); + await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).rejects.toThrow( + /not a SQLite header/, + ); + }); + + it("warns, but continues, when the page size is not the request size", async () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + const fetchImpl = async () => respond(header(4096)); + + await expect(probeDatabase("/db/2016.sqlite3", 1024, fetchImpl)).resolves.toBe(317096960); + expect(warn).toHaveBeenCalledWith(expect.stringMatching(/page size 4096/)); + warn.mockRestore(); + }); +}); diff --git a/web/src/lib/sqlite.svelte.js b/web/src/lib/sqlite.svelte.js index 19cdcaa..65868f7 100644 --- a/web/src/lib/sqlite.svelte.js +++ b/web/src/lib/sqlite.svelte.js @@ -1,6 +1,7 @@ import { createDbWorker } from "sql.js-httpvfs"; import workerUrl from "sql.js-httpvfs/dist/sqlite.worker.js?url"; import wasmUrl from "sql.js-httpvfs/dist/sql-wasm.wasm?url"; +import { probeDatabase } from "./db-probe.js"; /** * The database is read where it lies. SQLite asks for pages, the virtual file @@ -61,8 +62,17 @@ export class RemoteDatabase { async #open() { const opened = performance.now(); try { + // Supplied rather than left to the library, which would size the file + // with a HEAD request. GitHub Pages compresses that response and reports + // the compressed length, which the library refuses to use. See db-probe. + const fileLength = await probeDatabase(this.url, CHUNK_BYTES); const worker = await createDbWorker( - [{ from: "inline", config: { serverMode: "full", url: this.url, requestChunkSize: CHUNK_BYTES } }], + [ + { + from: "inline", + config: { serverMode: "full", url: this.url, requestChunkSize: CHUNK_BYTES, fileLength }, + }, + ], workerUrl, wasmUrl, this.budgetBytes,