110 Commits
Author SHA1 Message Date
tiennm99 03b1400992 docs: index merged crawlers and point their module paths at mttools 2026-10-03 11:16:25 +07:00
tiennm99 e0fb36cbdf docs: require Go 1.25 to match go.mod 2026-08-31 15:25:31 +07:00
tiennm99 06840c8e6f feat(pdf): export a novel as a phone-readable PDF
Add a pdfout package that renders the fetched chapters to a single PDF
and an export package holding the whole pipeline, so the CLI and an
embedding program share one code path and reject the same inputs.

The default page is 90x160mm rather than A4 with large type: a viewer
scales a whole page to fit the screen, so on a phone the page shape
decides legibility and a wide page just gets shrunk.

Embed DejaVu Sans in the binary. Vietnamese needs the Latin Extended
Additional block, and a headless host often has no fonts installed at
all, so rendering must not depend on finding one. An explicitly named
font is never silently substituted. Font data is passed as bytes because
fpdf resolves a font path against its own font directory.

Keep the original per-chapter text output available behind -txt.
2026-08-24 20:40:01 +07:00
tiennm99 50bbeb7c3c refactor!: generalise single-novel script into a hako library
The crawler took one hardcoded novel URL, wrote per-chapter text files
into ./data, and fetched serially until the site answered 429.

Rename the module to hako-crawler and split fetching, page parsing and
crawl orchestration into a hako package that takes any novel URL.

The chapter body is no longer served as markup: #chapter-content now
carries a shuffled, XOR-obfuscated payload that the site's own script
expands in the browser, so the previous p[id=<digits>] extractor
returned nothing at all. Decode it, keeping the plain-markup path as a
fallback. The scheme is read from the page and an unrecognised one is a
hard error, because the alternative is exporting a book of mojibake.

Space requests globally instead of per-goroutine, so concurrency no
longer raises the request rate, and retry only what a retry can fix.
Cache raw pages so re-runs cost no requests.

BREAKING CHANGE: module path is now github.com/tiennm99/hako-crawler,
and the novel URL is a required argument rather than a constant.
2026-08-24 20:39:50 +07:00
tiennm99 e56338e5e3 feat: parse novel tags and add a single-request metadata fetch
Novel gains Tags, read from the anchors carrying schema.org itemprop="genre".
Matching that microdata rather than the /the-loai/ href shape is what keeps the
site-wide genre navigation out: a sampled page links 69 categories against the
novel's own 6.

NovelInfo fetches just the landing page and parses it, so a caller that only
wants metadata pays one request instead of the two Novel needs for its chapter
list cross-check.
2026-07-30 00:45:20 +07:00
tiennm99 5b213da6be feat: lower the default body font size to 10pt
On the default 90x160mm phone page, 12pt fits about 36 characters per line;
10pt fits about 43 across 26 lines, which reads denser without becoming small
on a phone at 100% zoom. -font-size still overrides it.
2026-07-30 00:10:20 +07:00
tiennm99 beada1d17f fix: embed font data instead of a path, and bundle a fallback font
PDF rendering failed on Linux with "stat usr/share/fonts/...: no such file or
directory": fpdf joins the font path onto its own font directory, which it
defaults to ".", so path.Join turns an absolute path into a
working-directory-relative one. It only resolved when the process happened to
run from the filesystem root, which is why Windows was unaffected. Font data is
now read by pdfout and handed over as bytes.

LoadFont also falls back to a bundled DejaVu Sans, so a host with no fonts
installed still renders. An explicitly requested font remains a hard error when
unreadable rather than being silently substituted. Tests verify the bundled font
parses and covers Vietnamese, which is the reason it exists.
2026-07-29 23:41:30 +07:00
tiennm99 11207bb051 refactor: make crawler importable and add export pipeline
Move monkeyd and pdfout out of internal/ so other modules can import them,
and add export.Export, which holds the crawl-to-PDF sequence the CLI used to
inline. The CLI now parses flags and delegates, so an embedding program gets
the same defaults and validation.
2026-07-29 22:49:30 +07:00
tiennm99 8eb55da485 fix: drop sponsor, watermark and navigation blocks from chapter text
Three blocks the site injects inside the chapter body were being extracted
as prose: the sponsor block, a source watermark planted mid-chapter, and
the prev/next chapter links.

Anchors and images are now skipped, since inside a chapter body they are
always site chrome, and the sponsor wrapper and watermark are skipped by
class. Inline emphasis tags are left alone so italics in the prose survive.

The sponsor wrapper has a sibling that must NOT be skipped: on most
chapters the real text sits in a container hidden with display:none,
because the body is gated behind a click on the sponsor link and revealed
with JavaScript. Skipping hidden elements, or treating the two sibling
classes as one family, discards the entire chapter. A test pins the gated
body as kept while its sibling is dropped.

Removes 662 words of injected text across the sample novel and lets the
repeated chapter-number paragraph be recognised again, which the sponsor
block had been masking by displacing it from the top of the body.
2026-07-29 22:36:15 +07:00
tiennm99 1da14cd8ab feat: add monkeydd novel crawler with phone-sized PDF export
Fetches every chapter of a monkeydd.com novel and renders it as a single
PDF laid out for reading on a phone.

Two site behaviours drive the extractor design:

- Roughly a fifth of each chapter's words are not in the markup. The page
  emits empty spans and supplies the word from the stylesheet via
  ":before { content: ... }" rules, so reading DOM text alone drops them
  with no error. The extractor resolves those rules and substitutes the
  words back; a test asserts they disappear when the rule is removed.
- Chapter URLs cannot be generated. Numbering has gaps and slugs are not
  uniform across novels, so chapter links are always parsed from the page.

The chapter list is read from the landing page and cross-checked against
the dropdown embedded in each chapter page, so a truncated list cannot
silently shorten the export.

PDF defaults to a 90x160mm page rather than A4 with large type: viewers
scale a whole page to fit the screen, so a phone-shaped page fills it at
100% zoom where 12pt stays comfortable. A5 and A4 remain available.

Requests are spaced globally, so raising worker count does not raise the
request rate. Raw pages cache to disk so re-exporting at different font
or page settings needs no network.
2026-07-29 22:26:25 +07:00
tiennm99 5cc7df43a5 Initial commit 2026-07-29 21:50:24 +07:00
tiennm99 94f146810b feat(go): rewrite in Go 2025-11-30 20:22:58 +07:00
tiennm99 3a341a2bf6 [Init] 2024-10-11 21:03:49 +07:00
Tien Nguyen Minh 2a6d11f617 Initial commit 2024-10-11 19:54:02 +07:00
tiennm99 d6af707b9c Revert "Merge branch 'feature/circular-dependency-checker-go' into main"
Keeps the branch in history only; main's content is unchanged.
2026-09-30 20:04:14 +07:00
tiennm99 da88525bd8 Merge branch 'feature/circular-dependency-checker-go' into main 2026-09-30 20:04:14 +07:00
tiennm99 5c772b0273 Revert "Merge branch 'feature/export-chrome-cookies-go' into main"
Keeps the branch in history only; main's content is unchanged.
2026-09-30 20:04:13 +07:00
tiennm99 a7f20ae704 Merge branch 'feature/export-chrome-cookies-go' into main 2026-09-30 20:04:13 +07:00
tiennm99 2e30783a22 Revert "Merge branch 'feature/bulk-telegram-sticker-go' into main"
Keeps the branch in history only; main's content is unchanged.
2026-09-30 20:04:13 +07:00
tiennm99 ceb3904327 Merge branch 'feature/bulk-telegram-sticker-go' into main 2026-09-30 20:04:13 +07:00
tiennm99 eb63c418c0 docs: index merged tools and point their module paths and links at mttools 2026-09-30 17:10:08 +07:00
tiennm99 aef4274943 chore: stop tracking the installed dependency tree
Nodejs/node_modules was committed (4,275 files), which is also how three
upstream yarn.lock files ended up in the repository. package-lock.json already
pins everything, and all five dependencies still resolve from the registry, so
the tree is reproducible with npm ci and does not need to be stored here.

Add a .gitignore; the repository had none.
2026-08-17 16:58:56 +07:00
tiennm99 802c353def feat: replace Python and JS downloaders with a Go crawler
Port both scripts into one Go program with a command per behavior:
gallery walks the numbered ghibli.jp film galleries, scrape downloads
every image referenced by a page.

Downloads now run through a bounded worker pool, and non-200 responses
are no longer written to disk as image files. Drop the orphaned
package-lock.json, which pinned the deprecated request dependency and
was the source of the repository's Dependabot alerts.
2026-07-25 19:41:43 +07:00
tiennm99 31d0d3733b docs(readme): rename project to ghibli-gallery-crawler 2026-07-25 19:28:02 +07:00
tiennm99 57a06c60eb chore: remove dependabot version-update config 2026-07-25 14:13:41 +07:00
tiennm99 eeb154c5b1 fix(deps): bump golang.org/x/net to patch security advisories 2026-07-17 17:42:06 +07:00
tiennm99 e0251c71e6 chore: drop inapplicable dependabot ecosystems (#2) 2026-05-23 11:09:24 +07:00
tiennm99 a3f04b6dfe chore: add dependabot config (#1) 2026-05-23 10:48:19 +07:00
tiennm99 b4e1131dac docs(readme): add input format, output behavior, and variant notes 2026-05-11 21:42:38 +07:00
tiennm99 80023a9a6c chore: relicense from MIT to Apache-2.0 (sole-author repo) 2026-05-11 17:15:43 +07:00
tiennm99 2f366a2b7f docs: add README 2026-05-11 17:04:19 +07:00
tiennm99 dfa9698881 chore: add Apache-2.0 license 2026-04-29 21:33:48 +07:00
tiennm99 976137eea8 chore: migrate 2025-12-21 10:16:18 +07:00
tiennm99 e669898f97 feat(go): update desc 2025-11-30 21:36:32 +07:00
tiennm99 9f27ea6a85 feat(go): update deps 2025-11-30 20:57:33 +07:00
tiennm99 8fcc74f6eb feat(go): rewrite in go 2025-11-30 20:52:54 +07:00
tiennm99 1b5284af18 Update README.md 2025-11-29 13:19:06 +07:00
tiennm99 c35e2dc9d9 feat: add note 2025-11-29 13:11:13 +07:00
tiennm99 6a1eb53102 feat(go): convert to Go 2025-11-29 13:00:05 +07:00
tiennm99 c9dde6c993 Update .gitignore 2025-11-28 22:20:47 +07:00
tiennm99 841006a59f Update main.go 2025-11-28 22:05:06 +07:00
tiennm99 7b01dce943 feat(go): rewrite in Go 2025-11-28 21:52:05 +07:00
tiennm99 9615e32754 Update README.md 2025-11-28 21:50:47 +07:00
tiennm99 ba0d31680e feat(go): rewrite in Go 2025-11-28 20:59:16 +07:00
tiennm99 947178086d Update main.py 2025-07-20 18:59:06 +07:00
tiennm99 eee2999d9a feat: init 2025-03-28 20:35:09 +07:00
Tien Nguyen Minh 0935eae148 Initial commit 2025-03-28 18:45:40 +07:00
tiennm99 f5ac188f77 Update main.sh 2024-05-22 21:04:10 +07:00
tiennm99 a5c643133e Update main.sh 2024-05-22 21:00:16 +07:00
tiennm99 5b462e3731 Update main.sh 2024-05-22 20:46:10 +07:00