tiennm99 b459b1f5e9 fix: re-read each message for a live file reference before downloading it
The walk collects every item's file reference up front, and Telegram expires
those. The lifetime is undocumented; a run was seen to outlast an hour and fail
before two, which is less than a large chat spends downloading. So every fetch
past that point returned FILE_REFERENCE_EXPIRED, five in a row tripped the
breaker, and the run abandoned the rest of its todo list.

Telegram's contract is to cache a reference together with the source it came
from and re-read that source when it expires, so each message is re-read for a
live reference immediately before its own download. That is one round trip per
transfer and needs no guess at how long a reference lives. It runs after the
staging acquire rather than before: that acquire is the run's brake and blocks
for as long as the destination is slow, which on a stalled remote is long enough
to expire a fresh reference all over again.

A re-read that describes a different file — different name or different length —
is refused, because archiving it under the name this run recorded would store
the wrong bytes. That case, a deleted message, and one that no longer holds a
file are all skipped: nothing can be archived for them, and one must not strand
the rest of the chat.

Any other re-read failure is recorded as a failed download instead, because that
is what it is. The connection that serves the re-read is the one that serves the
file, so these arrive exactly when transfers are failing too. An item that never
reaches a worker never reaches OnAdd or OnDone, so on the skip path it would
move no counter, feed no streak and produce no outcome to retry — a source
outage would walk the whole list one dead round trip at a time and still exit as
merely incomplete.

tutil.GetSingleMessage is not reused for the re-read: it reports a deleted
message three different ways, two of them wrapping a nil error, and only one is
distinguishable — so deletions would be misclassified as transient and spend the
breaker's streak.
2026-09-08 22:47:52 +07:00
2026-09-06 20:51:08 +07:00
2026-08-11 14:07:26 +07:00

telegram-exporter

Archive a Telegram chat's media to any rclone remote — S3, Google Drive, Dropbox, Backblaze B2, SFTP, WebDAV, pikpak, or anything else rclone supports — using far less local disk than the chat's total size.

tgexport embeds tdl and rclone as libraries and runs both halves in one process. Files are downloaded into a small staging directory and uploaded the moment each one finishes, so local disk only ever holds what is in flight. A multi-terabyte chat archives fine on a small disk. Telegram caps a single file at 2 GB (4 GB from premium uploaders), so a few dozen GB of staging covers the worst case regardless of chat size.

This exists because tdl can only write to a local directory — it has no remote destination of any kind (tdl dl -d takes a filesystem path; tdl --storage is its session database, not an output target). tdl downloads over MTProto with a user account, so Bot API limits do not apply: full history is readable and there is no 20 MB download cap.

Requirements

Neither tool is invoked at run time; tgexport reads the session and config they write.

Setup

Both steps are one-time.

tdl login                 # writes the Telegram session tgexport reads
rclone config             # define the destination remote
go build -o tgexport ./cmd/tgexport
./tgexport doctor -r myremote:archive

doctor proves both halves work before a long run: it prints the logged-in account, resolves the destination, and reports free space.

Build variants

Build Backends Size
go build ./cmd/tgexport every rclone backend ~92 MB
go build -tags slim ./cmd/tgexport pikpak only ~49 MB

A backend that is not compiled in does not exist at run time, so use the default build unless the destination will never change.

Usage

./tgexport sync -c CHAT -r REMOTE:PATH [options]

A run states what it found, what it is about to do, and then shows each file as it moves:

reading mychannel
  12,000 messages read in 4m31s
indexing PikPak root 'mychannel'
  11,406 objects listed in 2m10s

  chat holds     12,000 media, 250.0 GiB
  archived       11,400
    never fetched 600
    wrong size    6

  fetching       606 files, 79.0 GiB (largest 2.0 GiB)
  into           PikPak root 'mychannel'
  staging        ./staging, capped at 40.0 GiB
  concurrency    2 download(s) x 4 thread(s), 2 upload(s)

  ↓ total    26/606 files [=>             ] 617.5 MiB / 79.0 GiB  2.5 MiB/s  8h47m
  ↑ total    24/606 files [=>             ] 598.0 MiB / 79.0 GiB  2.4 MiB/s  8h58m
  ↓ …7890_4242_1000000000000000001.mp4   [=======>       ]  41.2 MiB / 96.0 MiB  1.8 MiB/s
  ↑ …7890_4243_1000000000000000002.mp4   ⠹                         uploading 1.9 GiB

The two legs are counted separately because they run at different speeds and fail for different reasons. They normally track a file or two apart; a widening gap means the remote is falling behind and staging is filling up.

Redirected output gets the same two figures as plain periodic lines, plus one line per archived file, with no cursor movement — a captured log stays readable:

  download 1,200/606 files, 45.0 GiB of 79.0 GiB, 76.8 MiB/s, ETA 7m33s
  upload   1,190/606 files, 44.2 GiB of 79.0 GiB, 75.4 MiB/s, ETA 7m53s

CHAT accepts a numeric id as printed by tdl chat ls, a username with or without @, or a t.me/tg:// link. A Bot API -100… id is converted automatically. A link to a single message is refused — it names a message, not a chat.

# archive a chat, capping staging at 40 GiB
./tgexport sync -c @mychannel -r gdrive:telegram/media -m 40G

# check completeness without downloading anything
./tgexport verify -c @mychannel -r gdrive:telegram/media

# list what the chat holds
./tgexport list -c @mychannel

Options

Flag Default Meaning
-c — chat id, username, or link (required)
-r — rclone destination, REMOTE:PATH (required)
-d ./staging staging directory for files in flight
-m no cap cap staging at a size, e.g. 40G
--threads 4 connections per file
--limit 2 files downloading at once
--uploads 2 files uploading at once
--min-free 5 stop if the remote has fewer than this many GiB free, checked before and during the run
--limit-items 0 stop after N files; for smoke tests
--confirm true re-state each uploaded file to prove its size
--takeout true use a takeout session
-n default tdl session namespace

Exit codes

Code Meaning
0 complete
1 ran, but files remain — run again
2 usage error
3 remote or Telegram failure, including either leg refusing transfer after transfer
4 stalled: files remain, none of which can ever be fetched
130 / 143 interrupted (SIGINT / SIGTERM)

Only 1 is worth retrying. A driver looping until 0 should stop on anything else: 3 and 4 both mean the next pass would do exactly what this one did.

How it works

Re-running is the resume path. Each item is checked against a listing of the remote immediately before download, so an interrupted run picks up where it left off and a completed one downloads nothing.

Filenames. Every file is stored as {DialogID}_{MessageID}_{FileName}, where FileName is exactly what Telegram reports. One function derives that string, and the same string is used both to ask whether the file is already archived and to write it — so the two can never disagree.

A name that cannot survive that round trip is refused rather than rewritten: too long for the filesystem once .part is appended, not a single path element, or containing a character rclone's path encoder rewrites (control bytes, DEL, and the encoder's own escape character). Those files are reported under unarchivable and never counted as present. Rewriting them is what the next paragraph is about.

That last point is the reason this program exists. Its predecessor derived the name twice: tdl chat export wrote the raw name into a JSON, while tdl dl rendered it through a template applying filenamify, which rewrites characters a filesystem rejects and collapses runs of !. A file whose name contained !! was looked up under one name and stored under another, so the verifier never found it and re-fetched it on every pass — forever, at 966 MB a time.

Note the consequence: names are not run through filenamify, so they are not byte-compatible with what the old shell pipeline wrote. A file it stored under a rewritten name will not be recognised and gets fetched again.

Disk. -m is a byte budget. A download reserves its own size before starting and releases it only once the upload is confirmed, so when the remote is slow the downloads pause on their own. The cap must exceed the largest single file, and a cap that does not is refused at startup rather than discovered as a hang.

Integrity. A download is written to <name>.part and renamed only once its size matches what Telegram reported, so a file without the suffix is always whole. Uploads are re-stated afterwards to prove they arrived at the right size, before the local copy is gone, and an object that turns out short is deleted rather than left under a name a later run would trust.

verify compares stored sizes against what Telegram reports, so a truncated object is outstanding rather than "present". This is stricter than the shell verifier, which matched on name and non-zero size — on the archive this was built for it found six objects that had been counted complete for months, one of them 221 MiB standing in for a 2 GiB video. Re-running repairs them.

File references. Telegram hands out a short-lived token with every file location, and a run over a large chat outlives the ones its walk collected — the lifetime is undocumented, but an observed run stopped just under two hours in with FILE_REFERENCE_EXPIRED on every remaining file. So each message is re-read for a live token immediately before its own download, which costs one round trip per transfer and needs no guess at how long a token lasts. Retry passes go through the same path, so a token that dies during a single very large transfer is replaced rather than replayed.

The re-read has to describe the same file — same name, same length — or it is not the file this run recorded, and archiving it under that name would store the wrong bytes. A message that fails that check, or that has been deleted, or that no longer holds a file at all, is skipped and reported: nothing can be archived for it, one of them must not strand the rest of the chat, and the closing verify still lists it as outstanding.

A re-read that fails for any other reason is counted as a failed download, not a skip, because that is what it is — the connection that serves the re-read is the one that serves the file. So it is retried with the rest, it shows up in the progress report and the closing failure list, and enough of them in a row trip the same breaker below.

Failures. A failed download is retried inside the run: up to three passes over whatever is still missing, spaced a minute and then two apart. The failure this exists for is the transient one — a connection that dies takes every transfer in flight with it, and all of them are fetchable again minutes later — and leaving them to the next run costs a full re-walk of the chat and a re-index of the remote before a byte can move. Each failure is reported with the reason Telegram gave for it, not just the byte count that arrived.

Either leg gives up once five transfers in a row fail. A source or destination that is refusing everything will refuse the rest of the list too, in milliseconds, so the run stops and exits 3 instead of spending the whole outstanding list finding that out. A run that trips and then recovers on a retry reports nothing.

Replacing the shell pipeline

Earlier versions of this repo were three bash scripts — run.sh, export-until-complete.sh and verify-export.sh — driving tdl and rclone as separate processes. Everything expensive in them existed to work around the fact that neither process could see the other's state: a staging directory polled with du -sk, an --min-age guard, a *.tmp exclusion, SIGSTOP/SIGCONT to enforce the disk cap, a sweep-failure counter, and an outer loop that re-verified and re-narrowed a JSON export between passes.

One process needs none of it. Completion is a function returning; the cap is a semaphore. Some hard-won details were worth keeping, and are:

  • pikpak commits uploads as a server-side async task, and rclone abandons a still-pending one when its low-level retries run out. transfers=2 and low-level-retries=20 are the defaults here for that reason. Environment overrides still win.
  • A backend with no quota API is treated as unlimited, so it never blocks a run.
  • Zero-byte files count as missing — rclone overwrites a size-mismatched destination, so re-running repairs them — while files under 1 KiB are reported but trusted, since some real media genuinely is that small.
  • Indexing ignores RCLONE_* filters. The transfer tunables above are deliberately env-overridable; the listing is not. A stray RCLONE_EXCLUDE or RCLONE_MIN_SIZE left over from another job would otherwise narrow the index and re-download everything it hid.

Notes

  • tgexport and the tdl CLI share one session store and cannot run against the same namespace at once. Use -n for a second namespace if you need both.
  • A partially downloaded file is not resumable across restarts — tdl's library exposes no resume offset — so an interrupted run re-fetches whatever was in flight, bounded by --limit.
  • Everything is read-only against Telegram. Nothing is uploaded, deleted, or marked read.
S
Description
No description provided
Readme Apache-2.0
613 KiB
0 Stars 1 Watchers 0 Forks
Languages
Go 100%