The walk collects every item's file reference up front, and Telegram expires those. The lifetime is undocumented; a run was seen to outlast an hour and fail before two, which is less than a large chat spends downloading. So every fetch past that point returned FILE_REFERENCE_EXPIRED, five in a row tripped the breaker, and the run abandoned the rest of its todo list. Telegram's contract is to cache a reference together with the source it came from and re-read that source when it expires, so each message is re-read for a live reference immediately before its own download. That is one round trip per transfer and needs no guess at how long a reference lives. It runs after the staging acquire rather than before: that acquire is the run's brake and blocks for as long as the destination is slow, which on a stalled remote is long enough to expire a fresh reference all over again. A re-read that describes a different file — different name or different length — is refused, because archiving it under the name this run recorded would store the wrong bytes. That case, a deleted message, and one that no longer holds a file are all skipped: nothing can be archived for them, and one must not strand the rest of the chat. Any other re-read failure is recorded as a failed download instead, because that is what it is. The connection that serves the re-read is the one that serves the file, so these arrive exactly when transfers are failing too. An item that never reaches a worker never reaches OnAdd or OnDone, so on the skip path it would move no counter, feed no streak and produce no outcome to retry — a source outage would walk the whole list one dead round trip at a time and still exit as merely incomplete. tutil.GetSingleMessage is not reused for the re-read: it reports a deleted message three different ways, two of them wrapping a nil error, and only one is distinguishable — so deletions would be misclassified as transient and spend the breaker's streak.
telegram-exporter
Archive a Telegram chat's media to any rclone remote — S3, Google Drive, Dropbox, Backblaze B2, SFTP, WebDAV, pikpak, or anything else rclone supports — using far less local disk than the chat's total size.
tgexport embeds tdl and
rclone as libraries and runs both halves in one process.
Files are downloaded into a small staging directory and uploaded the moment each
one finishes, so local disk only ever holds what is in flight. A multi-terabyte
chat archives fine on a small disk. Telegram caps a single file at 2 GB (4 GB
from premium uploaders), so a few dozen GB of staging covers the worst case
regardless of chat size.
This exists because tdl can only write to a local directory — it has no
remote destination of any kind (tdl dl -d takes a filesystem path; tdl --storage is its session database, not an output target). tdl downloads over
MTProto with a user account, so Bot API limits do not apply: full history is
readable and there is no 20 MB download cap.
Requirements
- Go 1.25+ to build, or a prebuilt binary.
- tdl — only for
tdl login. https://docs.iyear.me/tdl/getting-started/installation/ - rclone — only to configure a remote. https://rclone.org/install/
Neither tool is invoked at run time; tgexport reads the session and config
they write.
Setup
Both steps are one-time.
tdl login # writes the Telegram session tgexport reads
rclone config # define the destination remote
go build -o tgexport ./cmd/tgexport
./tgexport doctor -r myremote:archive
doctor proves both halves work before a long run: it prints the logged-in
account, resolves the destination, and reports free space.
Build variants
| Build | Backends | Size |
|---|---|---|
go build ./cmd/tgexport |
every rclone backend | ~92 MB |
go build -tags slim ./cmd/tgexport |
pikpak only | ~49 MB |
A backend that is not compiled in does not exist at run time, so use the default build unless the destination will never change.
Usage
./tgexport sync -c CHAT -r REMOTE:PATH [options]
A run states what it found, what it is about to do, and then shows each file as it moves:
reading mychannel
12,000 messages read in 4m31s
indexing PikPak root 'mychannel'
11,406 objects listed in 2m10s
chat holds 12,000 media, 250.0 GiB
archived 11,400
never fetched 600
wrong size 6
fetching 606 files, 79.0 GiB (largest 2.0 GiB)
into PikPak root 'mychannel'
staging ./staging, capped at 40.0 GiB
concurrency 2 download(s) x 4 thread(s), 2 upload(s)
↓ total 26/606 files [=> ] 617.5 MiB / 79.0 GiB 2.5 MiB/s 8h47m
↑ total 24/606 files [=> ] 598.0 MiB / 79.0 GiB 2.4 MiB/s 8h58m
↓ …7890_4242_1000000000000000001.mp4 [=======> ] 41.2 MiB / 96.0 MiB 1.8 MiB/s
↑ …7890_4243_1000000000000000002.mp4 ⠹ uploading 1.9 GiB
The two legs are counted separately because they run at different speeds and fail for different reasons. They normally track a file or two apart; a widening gap means the remote is falling behind and staging is filling up.
Redirected output gets the same two figures as plain periodic lines, plus one line per archived file, with no cursor movement — a captured log stays readable:
download 1,200/606 files, 45.0 GiB of 79.0 GiB, 76.8 MiB/s, ETA 7m33s
upload 1,190/606 files, 44.2 GiB of 79.0 GiB, 75.4 MiB/s, ETA 7m53s
CHAT accepts a numeric id as printed by tdl chat ls, a username with or
without @, or a t.me/tg:// link. A Bot API -100… id is converted
automatically. A link to a single message is refused — it names a message, not
a chat.
# archive a chat, capping staging at 40 GiB
./tgexport sync -c @mychannel -r gdrive:telegram/media -m 40G
# check completeness without downloading anything
./tgexport verify -c @mychannel -r gdrive:telegram/media
# list what the chat holds
./tgexport list -c @mychannel
Options
| Flag | Default | Meaning |
|---|---|---|
-c |
— | chat id, username, or link (required) |
-r |
— | rclone destination, REMOTE:PATH (required) |
-d |
./staging |
staging directory for files in flight |
-m |
no cap | cap staging at a size, e.g. 40G |
--threads |
4 | connections per file |
--limit |
2 | files downloading at once |
--uploads |
2 | files uploading at once |
--min-free |
5 | stop if the remote has fewer than this many GiB free, checked before and during the run |
--limit-items |
0 | stop after N files; for smoke tests |
--confirm |
true | re-state each uploaded file to prove its size |
--takeout |
true | use a takeout session |
-n |
default |
tdl session namespace |
Exit codes
| Code | Meaning |
|---|---|
| 0 | complete |
| 1 | ran, but files remain — run again |
| 2 | usage error |
| 3 | remote or Telegram failure, including either leg refusing transfer after transfer |
| 4 | stalled: files remain, none of which can ever be fetched |
| 130 / 143 | interrupted (SIGINT / SIGTERM) |
Only 1 is worth retrying. A driver looping until 0 should stop on anything else: 3 and 4 both mean the next pass would do exactly what this one did.
How it works
Re-running is the resume path. Each item is checked against a listing of the remote immediately before download, so an interrupted run picks up where it left off and a completed one downloads nothing.
Filenames. Every file is stored as {DialogID}_{MessageID}_{FileName}, where
FileName is exactly what Telegram reports. One function derives that string,
and the same string is used both to ask whether the file is already archived and
to write it — so the two can never disagree.
A name that cannot survive that round trip is refused rather than rewritten: too
long for the filesystem once .part is appended, not a single path element, or
containing a character rclone's path encoder rewrites (control bytes, DEL, and
the encoder's own escape character). Those files are reported under
unarchivable and never counted as present. Rewriting them is what the next
paragraph is about.
That last point is the reason this program exists. Its predecessor derived the
name twice: tdl chat export wrote the raw name into a JSON, while tdl dl
rendered it through a template applying filenamify, which rewrites characters a
filesystem rejects and collapses runs of !. A file whose name contained !!
was looked up under one name and stored under another, so the verifier never
found it and re-fetched it on every pass — forever, at 966 MB a time.
Note the consequence: names are not run through filenamify, so they are not
byte-compatible with what the old shell pipeline wrote. A file it stored under a
rewritten name will not be recognised and gets fetched again.
Disk. -m is a byte budget. A download reserves its own size before starting
and releases it only once the upload is confirmed, so when the remote is slow the
downloads pause on their own. The cap must exceed the largest single file, and a
cap that does not is refused at startup rather than discovered as a hang.
Integrity. A download is written to <name>.part and renamed only once its
size matches what Telegram reported, so a file without the suffix is always
whole. Uploads are re-stated afterwards to prove they arrived at the right size,
before the local copy is gone, and an object that turns out short is deleted
rather than left under a name a later run would trust.
verify compares stored sizes against what Telegram reports, so a truncated
object is outstanding rather than "present". This is stricter than the shell
verifier, which matched on name and non-zero size — on the archive this was
built for it found six objects that had been counted complete for months, one
of them 221 MiB standing in for a 2 GiB video. Re-running repairs them.
File references. Telegram hands out a short-lived token with every file
location, and a run over a large chat outlives the ones its walk collected — the
lifetime is undocumented, but an observed run stopped just under two hours in
with FILE_REFERENCE_EXPIRED on every remaining file. So each message is re-read
for a live token immediately before its own download, which costs one round trip
per transfer and needs no guess at how long a token lasts. Retry passes go
through the same path, so a token that dies during a single very large transfer
is replaced rather than replayed.
The re-read has to describe the same file — same name, same length — or it is not
the file this run recorded, and archiving it under that name would store the
wrong bytes. A message that fails that check, or that has been deleted, or that
no longer holds a file at all, is skipped and reported: nothing can be archived
for it, one of them must not strand the rest of the chat, and the closing
verify still lists it as outstanding.
A re-read that fails for any other reason is counted as a failed download, not a skip, because that is what it is — the connection that serves the re-read is the one that serves the file. So it is retried with the rest, it shows up in the progress report and the closing failure list, and enough of them in a row trip the same breaker below.
Failures. A failed download is retried inside the run: up to three passes over whatever is still missing, spaced a minute and then two apart. The failure this exists for is the transient one — a connection that dies takes every transfer in flight with it, and all of them are fetchable again minutes later — and leaving them to the next run costs a full re-walk of the chat and a re-index of the remote before a byte can move. Each failure is reported with the reason Telegram gave for it, not just the byte count that arrived.
Either leg gives up once five transfers in a row fail. A source or destination that is refusing everything will refuse the rest of the list too, in milliseconds, so the run stops and exits 3 instead of spending the whole outstanding list finding that out. A run that trips and then recovers on a retry reports nothing.
Replacing the shell pipeline
Earlier versions of this repo were three bash scripts — run.sh,
export-until-complete.sh and verify-export.sh — driving tdl and rclone as
separate processes. Everything expensive in them existed to work around the fact
that neither process could see the other's state: a staging directory polled with
du -sk, an --min-age guard, a *.tmp exclusion, SIGSTOP/SIGCONT to
enforce the disk cap, a sweep-failure counter, and an outer loop that re-verified
and re-narrowed a JSON export between passes.
One process needs none of it. Completion is a function returning; the cap is a semaphore. Some hard-won details were worth keeping, and are:
- pikpak commits uploads as a server-side async task, and rclone abandons a
still-pending one when its low-level retries run out.
transfers=2andlow-level-retries=20are the defaults here for that reason. Environment overrides still win. - A backend with no quota API is treated as unlimited, so it never blocks a run.
- Zero-byte files count as missing — rclone overwrites a size-mismatched destination, so re-running repairs them — while files under 1 KiB are reported but trusted, since some real media genuinely is that small.
- Indexing ignores
RCLONE_*filters. The transfer tunables above are deliberately env-overridable; the listing is not. A strayRCLONE_EXCLUDEorRCLONE_MIN_SIZEleft over from another job would otherwise narrow the index and re-download everything it hid.
Notes
tgexportand thetdlCLI share one session store and cannot run against the same namespace at once. Use-nfor a second namespace if you need both.- A partially downloaded file is not resumable across restarts — tdl's library
exposes no resume offset — so an interrupted run re-fetches whatever was in
flight, bounded by
--limit. - Everything is read-only against Telegram. Nothing is uploaded, deleted, or marked read.