Schedule and backup tuning¶
The backup service runs on a cron schedule and backs up every account in turn.
Two ways to run the backup¶
| Command | Runs | Use it for |
|---|---|---|
schedule |
Until stopped | Normal operation. The stock compose file uses it. |
backup |
One run over every account, then exits | A manual run, or an external scheduler. |
The schedule command¶
The schedule command opens one shared Telegram connection per account and keeps it open. It starts the real-time listener when ENABLE_LISTENER=true. It runs one backup straight away, then waits for the next SCHEDULE tick. Every 30 seconds it writes the heartbeat file the Docker health check reads. If a listener dies, the scheduler restarts it after 5 seconds. SIGTERM and SIGINT cancel the running backup, stop the listeners and close the Telegram connections.
The backup command¶
The backup command runs one backup over every account and exits. It fails only when every account failed. It runs no gap-fill, starts no listener and writes no heartbeat.
Stop the scheduler before a one-shot backup
Stop the schedule process first and start it again afterwards. See One client per session.
A container that runs backup instead of schedule writes no heartbeat. Docker checks the heartbeat every 60 seconds once the 300-second start period ends. After three failed checks, about three minutes, it marks the container unhealthy. See Monitoring and troubleshooting.
When backups run¶
SCHEDULE is a cron expression with five fields: minute, hour, day, month and day of week. The default is 0 */6 * * *, which runs at minute 0 of every sixth hour. Any other number of fields stops the scheduler at startup.
SCHEDULE |
Runs |
|---|---|
0 */6 * * * |
Every 6 hours, on the hour |
0 * * * * |
Every hour |
30 3 * * * |
Every day at 03:30 |
0 2 * * sun |
Every Sunday at 02:00 |
Use day names
In the day-of-week field, 0 means Monday, not Sunday as in crontab. Names such as mon, sun or mon-fri avoid the confusion.
The schedule uses the local time of the backup process. The images run in UTC unless you set TZ on the backup service. VIEWER_TIMEZONE does not change it. The stock compose file passes the whole .env to the backup service, so one line there is enough:
A run that starts late, for example after the machine was asleep, still runs if it is less than 3600 seconds late. Several missed runs collapse into one. When a tick arrives while a backup is still running, the scheduler skips that tick and logs a warning.
Accounts are backed up one after another, in configuration order. With several accounts, one run takes as long as all the accounts' runs added together. See Multiple accounts.
What one run does¶
For each account, in order:
- Log in and store the account owner's id.
- Write the last backup time and set the "backup in progress" flag.
- Delete previously downloaded YouTube preview videos and their rows, when
YOUTUBE_VIDEOS_DELETE_EXISTINGis on andDOWNLOAD_YOUTUBE_VIDEOSis off. See Media downloads. - Correct configured chat ids that lack the
-100prefix, and refresh folder membership for folder-based filters. - Select the chats to back up, priority chats first. Choosing chats explains the rules.
- For each chat:
- Update the chat row. With
DOWNLOAD_CHAT_DESCRIPTION=true, this includes the description and member count, at the cost of one extra request per chat. If that request hits a FloodWait, descriptions are skipped for the rest of the run and fetched again on the next one; the chat keeps the values it already has. The member count is filled in for channels and supergroups. - Check the chat's avatar and download it when it changed.
- Fetch new messages, oldest first.
- Sync edits and deletions, if
SYNC_DELETIONS_EDITSis on. - Refresh pinned messages, up to 100 per chat.
- Re-check recent reactions, if the reaction re-sweep is on.
- Update the chat row. With
- Back up archived chats that the main loop did not already cover.
- Reconcile group-to-supergroup migrations.
- Refresh forum topics and folders.
- Recalculate statistics.
- Retry failed media downloads. See Media downloads.
- Verify media files, if
VERIFY_MEDIAis on. - Send pending media for transcription.
The run clears the "backup in progress" flag at the end, even when it fails.
When a chat is private or forbidden, or you are banned from it, the run logs it as skipped. Any other error in one chat is logged, and the run moves on to the next chat.
The last backup time is taken at the start of the run. The archive stores one value for all accounts, so with several accounts the last account to start wins.
Where a run picks up¶
Each chat has a stored position per account: the id of the newest message already archived. A run fetches only messages with a higher id. A chat with no position starts from its first message, so the first run fetches the whole history. The position never moves back.
Messages are written in batches of BATCH_SIZE messages, 100 by default. The position is saved every CHECKPOINT_INTERVAL batches. The default is 1, which is also the minimum. It is saved once more at the end of each chat. After a crash, only the messages since the last saved position are fetched again.
A message that fails to process holds the position just before it. The rest of the chat is still stored, and the failed message is tried again on the next run. Once the same message has failed on 2 separate runs, the position moves past it. The backup records its id in the metadata.
Older messages are never read again by a normal run. Edits and deletions of messages below the position reach the archive only through SYNC_DELETIONS_EDITS or the real-time listener. Two things are refreshed on every run anyway: pinned flags, and reactions on every message fetched in that run.
Catching changes to older messages¶
| Option | What it catches | Cost |
|---|---|---|
| Real-time listener | Edits, deletions, new messages and reactions as they happen | A process that stays connected. |
SYNC_DELETIONS_EDITS |
Edits and deletions of every archived message | Re-reads the whole archive on every run. |
REACTION_RESWEEP_DAYS |
Reaction changes on recent messages | Up to 500 messages per chat per run, spaced out. |
FILL_GAPS |
Messages missing between two archived messages | One scan of the stored ids per chat after each backup run. |
SYNC_DELETIONS_EDITS¶
With SYNC_DELETIONS_EDITS=true, every run re-reads every archived message of every chat, 100 ids per request. The sync applies an edit when Telegram reports a different edit date. It also reconciles reactions.
A message counts as deleted only when Telegram's answer matches the requested ids one for one. If the answer does not match, no message from that request is marked deleted in this run. A deletion follows DELETION_MODE. The default, soft, marks the message deleted and keeps it. hard removes the row.
This sync fires no event webhook.
No mass-operation limit
The mass-operation protection only covers the real-time listener. This sync has no such limit. With DELETION_MODE=hard, it removes every message that Telegram confirms as gone, however many there are. Keep DELETION_MODE=soft unless you want deleted messages gone from the archive too.
Reaction re-sweep¶
REACTION_RESWEEP_DAYS re-checks reactions on messages from the last N days, in every chat, on every run. It is off by default and works without the listener.
| Variable | Default | Effect |
|---|---|---|
REACTION_RESWEEP_DAYS |
0 |
Days to look back. 0 turns the re-sweep off. |
REACTION_RESWEEP_MAX_PER_CHAT |
500 |
Most messages re-checked per chat per run. Requests hold 100 ids. |
REACTION_RESWEEP_BATCH_DELAY_SECONDS |
2.0 |
Minimum gap between re-sweep requests, across all chats. |
After a FloodWait, the re-sweep waits the required time plus 2 seconds, then continues in the same run. After 3 FloodWaits in one run, it stops until the next run. Progress is saved, so the next run continues where this one stopped. Saved progress older than 48 hours, or from a different day window, is discarded.
Gap-fill¶
With FILL_GAPS=true, the scheduler looks for holes after the startup backup and after each scheduled backup. It never runs after a one-shot backup.
A gap is two consecutive stored message ids further apart than GAP_THRESHOLD, 50 by default. The missing range is fetched and stored. Gap-fill never moves the chat's position. When it recovers messages, it recalculates statistics.
A hole before a chat's earliest archived message is reported in the log but never fetched. To run gap-fill by hand, for one chat or with another threshold, use the fill-gaps command described in Import and maintenance tasks.
Rate limits and retries¶
Telegram answers too many requests with a FloodWait: an order to wait a number of seconds. The client never sleeps through a FloodWait on its own. There are two exceptions, and both absorb waits of up to 60 seconds by default:
| Variable | Default | Applies to |
|---|---|---|
MEDIA_FLOOD_SLEEP_THRESHOLD |
60 |
Media transfers. The transfer resumes at its current offset. |
DIALOG_FLOOD_SLEEP_THRESHOLD |
60 |
Listing your chats. The listing resumes on the same page. |
For how this interacts with the download timeout, see Media downloads.
Message history, chat lookups, the deletion and edit sync, pinned messages and forum topic pages go through a bounded retry:
| Variable | Default | Effect |
|---|---|---|
MAX_FLOOD_RETRIES |
5 |
Retries for one call. |
MAX_FLOOD_WAIT_SECONDS |
3600 |
A longer FloodWait fails at once. The chat or file is tried again on the next run. |
BACKOFF_MIN_SECONDS |
2.0 |
Base of the exponential backoff. |
BACKOFF_MAX_SECONDS |
300.0 |
Ceiling of the backoff. |
FLOOD_WAIT_LOG_THRESHOLD |
10 |
While fetching message history, a first FloodWait shorter than this is not logged. 0 logs them all. |
The folder fetch near the end of the run, the avatar download and the chat description fetch are not retried; a failure there is logged and the step waits for the next run.
The sleep before retry number n is:
Transient errors, such as timeouts and dropped connections, use the same exponential backoff with 0.5 to 1.5 s of jitter, and reconnect the client when it is down. These errors are never retried:
- an expired file reference
- a private or forbidden chat
- an invalid chat or peer id
- an unauthorized session
- an auth key error
- a ban from the channel
While fetching a chat's history, a FloodWait or a dropped connection resumes from the last message it already received instead of starting the chat over. The retry counter resets each time a message arrives, so a long history is not cut short by scattered waits.
An invalid value in one of these retry variables logs a warning and falls back to its default. An invalid value in most other settings stops startup with the variable's name.
Statistics¶
Statistics are recalculated at the end of every backup run, and after a gap-fill that recovered messages. STATS_CALCULATION_HOUR does not affect the backup. It sets the hour of the viewer's own daily recalculation. See Using the viewer.
Forum topics and folders¶
Forum topics are fetched early in each forum chat's backup and again at the end of the run. Telegram returns them 100 per page, and the backup reads up to 50 pages. Deleted topics and topics listed in SKIP_TOPIC_IDS are not stored. Custom emoji icons are resolved to their emoji. If the topic request fails before any page arrives, topics are inferred from the topic ids on the stored messages.
Near the end of each run, your Telegram folders are stored so the viewer can show them as tabs. Membership is resolved against the archived chats and includes category flags such as "all groups". Mute and read state are not archived, so the "exclude muted" and "exclude read" flags are not applied. Folders you deleted in Telegram are removed from the archive.