Telegram Archiver is a tool for live archival of telegram messages in channels, chats, or dialogs. It saves incoming messages (filtered by config) into human-readable JSON files. Those JSON files are similar in format to one used by Telegram Desktop export, and, while designed to be familiar, those formats are not designed to be compatible.
This tool is designed to survive any disruption because it always downloads older messages first and does that sequentally, allowing it to catch up in case of restarts of any nature. Due to that, its performance is limited.
Due to Telethon limitations, this tool crashes on any error. However, this is done for good: there's no retry mechanisms available nor there's any possibility to stop updates from coming, so this tool relies on external restart for a recovery catch-up. If ntfy is enabled in config, it will alert before crashing.
Use pipenv install and pipenv shell (having pipenv installed) to download dependencies.
Then use python -m tg_archiver --help. If you place config.yml file in current working directory, you can use
python -m tg_archiver -i right away. -i flag is required for authentication only, and can be removed after that.
Configuration example is below. Commented values are defaults. Also see config.py file.
# Read https://core.telegram.org/api/obtaining_api_id to understand how to get those values
api_id: <API_ID>
api_hash: <API_HASH>
# Path to telethon session file
#session_path: "session.session"
# Path to message store directory
#data_dir: "data"
# Includes all peers if included here. This example only fetches new messages from subscribed channels.
include_all:
# all three share the same options
#dialogs:
# false - don't prefetch all preceding messages
# exclude_archived - prefetch all preceding messages in each non-archived peer
# true - prefetch all preceding messages in all peers
#fetch_old: false
# none - download all forwarded media
# all - avoid downloading forwarded media
# from_channels - avoid downloading forwarded media from channels
#exclude_forwarded_media: none
# If document file is bigger than this, it is not downloaded
#max_document_size_bytes: 9223372036854775807
#groups: {}
# you can use {} for defaults
channels: {}
include_peers:
# either id or username. Including both is not supported
- id: 0
username: "username"
# Same as in include_all but there's no "exclude_archived" option
#fetch_old: false
# Same as in include_all
#exclude_forwarded_media: "none"
# noinspection YAMLIncompatibleTypes
# This block supports both IDs and usernames, triggering YAMLIncompatibleTypes inspection in IDEA :-)
# Use it to exclude peers entirely. Such peers are ignored by Telegram Archiver.
exclude_peers:
- 0
- "username"
# If ntfy block is included, Telegram Archiver will alert about failures and notify about startups
#ntfy:
# endpoint: "https://some.ntfy.endpoint"
# token: "somebearertoken"