LogBus2 guide
1. About LogBus2
LogBus2 is a log synchronization tool redeveloped on the basis of the original LogBus. Compared with the original LogBus, it uses one-fifth of the memory and is 5 times faster.
LogBus2 imports backend log data into the AE backend in real time. Its core working principle is similar to Flume and Loggie: it monitors the file streams in the log directories on the server. When new data is written to any log file in a directory, LogBus2 validates the new data and sends it to the AE backend in real time.
We recommend LogBus2 for the following types of users:
- Users who store data in AE format with server-side SDKs, Kafka, or SLS, and upload it through LogBus2
- Users with high requirements for data accuracy and dimensions, whose data needs cannot be met by client SDKs alone, or for whom integrating client SDKs is inconvenient
- Users who do not want to develop their own backend data push process
- Users who need to transfer large volumes of historical data
- Users with specific requirements for memory usage and transfer efficiency
This document applies to Logbus v2. For Logbus v1, see the LogBus guide and the LogBus Windows version guide
To migrate from Logbus v1 to Logbus v2, contact ThinkingAI technical support.
2. Download LogBus2
Latest version: 2.1.2.2
Update time: 2025-11-07
Download link for Mac Apple Silicon
3. Preparation
File type
- Determine the directory where the data files to upload are stored, and configure LogBus2 accordingly. LogBus2 monitors file changes in the directory (new files are created or existing files are tailed).
- Do not directly rename data logs that are stored in the monitored directory and have already been uploaded. Renaming a log is equivalent to creating a new file, and LogBus2 may upload these files again, which causes duplicate data.
- The LogBus2 running directory contains snapshots of the current log transfer progress. Do not operate on the files in the runtime directory yourself
Note: Do not store data files in object storage or in mount directories of cloud file systems such as NFS and SMB, to avoid read errors.
Kafka
- Determine the Kafka message format. LogBus only processes the value of Kafka messages
- Make sure that data with the same user ID is written to the same partition to avoid out-of-order data
- Enable free use of Kafka consumer groups to prevent failures when multiple LogBus instances consume data
- Consumption starts from earliest by default. To consume from a specific offset, first create the consumer group and the specific offset
SLS
- Contact Alibaba Cloud to enable consumption over the Kafka protocol
CLS
- Make sure that automatic splitting is disabled for CLS.
- Apply for the AK&SK of the corresponding resources
4. Install and upgrade LogBus2
Install
Download the LogBus2 installation package and unzip it.
Directory structure after unzipping:
- Logbus: LogBus2: Binary file
- conf:
- daemon.json: Configuration file template 2
- tools:
- configConvert: Configuration conversion tool
Upgrade
Requirement: LogBus2 version ≥ 2.0.1.7
Run
./logbus update directly. After the upgrade is complete, run
./logbus start
5. Use and configure LogBus2
Startup parameters
Start
./logbus start
Stop
./logbus stop
Restart
./logbus restart
Check the configuration and connectivity with the AE system
./logbus env
Reset LogBus read records
./logbus reset
# Currently unavailable for Kafka
View transfer progress
./logbus progress
Validate the file format
./logbus dev
# Currently unavailable for Kafka
Configuration file guide
Default configuration template
{
"datasource": [
{
"file_patterns": [
"/data/log1/*.txt", // #Change this to the absolute path of your source data files
"/data/log2/*.log"
],//File patterns
"app_id": "app_id"//app_id is the token from the AE official website. Get the APPID of the project from the project configuration page in the AE backend and enter it here
}
],
"push_url": "http://RECEIVER_URL"//For HTTP transfer, use http://receiver.ta.thinkingdata.cn/. If you use an on-premises deployment, change the transfer URL to: http://YOUR_RECEIVER_URL/
}
Common configuration
File
{
"datasource": [
{
"type":"file",
"file_patterns": ["/data/log1/*.txt", "/data/log2/*.log"], //File Glob matching rules
"app_id": "app_id", //APPID is the token from the AE official website. Get the APPID of the project from the project configuration page in the AE backend and enter it here
"unit_remove": "day", //File deletion unit. Supports "day", "hour"
"offset_remove": 7,//unit_remove*offset_remove gives the final removal time **offset must be greater than 0; otherwise, it does not take effect
"remove_dirs": true,//Whether to enable folder deletion. Default: false NOTE: A folder is deleted only after all files in it have been consumed
"http_compress": "gzip" //Whether to enable HTTP compression. Default: "none"
}
],
"cpu_limit": 4, //Limit the number of CPU cores that LogBus2 uses
"push_url": "http://RECEIVER_URL"
}
Kafka
{
"datasource": [
{
"type":"kafka", //The type is Kafka
"topic":"ta", //The topic to consume
"brokers":[
"localhost:9091" //Kafka broker addresses
],
"consumer_group":"logbus", //Consumer group name
"cloud_provider":"ali", //Cloud provider name. Supports "ali", "tencent", "huawei"
"username":"", //Kafka username
"password":"", //Kafka authentication password
"instance":"", //Cloud provider instance name
"protocol":"none", //Authentication protocol. Supports "none"|"plain"|"scramsha256"|"scramsha512"
"block_partitions_revoked":true,
"app_id":"YOUR_APP_ID"
}
],
"cpu_limit": 4, //Limit the number of CPU cores that LogBus2 uses
"push_url": "http://RECEIVER_URL"
}
SLS
NOTE: Before you consume with SLS, contact Alibaba Cloud to enable the SLS Kafka consumption protocol
{
"datasource": [
{
"type":"kafka",
"brokers":["{PROJECT}.{ENTRYPOINT}:{PORT}"], //NOTE: For details, see https://help.aliyun.com/document_detail/29008.htm#reference-wgx-pwq-zdb
"topic":"{SLS_Logstore_NAME}", //Logstore name
"protocol":"plain",
"consumer_group":"{YOUR_CONSUMER_GROUP}", // ConsumerGroup
"username":"{PROJECT}", // Project name
"disable_tls":true,
"password":"{ACCESS_ID}#{ACCESS_PASSWORD}", // Alibaba Cloud RAM authorization
"app_id":"YOUR_APP_ID"
}
],
"push_url": "http://RECEIVER_URL"
}
CLS
NOTE: Before use, make sure that the consumption/write throughput is greater than the log retention time
{
"datasource":[
{
"type": "kafka",
"brokers":["YOUR_AZ_ENDPOINT"],
"session_timeout": 9000,
"fetch_max_bytes": 104857600,
"topic": "YOUR_TOPIC",
"protocol": "plain",
"consumer_group": "YOUR_GROUP",
"username": "",
"password": "",
"block_paritions_revoked":"true",
"app_id":"YOUR_APP_ID"
}
],
"push_url":"http://RECEIVER_URL"
}
Full configuration items
Configuration items and descriptions
| Setting | Type | Example | Required | Description |
|---|---|---|---|---|
| cpu_limit | Number | 4 | Maximum number of CPU cores that LogBus2 can use | |
| push_url | String | ✔️ | Receiver address. Must start with http/https. | |
| datasource | Object list | ✔️ | Data source list | |
min_disk_free_space | uint64 | 1024 | Checks the free space in the directory where logbus2 is located, and shuts down automatically when it is less than the configured value. Unit: KB. Default: 1*1024*1024 |
datasource (data source configuration)
File
| Setting | Type | Example | Required | Default value | Description |
| app_id | String | ✔️ | "" | appid of the project that data is reported to | |
| appid_in_data | Bool | false | false | Enable this option to distribute data by the appid in the files. LogBus2 then no longer uses the appid in app_id for distribution | |
specified_push_url | Bool | false | true: push_url is not parsed, and data is sent to the push_url exactly as you configured it, that is, http://yourhost:yourport. false: push_url is parsed, and data is sent to the logbus url specified by the receiver, that is, http://yourhost:yourport/logbus. | ||
| add_uuid | Bool | false | true: Whether to add a uuid property to each record (enabling it reduces transfer efficiency). | ||
| file_patterns | String list | ✔️ | [""] | Supports directory wildcards; regex is not supported yet. Unless specially configured, files with the .gz/.iso/.rpm/.zip/.bz/.rar/.bz2 suffixes are skipped by default | |
| ignore_files | String list | [""] | Files to filter out of file_patterns | ||
| unit_remove | String | "" | User file deletion. Delete by day (day) or hour (hour). Note: If automatic file deletion is not configured, LogBus2 memory usage gradually increases | ||
| offset_remove | Int | 0 | User file deletion. File deletion takes effect when offset_remove>0 and unit_remove is set to day or hour. | ||
| remove_dirs | Bool | true|false | false | Whether to enable folder deletion | |
| http_timeout | String | 500ms | 600s | Timeout for sending data to the receiver. Default: 600s. Range: 200ms - 600s. Supports milliseconds "ms", seconds "s", minutes "m", and hours "h". | |
| iops | int | 20000 | 20000 | Rate limit on LogBus data throughput per second (number of records) | |
| limit | bool | true|false | false | Enables rate limiting | |
| http_compress | String | none | gzip | none | Compression format for data sent over HTTP. none = no compression. Default: none. | |
| filters | object list | Event filters. Multiple filters are combined with OR | |||
| filters[0].key | string | #event_name | ✔️ | Key to filter on | |
| filters[0].value | interface{} | register | ✔️ | Value to filter on | |
| filter[0].type | string | string | string | Type of the value. Default: string. Supports string | boolean | int64 | regex |
Kafka
NOTE: Before you use LogBus in Kafka mode, make sure to enable free use of consumer groups
| Setting | Type | Example | Required | Default value | Description |
|---|---|---|---|---|---|
| brokers | String List | ["localhost:9092"] | ✔️ | [""] | Kafka Brokers |
| topic | String | "ta-msg-chan" | ✔️ | "" | Kafka topic to consume |
| consumer_group | String | "ta-consumer" | ✔️ | "" | Kafka Consumer Group |
| protocol | String | "plain" | "none" | Kafka authentication mode | |
| username | String | "ta-user" | "" | Kafka username | |
| password | String | "ta-password" | "" | Kafka password | |
| instance | String | "" | "" | Instance ID required by CKafka | |
| fetch_count | Number | 1000 | 10000 | Number of messages per poll | |
| fetch_time_out | Number | 30 | 5 | Poll timeout | |
| read_committed | Bool | true | false | Whether to consume only Kafka Committed data | |
| disable_tls | Bool | true | false | Disables TLS verification | |
| cloud_provider | String | "tencent" | "" | Enable when accessing Kafka over the public network. Currently supported cloud providers: tencent, huawei, ali | |
| block_partitions_revoked | Bool | false | false | Whether to block consumption. If it is not enabled, duplicate data occurs when multiple LogBus instances are in the same consumer_group | |
| auto_reset_offset | String | "earliest" | "earliest" | Default behavior when no offset has been committed |
NOTE: LogBus v2 currently consumes Kafka in load-balancing mode. The number of LogBus v2 deployments must be ≤ the number of partitions
Monitoring configuration and dashboard setup
See Monitoring configuration demo
Alert configuration
Plugins
6. Advanced usage
Report multiple events with a single LogBus
When a single LogBus is deployed, some data may be consumed with a delay because of I/O limits. For example
.
├── event_1
│ ├── log.1
│ ├── log.2
│ ├── log.3
│ ├── log.4
│ └── log.5
├── event_2
│ ├── log.1
│ ├── log.2
│ ├── log.3
│ ├── log.4
│ └── log.5
└── event_3
├── log.1
├── log.2
├── log.3
├── log.4
└── log.5
Because of polling, the consumption order is event_*/log.1 -> event_*/log.2 -> event_*/log.3. In this case, file consumption progresses slowly. You can start multiple LogBus instances and split logs that have no contextual semantics with Glob patterns, so that the files matched by each Glob are uploaded in parallel
Multi-pipeline configuration
NOTE: appids cannot be repeated across multiple pipelines
{
"datasource": [
{
"file_patterns": ["/data/log1/*.txt", "/data/log2/*.log"], //File Glob matching rules
"app_id": "app_id", //APPID is the token from the AE official website. Get the APPID of the project from the project configuration page in the AE backend and enter it here
"unit_remove": "day", //File deletion unit. Supports "day", "hour"
"offset_remove": 7,//unit_remove*offset_remove gives the final removal time **offset must be greater than 0; otherwise, it does not take effect
"remove_dirs": true,//Whether to enable folder deletion. Default: false NOTE: A folder is deleted only after all files in it have been consumed
"http_compress": "gzip"//Whether to enable HTTP compression. Default: none
},
{
"file_patterns": ["/data/log1/*.txt", "/data/log2/*.log"], //File Glob matching rules
"app_id": "app_id", //APPID is the token from the AE official website. Get the APPID of the project from the project configuration page in the AE backend and enter it here
"unit_remove": "day", //File deletion unit. Supports "day", "hour"
"offset_remove": 7,//unit_remove*offset_remove gives the final removal time **offset must be greater than 0; otherwise, it does not take effect
"remove_dirs": true,//Whether to enable folder deletion. Default: false NOTE: A folder is deleted only after all files in it have been consumed
"http_compress": "gzip"//Whether to enable HTTP compression. Default: none
}
],
"cpu_limit": 4, //Limit the number of CPU cores that LogBus2 uses
"push_url": "http://RECEIVER_URL"
}
LogBus2 On Docker
Pull the latest image
docker pull thinkingdata/ta-logbus-v2:latest
Create persistent folders on the host and initialize the configuration file
mkdir -p /your/folder/path/{conf,log,runtime}
touch /your/folder/path/conf/daemon.json
vim /your/folder/path/conf/daemon.json
⚠️Warning: Do not delete any files in the runtime directory yourself
Modify the configuration template and write it to daemon.json
{
"datasource": [
{
"type":"file",
"app_id": "YOUR APP ID",
"file_patterns": ["/test-data/*.json"]
},
{
"type":"kafka",
"app_id": "YOUR APP ID",
"brokers": ["localhot:9092"],
"topic":"ta-message",
"consumer_group":"ta"
}
],
"push_url": "YOUR PUSH URL WITHOUT SUFFIX OF/logbus"
}
Mount the data folder and start LogBus
docker run -d \
--name logbus-v2 \
--restart=always \
-v /your/data/folder:/test-data/ \
-v /your/folder/path/conf/:/ta/logbus/conf/ \
-v /your/folder/path/log/:/ta/logbus/log/ \
-v /your/folder/path/runtime/:/ta/logbus/runtime/ \
thinkingdata/ta-logbus-v2:latest
LogBus2 On K8s
Environment preparation
- kubectl can connect to the k8s cluster and has deployment permissions.
- Install dependencies: Install helm to your local command line by following the helm documentation https://helm.sh/zh/docs/intro/install/
Download the logbus v2 helm files
📎 ta-logBusv2-2.0.1.8-helm.tar.gz(2 KB)tar xvf logBusv2-helm.tar && cd logbusv2
Configure logbus
Preparation
- Create the PVC for the logs to upload in the console
- Get the PVC name and confirm the namespace
- Get the AE app id and receiver url
Modify values.yaml
pvc:
name: YOUR_PVC_NAME # PVC name
logbus_version: 2.1.0.2
namespace: YOUR_NAMESPACE # namespace name
logbus_configs:
- push_url: "http://RECEIVER_URL" # receiver address for uploading data to AE
datasource:
- file_patterns:
- "container:wildcard for the relative path of the files" # Note: Do not delete the "container:" prefix
- "container:wildcard for the relative path of the files" # Note: Do not delete the "container:" prefix
app_id: YOUR_APP_ID # app id of the AE system
Preview the rendered yaml
helm install --dry-run -f values.yaml logbus .
Deploy logbusv2 with helm
helm install -f values.yaml logbus-v2 .
Check the created statefulset
kubectl get statefulset
Check the created pod
kubectl get pods
Update the LogBus version in K8s
vim values.yaml # Modify the existing values.yaml file
# Change logbus_version to the latest version NOTE: For backward compatibility, it is best not to use latest!
logbus_version:2.0.1.8 -> logbus_version:2.1.0.2
# Save and exit
helm upgrade -f values.yaml logbus .
# Wait for the rolling update
Note
logbusv2 needs read and write permissions on the PVC of the mounted logs.
logbusv2 writes file consumption records and runtime logs to the PVC separately for each pod. If the PVC is deleted, the LogBus records are lost, and data may be re-transferred.
Configuration details
Run the command:
helm show values .
The available configuration is displayed:
# Default values for logbusv2.
# This is a YAML-formatted file.
# Declare variables to be passed into your templates.
pvc:
name: pvc-logbus
logbus_version: 2.1.0.2
namespace: big-data
logbus_configs:
#### pod 1
#### push_url: receiver url, need http:// https:// prefix
- push_url: "http://192.0.2.10:8992/"
datasource:
- file_patterns:
#### target files relative path in pvc
- "container:/ta-logbus-0/data_path/*"
#### AE app_id
app_id: "thinkingAnalyticsAppID"
#### pod 2
- push_url: "http://192.0.2.11:8992/"
datasource:
- file_patterns:
- "container:/ta-logbus-1/data_path/*"
app_id: "thinkingAnalyticsAppID"
#### pod 3
- push_url: "http://192.0.2.11:8992/"
datasource:
- file_patterns:
- "container:/ta-logbus-2/data_path/*"
app_id: "thinkingAnalyticsAppID"
#### logbus pod requests
#requests:
# cpu: 2
# memory: 1Gi
If the requests section is not explicitly configured, it does not appear in the yaml.
Single directory in a PVC
pvc:
name: YOUR_PVC_NAME # Enter the actual PVC name
namespace: YOUR_NAMESPACE # An existing namespace
logbus_configs:
- push_url: "http://RECEIVER_URL" # http or https; enter an AE receiver URL that the pod can access
datasource:
- file_patterns:
- "container:/ta-logbus-0/data_path/*" # "container:" is a placeholder. During YAML deployment, the relative path is replaced with an absolute path that the container can access. When you configure a directory, add the container: prefix to it.
app_id: "thinkingAnalyticsAppID" # app id of the AE system
Multiple directories in a PVC
When you read multiple directories in a PVC, we recommend deploying separate pods, with each pod responsible for one folder. This gives better performance and security.
pvc:
name: pvc-logbus
namespace: big-data
logbus_configs:
#### pod 1
#### push_url: receiver url, need http:// https:// prefix
- push_url: "http://192.0.2.10:8992/"
datasource:
- file_patterns:
#### target files relative path in pvc
- "container:/ta-logbus-0/data_path/*"
#### AE app_id
app_id: "thinkingAnalyticsAppID"
#### pod 2
- push_url: "http://192.0.2.11:8992/" # Note: Configure each app id and push url separately
datasource:
- file_patterns:
- "container:/ta-logbus-1/data_path/*"
app_id: "thinkingAnalyticsAppID"
#### pod 3
- push_url: "http://192.0.2.11:8992/"
datasource:
- file_patterns:
- "container:/ta-logbus-2/data_path/*"
app_id: "thinkingAnalyticsAppID"
Multiple PVCs
Currently, only single-PVC deployment is supported. For multiple PVCs, configure the values.yaml file multiple times
7. FAQ
Q: Why doesn't LogBus delete the folder even though folder deletion is enabled?
A: LogBus deletes a folder only when the files in the folder have been read by LogBus and no files remain in the folder
Q: Why can't logs be uploaded?
A: A single record in the data files that LogBus reads must not contain line breaks. The configured data files do not support regex; only wildcards (Glob) can be used. Check whether the configured data file rules match any files
Q: Why are files uploaded repeatedly?
A: LogBus stores its file reading progress in the runtime directory. If this directory is deleted by mistake, files that have already been uploaded are uploaded again.
Q: Why is the data skewed?
A: AE currently uses the customer's distinct_id as the data UUID for shuffling. When a massive amount of data uses the same string as distinct_id, the memory pressure on a single machine may increase, which increases the risk of data skew
8. Release notes
Version: 2.1.2.2 --- 2025.11.7
New
- Added an option to skip parser data integrity validation
Version: 2.1.2.1 --- 2025.9.26
New
- Added a Kafka connection timeout setting. Default: 15000ms
Optimized
- Reduced how often Kafka consumption information monitoring metrics are fetched
Version: 2.1.2.0 --- 2025.5.8
Optimized
- Updated the franz-go version
Version: 2.1.1.9 --- 2025.3.24
New
- Added whole-line filtering
Version: 2.1.1.8 --- 2025.2.23
Optimized
- Optimized the handling logic for waiting after reading empty lines
Version: 2.1.1.7 --- 2025.1.20
New
- Added regex to the filter types, allowing filtering with regular expressions
Version: 2.1.1.6 --- 2024.12.20
Optimized
- Supported HTTP proxies configured in env
Version: 2.1.1.4 --- 2024.5.15
Optimized
- Kafka sources support fetch_max_partition_bytes
- DataSource supports skipping remote validation (disabled by default)
- Used SIMD to accelerate JSON processing
Version: 2.1.1.3 --- 2024.2.23
Optimized
- Improved the Kafka progress experience
- Supported the auto_reset_offset parameter
- Logged Kafka client information
Fixed
- LogBus killed the wrong process after a forced shutdown
- Memory leaks
- Goroutine leaks
- Kafka stopped consuming after timeouts
- Windows version failed to start
- Windows version could not be updated
Version: 2.1.1.2 --- 2024.2.2
Optimized
- Log format and partially mixed i18n issues
Version: 2.1.1.1 --- 2024.1.10
Optimized
- Accuracy of the
progresscommand
Version: 2.1.1.0 --- 2023.12.11
Fixed
- Corrected the log defaults: logs are kept for 7 days, split at 100M per file, with up to 30 log files retained. If you don't upgrade, you can also set the log configuration item.
Impact
- When the data source produces a large amount of unparsable erroneous data, the log files keep growing because the logs record the erroneous data.
Version: 2.1.0.9 --- 2023.10.26
New
- Supported event filters to filter data on the client
- Supported space detection for the directory where logbus is located
Version: 2.1.0.8 --- 2023.6.06
Optimized
- Ensured that plugins are shut down correctly
- Optimized process communication and log output
Fixed
- cpu limit log output
- Linux ARM architecture build process
Version: 2.1.0.7 --- 2023.4.07
New
- Supported the progress command for Kafka data sources
Optimized
- Custom labels can read environment variables
Version: 2.1.0.6 --- 2023.3.28
Optimized
- Supported custom labels to trace data sources
- Supported custom plugin separators
Version: 2.1.0.5 --- 2023.2.20
Optimized
- Data routing projects support arrays and are compatible with the number and string basic types
- Monitoring metric calculation logic
Version: 2.1.0.4 --- 2023.1.12
New
- Supported configuring data sources that read files without line breaks
- Supported configuring the number of read loops and the interval
Fixed
- Fixed a concurrency bug in monitoring metric statistics
Version: 2.1.0.3 --- 2022.12.23
New
- Allowed the configuration to override the appid inside the data
Fixed
- When appid_in_data is used, the default value of appid is removed
Version: 2.1.0.2 --- 2022.12.13
New
- Plugins support property splitting
Fixed
- meta_name creation with multiple pipelines
Version: 2.1.0.1 --- 2022.11.29
New
- Kafka data sources support transactional read committed
- Plugin commands support sh environment dependencies
Fixed
- meta_name creation when no appid is configured
Version: 2.1.0.0 --- 2022.11.22
New
- Data routing: routes data to different projects based on the configured appidMap
- Kafka data sources support consuming multiple topics
- Rate limiter that limits the reporting speed to reduce server-side pressure
- Data reporting with multiple pipelines
- Custom plugin parser based on gRPC
- Real-time performance monitoring (prometheus, pushgateway, grafana)
Fixed
- File consumption blocking issue
- Fixed a bug where logbus could not be stopped with Kafka data sources
- Fixed a bug where highly active files could not be exited with file data sources
- Fixed file watcher shutdown issues
Version: 2.0.1.8 --- 2022.07.20
New
- dev (format validation command)
- Kafka Source
- Multi-platform support
Fixed
- Issue where the file transfer progress process was woken up multiple times
- Reduced log volume
Progresssorts by file transfer time- Multi-pipeline optimization
- Slimmed down the
Docker image
Version: 2.0.1.7 --- 2022.03.01
Optimized
- Running efficiency and performance
- File deletion logic
- Offset file export logic
- Memory usage

