Ta-Datax-Writer plugin user guide
1. Introduction
Ta-Datax-Writer is a DataX writer plugin that lets you transfer data to the AE cluster within the DataX ecosystem. You can deploy DataX on a data transfer server and use any reader plugin supported by DataX together with this plugin to sync data from multiple data sources to the AE cluster.
To learn more about DataX, visit the DataX GitHub page
Data is transferred by sending it to the AE receiver
2. Features and limitations
TaDataWriter converts data from the DataX protocol into the internal data format of the AE cluster. TaDataWriter has the following conventions:
- Writes only to the AE cluster.
- Supports data compression. The currently supported compression formats are gzip and snappy.
- Supports multi-threaded transfer.
3. Usage
3.1 Download DataX
- Visit the DataX official site
- Download the DataX toolkit: DataX download link
wget https://download.thinkingdata.cn/tools/datax/datax.tar.gz
When the download completes, the output looks like this:
HTTP request sent, awaiting response... 200 OK
Length: 829372407 (791M) [application/x-gzip]
Saving to: ‘datax.tar.gz’
2019-12-23 10:34:17 (12.4 MB/s) - ‘datax.tar.gz’ saved [829372407/829372407]
3.2 Extract DataX
tar -zxvf datax.tar.gz
3.3 Install the ta-datax-writer plugin
- Download the ta-datax-writer plugin: ta-datax-writer download link
wget https://download.thinkingdata.cn/tools/release/ta-datax-writer.tar.gz
- Copy ta-datax-writer.tar.gz to the datax/plugin/writer directory
cp ta-datax-writer.tar.gz datax/plugin/writer
- Go to the datax/plugin/writer directory and run ll. The directory now contains ta-datax-writer.tar.gz
cd datax/plugin/writer/
ll
-rw-r--r-- 1 root root 8575618 Dec 23 10:38 ta-datax-writer.tar.gz
- Extract the plugin package
tar -zxvf ta-datax-writer.tar.gz
- Delete the package
rm -rf ta-datax-writer.tar.gz
- Run ll. The directory now contains the extracted ta-datax-writer plugin directory
ll
drwxr-xr-x 4 ta ta 4096 Nov 25 19:17 ta-datax-writer
4. Features
4.1 Configuration example
{
"job": {
"setting": {
"speed": {
"channel": 1
}
},
"content": [
{
"reader": {
"name": "streamreader",
"parameter": {
"column": [
{
"value": "123123",
"type": "string"
},
{
"value": "testbuy",
"type": "string"
},
{
"value": "2019-08-16 08:08:08",
"type": "date"
},
{
"value": "2222",
"type": "string"
},
{
"value": "2019-08-16 08:08:08",
"type": "date"
},
{
"value": "test",
"type": "bytes"
},
{
"value": true,
"type": "bool"
}
],
"sliceRecordCount": 10
}
},
"writer": {
"name": "ta-datax-writer",
"parameter": {
"thread": 3,
"type": "track",
"pushUrl": "http://{data receiving URL}",
"appid": "6f9e64da5bc74792b9e9c1db4e3e3822",
"column": [
{
"index": "0",
"colTargetName": "#distinct_id"
},
{
"index": "1",
"colTargetName": "#event_name"
},
{
"index": "2",
"colTargetName": "#time",
"type": "date",
"dateFormat": "yyyy-MM-dd HH:mm:ss.SSS"
},
{
"index": "3",
"colTargetName": "#account_id",
"type": "string"
},
{
"index": "4",
"colTargetName": "testDate",
"type": "date",
"dateFormat": "yyyy-MM-dd HH:mm:ss.SSS"
},
{
"index": "5",
"colTargetName": "os_1",
"type": "string"
},
{
"index": "6",
"colTargetName": "testBoolean",
"type": "boolean"
},
{
"colTargetName": "add_clo",
"value": "123123",
"type": "string"
}
]
}
}
}
]
}
}
4.2 Parameters
-
thread
- Description: The number of threads, used for concurrency inside each channel. It is unrelated to the number of DataX channels.
- Required: No
- Default value: 3
-
pushUrl
- Description: The endpoint URL.
- Required: Yes
- Default: None
-
uuid
- Description: Adds "#uuid": "uuid value" to the transferred data. Use it together with the data unique ID feature.
- Required: No
- Default value: false
-
type
- Description: The type of data to write: user_set or track.
- Required: Yes
- Default: None
-
compress
- Description: The text compression type. If you leave it empty, no compression is applied. Supported compression types are gzip and snappy.
- Required: No
- Default value: No compression
-
appid
- Description: The appid of the corresponding project.
- Required: Yes
- Default: None
-
column
- Description: The list of fields to read. type specifies the data type. index specifies which column of the reader the current column corresponds to (starting from 0). value makes the current column a constant: instead of reading data from the reader, the column is generated automatically from the value.
You can specify Column field information with the following configuration:
[
{
"type": "Number",
"colTargetName": "test_col", //Column name for the generated data
"index": 0 //Get the Number field from the first column transferred from the reader to DataX
},
{
"type": "string",
"value": "testvalue",
"colTargetName": "test_col"
//Generate a string field with the value testvalue inside TaDataWriter and use it as the current field
},
{
"index": 0,
"type": "date",
"colTargetName": "testdate",
"dateFormat": "yyyy-MM-dd HH:mm:ss.SSS"
}
]
-
When you specify Column information, you must set either index or value. type is optional. When you set the date type, you can optionally set dateFormat.
- Required: Yes
- Default value: All columns are read using the reader types
4.3 Type conversion
The types are defined by TaDataWriter:
| DataX internal type | TaDataWriter data type |
|---|---|
| Int | Number |
| Long | Number |
| Double | Number |
| String | String |
| Boolean | Boolean |
| Date | Date |

