DataX engine
1. DataX overview
DataX is an open-source project released by Alibaba (for details, visit the DataX GitHub homepage). It's an efficient offline data sync tool, commonly used to sync data between heterogeneous data sources.
DataX uses a Framework + plugin architecture. Reading from and writing to data sources correspond to Reader and Writer plugins respectively, and each data source has a corresponding Reader or Writer. By default, DataX provides a wide range of Readers and Writers to fit many mainstream data sources. The Framework connects Readers and Writers and handles core processes in sync tasks such as data processing and transfer.
A DataX data sync task is mainly controlled by a configuration file. The most important configurations in it are the Reader and Writer configurations, which define how to extract data from a data source and how to write the extracted data to a data source. By using the Reader and Writer for the corresponding data sources in the configuration file, you can sync data between heterogeneous data sources.
In ta-tool, we integrated the DataX engine and wrote plugins for the AE cluster (that is, the AE cluster Reader and Writer). With the AE cluster plugins, the AE cluster can be used as a DataX data source.
With the DataX engine in ta-tool, you can perform the following data syncs:
- Import data from other databases into the AE cluster. This requires an existing DataX Reader plugin and the AE Writer
- Export data from the AE cluster to other databases. This requires the AE Reader and an existing DataX Writer plugin
2. How to use the DataX engine
To use the DataX engine in ta-tool for multi-data-source sync tasks, first write the Config file for the DataX task in the AE cluster, and then run the DataX command in the secondary development component to read the Config file and run the data sync task.
2.1 Configuration file example
The Config file of a DataX task must be a JSON file. The JSON configuration template is as follows:
{
"job": {
"content": [
{
"reader": {
"name": "streamreader",
"parameter": {
"sliceRecordCount": 10,
"column": [
{
"type": "long",
"value": "10"
},
{
"type": "string",
"value": "hello, world-DataX"
}
]
}
},
"writer": {
"name": "streamwriter",
"parameter": {
"encoding": "UTF-8",
"print": true
}
}
}
],
"setting": {
"speed": {
"channel": 5
}
}
}
}
The entire configuration file is a JSON object. The outermost element is "job", which contains two elements: "content" and "setting". The elements in "content" contain the reader and writer information. You can find the AE cluster Reader and Writer later in this document. "channel" in "speed" under "setting" is the number of tasks run concurrently.
The main parts to configure in the configuration file are the "reader" and "writer" elements in "content", which configure the Reader plugin that reads data and the Writer plugin that writes data respectively. For how to configure the Reader and Writer plugins preset in DataX, see the Support Data Channels section of DataX.
2.2 Run the DataX command
After you finish writing the configuration file, run the following command to read the configuration file and start the data sync task.
ta-tool datax_engine -conf <configPath> [--date <date>]
The parameter passed in is the path of the configuration file.
3. DataX plugins for the AE cluster
3.1 Plugins used inside the cluster
| Type | Data Source | Reader (read) | Writer (write) | Documentation |
|---|---|---|---|---|
| AE system | TE | √ | √ | Read, write |
| Custom table | TE | √ | Write | |
| JSON text | TE | √ | Write |
3.2 Plugins used outside the cluster
| Type | Data Source | Reader (read) | Writer (write) | Documentation |
|---|---|---|---|---|
| AE system | TE | √ | Write |
3.3 DataX native plugins
| Type | Data Source | Reader (read) | Writer (write) | Documentation |
|---|---|---|---|---|
| RDBMS relational databases | MySQL | √ | √ | Read, Write |
| Oracle | √ | √ | Read, Write | |
| SQLServer | √ | √ | Read, Write | |
| PostgreSQL | √ | √ | Read, Write | |
| DRDS | √ | √ | Read, Write | |
| General RDBMS (supports all relational databases) | √ | √ | Read, Write | |
Alibaba Cloud data warehouse storage | ODPS | √ | √ | Read, Write |
| ADS | √ | Write | ||
| OSS | √ | √ | Read, Write | |
| OCS | √ | √ | Read, Write | |
| NoSQL data storage | OTS | √ | √ | Read, Write |
| Hbase0.94 | √ | √ | Read, Write | |
| Hbase1.1 | √ | √ | Read, Write | |
| Phoenix4.x | √ | √ | Read, Write | |
| Phoenix5.x | √ | √ | Read, Write | |
| MongoDB | √ | √ | Read, Write | |
| Hive | √ | √ | Read, Write | |
| Cassandra | √ | √ | Read, Write | |
| Unstructured data storage | TxtFile | √ | √ | Read, Write |
| FTP | √ | √ | Read, Write | |
| HDFS | √ | √ | Read, Write | |
| Elasticsearch | √ | √ | Read, Write | |
| Time series databases | OpenTSDB | √ | Read | |
| TSDB | √ | √ | Read, Write |

