DataX integration for S3 data
DataX plugin: s3reader
Note that DataX treats all the Objects synced in one job as the same data table. You must make sure that all the Objects fit the same schema.
Sample JSON:
{
"job": {
"setting": {
"speed": {
"byte": 10485760
}
},
"content": [
{
"reader": {
"name": "s3reader",
"parameter": {
"accessKey": "xx",
"secretKey": "xx",
"tempPath": "xxx",
"region": "xxxxx",
"bucket": "xxxxx",
"object": [
"aa/bb/*"
],
"compress": "gzip",
"column": [
{
"type": "string",
"index": 0
}
],
"encoding": "UTF-8",
"fieldDelimiter": "\t"
}
},
"writer": {
"name": "streamwriter",
"parameter": {
"encoding": "",
"print": true
}
}
}
]
}
}
S3Reader currently supports the following features:
- Supports reading TXT files only, and the schema in the TXT file must be a two-dimensional table.
- Supports CSV-like files with custom delimiters.
- Supports reading multiple data types (represented as String), column pruning, and column constants
- Supports recursive reading and file name filtering.
- Supports text compression. The available compression formats are zip, gzip, and bzip2. Note that one archive can't contain multiple files.
- Multiple objects can be read concurrently.
Parameter description
-
accessKey
- Description: the accessKey of S3
- Required: Yes
- Default: None
-
secretKey
- Description: the secretKey of S3
- Required: Yes
- Default: None
-
region
- Description: the region of S3
- Required: Yes
- Default: None
-
tempPath
- Description: the absolute path of the local cache
- Required: Yes
- Default: None
- Fault tolerance: objectKeys that can't be retrieved are written to FailedFile.txt in this directory
-
bucket
- Description: the bucket of S3
- Required: Yes
- Default: None
-
object
- Description: the object information of S3. Note that you can enter multiple Objects here.
When a single S3 Object is specified, S3Reader can currently extract data with only a single thread.
When multiple S3 Objects are specified, S3Reader can extract data with multiple threads. The number of concurrent threads is specified by the number of channels.
When the wildcard
*is specified, S3Reader tries to traverse multiple Objects.- Required: Yes
- Default: None
-
column
- Description: the list of fields to read. type specifies the type of the source data, index specifies which column of the text the current column comes from (starting from 0), and value specifies that the current column is a constant: the data isn't read from the source file, and the column is generated automatically from the value.
By default, you can read all data as the String type, configured as follows:
"column": ["*"]You can specify the Column field information, configured as follows:
"column": [{"type": "long","index": 0},{"type": "string","value": "alibaba"}]When you specify Column information, type is required, and you must choose either index or value.
- Required: Yes
- Default value: read all as the string type
-
fieldDelimiter
- Description: Field delimiter for reading
- Required: Yes
- Default value: ,
-
compress
- Description: the text compression type. Leaving it empty (the default) means no compression. The supported compression types are zip, gzip, and bzip2.
- Required: No
- Default value: no compression
-
encoding
- Description: the encoding of the files to read. Currently, only utf-8/gbk are supported.
- Required: No
- Default value: utf-8
-
nullFormat
- Description: in text files, null (a null pointer) can't be represented by a standard string, so DataX provides nullFormat to define which strings can represent null.
For example, if you configure nullFormat="\N", DataX treats source data "\N" as a null field.
- Required: No
- Default value: \N
-
skipHeader
- Description: CSV-like files may have a header row of titles that needs to be skipped. Not skipped by default.
- Required: No
- Default value: false
-
csvReaderConfig
- Description: the parameter configuration for reading CSV files, of the Map type. CSV files are read with CsvReader, which has many settings. If you don't configure them, the defaults are used.
- Required: No
- Default: None
Common configuration:
"csvReaderConfig":{
"safetySwitch": false,
"skipEmptyRecords": false,
"useTextQualifier": false
}
All configuration items and their default values. When you configure them, use exactly the following field names in the csvReaderConfig map:
boolean caseSensitive = true;
char textQualifier = 34;
boolean trimWhitespace = true;
boolean useTextQualifier = true;//Whether to use the CSV text qualifier
char delimiter = 44;//Delimiter
char recordDelimiter = 0;
char comment = 35;
boolean useComments = false;
int escapeMode = 1;
boolean safetySwitch = true;//Whether to limit a single column to 100000 characters
boolean skipEmptyRecords = true;//Whether to skip empty lines
boolean captureRawRecord = true;
Type conversion
S3 itself doesn't provide data types. The types are defined by DataX S3Reader:
| DataX internal type | S3 data type |
|---|---|
| Long | Long |
| Double | Double |
| String | String |
| Boolean | Boolean |
| Date | Date |
Where:
- S3 Long is the string representation of an integer in S3 text, such as "19901219".
- S3 Double is the string representation of a Double in S3 text, such as "3.1415".
- S3 Boolean is the string representation of a Boolean in S3 text, such as "true" and "false". Case-insensitive.
- S3 Date is the string representation of a Date in S3 text, such as "2014-12-31". You can specify the format of Date.

