Importing external user properties by association
1. Introduction
In some cases, you need to import external user data into the AE cluster, but the user identifier in the data isn't the #account_id or #distinct_id in the AE system. For example, the data uses phone numbers, ID card numbers, or other identifier IDs as the primary key. To import this data into the AE system as user properties, use the update_user_by_foreignkey command to set up the association and update the external user properties in the AE system. All data sources supported by datax are currently supported:
2. Usage
2.1 Command
The data import command is as follows:
ta-tool update_user_by_foreignkey -conf <config files> [--date xxx]
2.2 Command parameters
2.2.1 -conf
The parameter passed in is the path of the import task's configuration file. Each task is one configuration file. You can import multiple tasks at the same time and use wildcards, for example: /data/config/* or ./config/*.json
2.2.2 --date
Optional parameter --date: Optional. Specifies the data date, which time macros use as the base time for replacement. If omitted, the current date is used. The format is YYYY-MM-DD. For details on using time macros, see Using time macros
2.3 Using time macros
You can use time macros in configuration files to replace time parameters. The ta-tool tool uses the import start time as the base, calculates time offsets based on the time macro parameters, and replaces the time macros in the configuration file. Supported time macro formats include @[{yyyyMMdd}], @[{yyyyMMdd}-{nday}], @[{yyyyMMdd}+{nday}], and more
-
yyyyMMddcan be replaced with any date format that JavadateFormatcan parse, for example:yyyy-MM-dd HH:mm:ss.SSS,yyyyMMddHH000000 -
n can be any integer and represents the time offset
-
day is the unit of the time offset and can be one of the following:
day,hour,minute,week,month -
Example: Assume the current time is
2018-07-01 15:13:23.234@[{yyyyMMdd}]is replaced with20180701@[{yyyy-MM-dd}-{1day}]is replaced with2018-06-30@[{yyyyMMddHH}+{2hour}]is replaced with2018070117@[{yyyyMMddHHmm00}-{10minute}]is replaced with20180701150300
3. Features
3.1 Configuration example
{
"job": {
"content": [{
"reader": {
"name": "mysqlreader",
"parameter": {
"username": "username",
"password": "password",
"connection": [
{
"querySql": [
"SELECT card_id, property1, property2,property3 FROM table1;"
],
"jdbcUrl": [
"jdbc:mysql://ip:port/database"
]
}
]
}
},
"writer": {
"parameter": {
"appid": "6f9e64da5bc74792b9e9c1db4e3e3822",
"column": [{
"type": "string",
"name": "card_id"
},
{
"type": "string",
"name": "property1"
},
{
"type": "string",
"name": "property2"
},
{
"type": "double",
"name": "property3"
}
],
"joinkey":{
"importDataKey": ["card_id"],
"taUserTableKey": ["card_id"]
}
}
}
}]
}
}
3.2 Parameters
3.2.1 reader section
- The reader configuration is the same as the readers supported by datax
3.2.2 writer section
-
appid
- Description: The appid of the corresponding project.
- Required: Yes
- Default: None
-
column
- Description: List of fields to read.
typespecifies the data type, andnamespecifies the column at the corresponding position in the reader, which is used as the property name when imported into the AE system.
- Description: List of fields to read.
You can specify the Column field information as follows:
[
{
"type": "double",
"name": "property1"
},
{
"type": "string",
"name": "property2"
},
{
"type": "bigint",
"name": "property3"
}
]
-
joinkey.importDataKey
- Description: The column name in the writer's column configuration used for the association.
- Required: Yes
- Default: None
-
joinkey.taUserTableKey
- Description: The column name in the user table of the AE system used for the association.
- Required: Yes
- Default: None
3.3 Type conversion
| DataX internal type | HIVE data type |
|---|---|
| Long | TINYINT,SMALLINT,INT,BIGINT |
| Double | FLOAT,DOUBLE |
| String | STRING,VARCHAR,CHAR |
| Boolean | BOOLEAN |
| Date | DATE,TIMESTAMP |

