manuallatestProduct Manual / Data archiving and reloading
Data archiving and reloading
1. Data archiving (ta_data_archive)
Data archiving migrates historical data, or data you don't need for now, to low-cost storage. This frees up disk resources on the AE cluster and reduces costs.
1.1 Archive command
#Start
ta-tool data_archive start
#Stop
ta-tool data_archive stop
#Rerun after failure
ta-tool data_archive retry -jobid *******
1.2 Archiving methods
1.2.1 S3 method
1.2.1.1 Environment setup
- Sign up for the Amazon S3 service
- Create the bucket for archiving. We recommend placing the bucket in the same region as the AE cluster servers
- Create a key that can access the bucket
1.2.1.2 Command example
[ta@ta1 ~]$ ta-tool data_archive start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 5487f6b**********f9c379aa9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > s3
------------------------------------------------------------
Enter the S3 AccesskeyID> YOUR_ACCESS_KEY_ID
------------------------------------------------------------
Enter the S3 secretAccessKey> YOUR_SECRET_ACCESS_KEY
------------------------------------------------------------
Enter the S3 region code> cn-****-1
------------------------------------------------------------
Enter the S3 bucket name> ta************ive
------------------------------------------------------------
Enter the S3 file storage class (default: STANDARD)> S*****D
------------------------------------------------------------
Enter the target directory of the project archive> data*****_test
------------------------------------------------------------
1.2.1.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. You need to specify the jobid to rerun the job if it fails.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select S3 as the archive storage type
- Enter the S3 accesskeyid
- Enter the secretAccessKey (managed in the S3 IAM service)
- Specify the region code of the bucket
- Enter the bucket name
- Select the storage class (standard by default). The GLACIER and DEEP_ARCHIVE storage classes are designed for low-cost data archiving, but the data must be restored before it can be recovered, which is cumbersome.
- The target directory of the archive (this directory is created in the target bucket, and the archived data is placed in it)
1.2.2 HDFS method
1.2.2.1 Environment setup
- Prepare an HDFS environment that can reach the AE cluster over the network
1.2.2.2 Command example
[ta@ta1 ~]$ ta-tool data_archive start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 5487************a9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > hdfs
------------------------------------------------------------
Enter the HFDS URL of the project archive> hdfs-nm-url
------------------------------------------------------------
Enter the HFDS user name of the project archive> hdfsUserName
------------------------------------------------------------
Enter the target directory of the project archive> hdfs******test
------------------------------------------------------------
1.2.2.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. It is used to specify the job when you rerun it after a failure.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select hdfs as the archive storage type
- Enter the HDFS address of the writer side. If the port is the default one, enter only the hostname
- Enter the HDFS user name of the writer side
- Enter the target directory of the archive. Use an absolute path. Otherwise, the data is stored in /user/hdfs user directory/target directory/
1.2.3 rsync method
1.2.3.1 Environment setup
- Set up the server in rsync daemon mode, and copy the key file to the node in the AE cluster where you run the command
1.2.3.2 Command example
[ta@ta1 ~]$ ta-tool data_archive start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 548*****************9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > rsync
------------------------------------------------------------
Enter the IP address of the target RSYNC server> rsyncIp
------------------------------------------------------------
Enter the port of the target RSYNC server> rsyncPort
------------------------------------------------------------
Enter the user name of the target RSYNC server> rsyncUser
------------------------------------------------------------
Enter the key file location of the target RSYNC server> passwordFilePath
------------------------------------------------------------
Enter the module name of the target RSYNC server> modelName
------------------------------------------------------------
sending incremental file list
/tmp/
/tmp/d41d8c*****ecf8427e.data
sent 99 bytes received 15 bytes 228.00 bytes/sec
total size is 11 speedup is 0.10 (DRY RUN)
Enter the target directory of the project archive> rsync******test_dir
1.2.3.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. It is used to specify the job when you rerun it after a failure.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select rsync as the archive storage type
- Enter the IP of the rsync server
- Enter the port of the rsync server
- Enter the rsync user name
- Enter the location of the rsync key file. Put the file in a directory and make sure its permission is set with chmod 600
- Enter the rsync module name (this step uses the information you entered earlier to verify that rsync is available)
- Enter the target directory of the archive
2. Data reloading (ta_data_reload)
Data reloading imports previously archived data back into the AE cluster so that you can use it again. It is typically used when you view trends over the years.
Before importing, make sure there is enough disk space.
2.1 Reload command
#Start
ta-tool data_reload start
#Stop
ta-tool data_reload stop
#Rerun after failure
ta-tool data_reload retry -jobid *******
2.2 Reloading methods
2.2.1 S3 method
2.2.1.1 Environment setup
- Sign up for the Amazon S3 service
- Create the bucket for archiving. We recommend placing the bucket in the same region as the AE cluster servers
- Create a key that can access the bucket
2.2.1.2 Command example
[ta@ta1 log]$ ta-tool data_reload start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 5487f6************a9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > s3
------------------------------------------------------------
Enter the S3 AccesskeyID> YOUR_ACCESS_KEY_ID
------------------------------------------------------------
Enter the S3 secretAccessKey> YOUR_SECRET_ACCESS_KEY
------------------------------------------------------------
Enter the S3 region code> cn*****-1
------------------------------------------------------------
Enter the S3 bucket name> ta*****ve
------------------------------------------------------------
Enter the target directory of the project archive> data*******t_1
------------------------------------------------------------
2.2.1.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. It is used to specify the job when you rerun it after a failure.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select S3 as the archive storage type
- Enter the S3 accesskeyid
- Enter the secretAccessKey (managed in the S3 IAM service)
- Specify the region code of the bucket
- Enter the bucket name
- Select the storage class (standard by default). If the storage class is GLACIER or DEEP_ARCHIVE, restore the data in S3 in advance. Otherwise, the data can't be retrieved
- The target directory of the archive (this directory is created in the target bucket, and the archived data is placed in it)
Note: When you enter the parameters, make sure the bucket name and directory path are the same as those used for archiving.
2.2.2 HDFS method
2.2.2.1 Environment setup
- Prepare an HDFS environment that can reach the AE cluster over the network
2.2.2.2 Command example
[ta@ta1 log]$ ta-tool data_reload start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 5487*******************9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > hdfs
------------------------------------------------------------
Enter the HFDS URL of the project archive> hdfs-nm-url
------------------------------------------------------------
Enter the target directory of the project archive> hdfs******test
------------------------------------------------------------
2.2.2.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. It is used to specify the job when you rerun it after a failure.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select hdfs as the archive storage type
- Enter the HDFS address of the writer side. If the port is the default one, enter only the hostname
- Enter the HDFS user name of the writer side
- Enter the target directory of the archive
Note: When you enter the parameters, make sure the directory path is the same as the one used for archiving.
2.2.3 rsync method
2.2.3.1 Environment setup
- Set up the server in rsync daemon mode, and copy the key file to the node in the AE cluster where you run the command
2.2.3.2 Command example
[ta@ta1 log]$ ta-tool data_reload start
Enter the JobId for this job. If you leave it empty, one is generated randomly in the background>
------------------------------------------------------------
Enter the appid of the project to archive> 54****************9bb
------------------------------------------------------------
Enter the start time of the project archive: YYYY-MM-DD > 2018-01-01
------------------------------------------------------------
Enter the end time of the project archive: YYYY-MM-DD > 2018-12-31
------------------------------------------------------------
Enter the event type of the project archive (optional)>
------------------------------------------------------------
Enter the archive storage type: hdfs or rsync or s3 > rsync
------------------------------------------------------------
Enter the IP address of the target RSYNC server> rsyncIp
------------------------------------------------------------
Enter the port of the target RSYNC server> rsyncPort
------------------------------------------------------------
Enter the user name of the target RSYNC server> rsyncUser
------------------------------------------------------------
Enter the key file location of the target RSYNC server> passwordFilePath
------------------------------------------------------------
Enter the module name of the target RSYNC server> modelName
------------------------------------------------------------
sending incremental file list
/tmp/
/tmp/d41d8cd98f00b204e9800998ecf8427e.data
sent 99 bytes received 15 bytes 20.73 bytes/sec
total size is 11 speedup is 0.10 (DRY RUN)
Enter the target directory of the project archive> rsync******test_dir
2.2.3.3 Steps
- Enter the jobid. You can define it yourself or let it be generated in the background. It is used to specify the job when you rerun it after a failure.
- Enter the project appid
- Enter the start date (outside the most recent month)
- Enter the end date (outside the most recent month)
- Enter a specific event type (optional) to archive only that event type
- Select rsync as the archive storage type
- Enter the IP of the rsync server
- Enter the port of the rsync server
- Enter the rsync user name
- Enter the location of the rsync key file. Put the file in a directory and make sure its permission is set with chmod 600
- Enter the rsync module name (this step uses the information you entered earlier to verify that rsync is available)
- Enter the target directory of the archive
Note: When you enter the parameters, make sure the directory path is the same as the one used for archiving.
Was this page helpful?

