Data deduplication tool
1. Overview
The data deduplication tool deduplicates duplicate event data in AE. You can deduplicate event data by time period and event type.
Because deduplication uses cluster computing resources, we recommend deduplicating data only when the data is abnormal, not as a frequent routine. Use this tool with caution.
2. Instructions
The data deduplication tool is available only to users of on-premises deployments. Log in as the root user to a server in the on-premises cluster where the secondary development tools are deployed (by default, the first three servers in a distributed deployment), and run su - ta.
Then run ta-tool dupevent_del to open the data deduplication tool.
2.1 Enter the appid of the project to process
You can find the project's appid on the project information card in Project Settings → Settings → Implementation in the AE console, or at the top of the Data → Source → Data Collection Guide page.
2.2 Confirm the project name
After you enter it, the tool shows the name of the project to deduplicate. Enter "y" to confirm or "n" to cancel.
2.3 Enter the names of the events to deduplicate
Next, enter the Event name of the events to deduplicate. The event name here is the key used when the data is sent, not the display name. You can look up event names in Data > Management > Events. To deduplicate multiple events, separate them with ",". After you finish, the tool shows the names of the events to deduplicate.
If you press Enter without typing anything, all event data is deduplicated:
2.4 Enter the column names to ignore in the deduplication logic
Next, enter the names of the fields to ignore when determining duplicates. Fields used internally by AE are already excluded from the duplicate check by default. For example, the "#server_time" and "#kafka_offset" fields aren't used to determine duplicates. To ignore multiple fields, separate them with ",". After you finish, the tool shows the names of the fields to ignore.
2.5 Enter the time range of the event data to deduplicate
Next, enter the time period of the data to deduplicate. The available time granularity is "day". Enter dates in the yyyy-MM-dd format. This item is required:
2.6 Final confirmation
Before deduplication starts, the tool asks for a final confirmation of the project name, the event names, and the time period to deduplicate. You need to confirm twice (the second time reminds you again that the data can't be recovered after deduplication). Processing starts only after you enter "y" both times. If anything is wrong, enter "n" to exit the tool and enter the information again:
2.7 Complete execution flow
After you confirm, data deduplication runs. The following screenshot shows the complete deduplication flow:
3. Notes
- Before you use the data deduplication tool, identify the cause of the duplicate data. If duplicate data keeps coming in while you deduplicate, the result can't be guaranteed.
- Deduplication uses cluster computing resources, so we don't recommend using it frequently.
- If the following message appears, the cluster is merging data. You can wait for it to finish on its own. If it stays stuck here for a long time, contact your operations team to investigate.
Deduplication will run after the background data sync task finishes. Retrying in 5 minutes. Please wait! Attempt 1

