User identification rules
Users may use your product on different devices and may also use it without logging in, which makes accurate user identification quite complex. AE adopts a solution that is relatively accurate and easy to understand. This document describes the user identification rules in detail and provides cases to help you understand them quickly.
1. A user is identified by three IDs:
- AE user ID (#user_id): The unique user ID in the AE system
- Account ID (#account_id): The user's login ID
- Distinct ID (#distinct_id): The user's ID when not logged in
2. The key to identifying a user is the AE user ID. For every record that the AE system receives, the system associates the corresponding AE user ID based on the record's account ID and distinct ID. If a record contains both an account ID and a distinct ID, the AE user ID is associated based on the account ID first. If there is no account ID, the AE user ID is associated based on the distinct ID
3. When data received by the AE backend contains a new account ID:
- If the data contains a distinct ID that is already associated with an AE user ID but not bound to an account ID, the account ID is bound to the distinct ID, and the two share one AE user ID
- If the distinct ID doesn't exist or is already bound to another account ID, no binding occurs, and the account ID is associated with a new AE user ID
4. When the AE backend receives data containing a new distinct ID:
- If the current data contains an account ID, the distinct ID is bound to that account ID, and both IDs are associated with the same AE user ID
- If the current data doesn't contain an account ID, no binding occurs, and the distinct ID is associated with a new AE user ID
5. AE user IDs and account IDs map one to one. One account ID can be bound to multiple distinct IDs, but a distinct ID can be bound to only one account ID.
1. Types of user identifiers
The AE platform mainly uses three user identifiers: Distinct ID (#distinct_id), Account ID (#account_id), and AE user ID (#user_id). This section briefly describes what each of these IDs means:
1.1 Distinct ID (#distinct_id)
The distinct ID identifies a user who isn't logged in. It is used to identify a user's data before login or outside the game, such as pre-registration data and ad data.
If you integrate with a client SDK, the SDK automatically assigns a unique distinct ID to the user. To customize a user's distinct ID, call identify to set it immediately after the SDK is initialized.
Avoid calling identify again to change the distinct ID after events have been uploaded. Doing so may cause serious data issues, such as users failing to match or duplicate users.
1.2 Account ID (#account_id)
The account ID identifies a logged-in user and is used to identify the user's data after login. Most games identify users along two dimensions: account and character. In general, we recommend using the finer granularity, that is, the character ID, as the account ID. If there is no character dimension, use the account login ID as the account ID.
If you integrate with a client SDK, call login to set the account ID when the user registers or logs in, or when the user creates a character or enters a server. The SDK saves the account ID, and every subsequent record carries it. If you call login again to set the account ID, the newly passed value becomes the account ID. You can also call logout to clear the account ID. Data sent after the account ID is cleared doesn't carry an account ID.
1.3 AE user ID (#user_id)
The AE user ID is the unique identifier that the AE system uses internally to identify users. When any valid record is stored, the system generates the record's AE user ID based on the account ID and distinct ID, which determines which user the record belongs to.
The AE user ID plays a very important role in analysis. Data tables such as event data, user properties, and cohort and tag tables are joined through the AE user ID. The deduplicated user count calculated in analysis models is essentially the deduplicated count of AE user IDs.
You can think of the user identification rules as the rules for generating the AE user ID of each record. The logic for generating the AE user ID has two steps:
- Update the ID mapping table: When data contains a new account ID or distinct ID, the AE system updates its internal ID mapping table
- Associate data with an AE user ID: The system looks up the user ID that corresponds to the account ID or distinct ID in the data in the ID mapping table, and associates each record with an AE user ID
2. Update the ID mapping table
Inside the AE system, there is an ID mapping table that is independent of the event table and the user table. It records how AE user IDs are associated with account IDs and distinct IDs. When data received by the system contains a new account ID or distinct ID, this ID mapping table is updated.
- If the data contains only an account ID or only a distinct ID, and the ID is received for the first time, the system creates a new AE user ID and associates it with the incoming ID.
If the data contains both an account ID and a distinct ID, an ID binding mechanism also applies. ID binding means binding the account ID to the distinct ID so that both are associated with the same AE user ID, which effectively links a user's data before and after login.
-
If both the account ID and the distinct ID are received for the first time, the two IDs are bound and associated with a new AE user ID
-
If the account ID exists in the ID mapping table and the distinct ID is new, the distinct ID is bound to the account ID. One account ID can be bound to multiple distinct IDs.
-
If the distinct ID exists in the ID mapping table and the account ID is new, there are two cases:
- If the distinct ID is already bound to another account ID, the distinct ID isn't bound to the new account ID, and the account ID is associated with a new AE user ID
- If the distinct ID isn't bound to any other account ID, the distinct ID is bound to the account ID
If both the account ID and the distinct ID in the data already exist in the ID mapping table, the ID mapping table isn't changed.
3. Associate data with an AE user ID
Next comes associating data with an AE user ID, which again follows two rules:
- If the data contains only an account ID or only a distinct ID, the AE user ID associated with that ID is used directly
- If the data contains both an account ID and a distinct ID, the AE user ID associated with the account ID is used
In short, the account ID takes priority when determining the associated user ID. When there is an account ID, the ID associated with the account ID is used. When there is no account ID, the ID associated with the distinct ID is used.
4. Case analysis
To help you better understand AE's user identification solution, this section uses cases to show how the identification rules work. The cases show how the backend assigns user IDs after receiving data. Focus on the value of #user_id at each step and the logic behind the association.
4.1 Only a distinct ID
When there is only a distinct ID, the user ID is generated from #distinct_id only
| #account_id | #distinct_id | #user_id |
|---|---|---|
| null | A | 1 |
| null | B | 2 |
| null | C | 3 |
| null | A | 1 |
In this scenario, the backend received three new distinct IDs, so it created a user ID three times. In step 4, distinct ID "A" already has a corresponding user ID "1", so no new user ID is created. It is treated as a previously created user, and its user ID is "1".
4.2 Distinct ID bound to a user ID but not to an account ID
When a distinct ID has a corresponding user ID but isn't bound to an account ID, passing in an account ID binds the account ID to the distinct ID
| #account_id | #distinct_id | #user_id |
|---|---|---|
| null | A | 1 |
| X | A | 1 |
In this scenario, the backend received a new distinct ID and created a new user ID for it. It then received a new account ID. Because the distinct ID wasn't bound to an account ID, the new account ID was bound to the distinct ID.
4.3 Distinct ID bound to both a user ID and an account ID
When a distinct ID is already associated with a user ID and bound to an account ID, a new account ID can't be bound to that distinct ID. The account ID can later try to bind to other distinct IDs:
| #account_id | #distinct_id | #user_id |
|---|---|---|
| X | A | 1 |
| Y | A | 2 |
| Y | B | 2 |
| null | B | 2 |
| null | A | 1 |
| Z | B | 3 |
In this scenario, distinct ID "A" is already bound to account ID "X", so the new account ID "Y" can't be bound to distinct ID "A" and is instead associated with a new user ID "2". In step 3, account ID "Y" and distinct ID "B" are passed in together. Distinct ID "B" hasn't been bound to an account ID, so the two are bound. As a result, distinct ID "B" passed in at step 4 is associated with user ID "2". Finally, the new account ID "Z" and distinct ID "B" are passed in together. Again, no binding occurs, and account ID "Z" is associated with a new user ID "3". The final state of the ID mapping table is as follows:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | B |
| 3 | Z | null |
5. Complex scenario analysis
Finally, here is user identification in a complex scenario. For easier understanding, we show the User table structure at key steps. Read it alongside the explanation of each step:
| Step | #account_id | #distinct_id | #user_id |
|---|---|---|---|
| 1 | null | A | 1 |
| 2 | X | A | 1 |
| 3 | Y | A | 2 |
| 4 | null | B | 3 |
| 5 | Y | B | 2 |
| 6 | Z | B | 3 |
| 7 | Z | C | 3 |
| 8 | Y | C | 2 |
| 9 | W | D | 4 |
| 10 | null | C | 3 |
Let's analyze the complex scenario above step by step:
(1) A new distinct ID "A" is passed in and bound to a newly created user ID "1"
(2) A new account ID "X" is added. Distinct ID "A" isn't bound to an account ID, so "X" and "A" are bound and associated with user ID "1".
(3) A new account ID "Y" is added. Distinct ID "A" is already bound to account ID "X", so a new user ID "2" is created and associated with "Y". At this point, account ID "Y" isn't bound to any distinct ID. The ID mapping table is as follows:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | null |
(4) A new distinct ID "B" is added, and a new user ID "3" is created and associated with it
(5) Account ID "Y" and distinct ID "B" both exist in the ID mapping table, so no binding occurs. The ID mapping table is as follows:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | null |
| 3 | null | B |
(6) A new account ID "Z" is added. Distinct ID "B" isn't bound to an account ID, so "Z" and "B" are bound and associated with user ID "3". The ID mapping table is as follows:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | null |
| 3 | Z | B |
(7) A new distinct ID "C" is added and bound to account ID "Z". The ID mapping table is as follows:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | null |
| 3 | Z | B, C |
(8) Account ID "Y" and distinct ID "C" both exist in the ID mapping table, so the ID mapping table doesn't change
(9) A new account ID "W" and a new distinct ID "D" are added. The two are bound and associated with a new user ID "4"
(10) Finally, the data contains only distinct ID "C", and its associated user ID "3" is returned
The final ID mapping table structure is:
| #user_id | #account_id | #distinct_id |
|---|---|---|
| 1 | X | A |
| 2 | Y | null |
| 3 | Z | B, C |
| 4 | W | D |

