Pre-integration preparation
The Basics section covers the AE data rules you must understand before integration.
The Essential integration information section lists the system parameters you need to prepare before integration.
The AE (Agentic Engine) system provides a data integration solution that covers all platforms.
Integrating with AE usually takes three steps: first, sort out your business requirements and work out a data collection plan, with the help of ThinkingAI analysts; next, your developers integrate the data according to the data collection plan; finally, verify that the data is integrated correctly. The integration process is shown in the figure below:
Before integrating data, it is important to understand the basics of the AE system. This document gives an overview of the integration knowledge you need, and also explains how to get help when you want to learn more.
This document is intended for everyone involved in the integration, including business staff, developers, and testers.
1. Overview
AE provides a data integration solution that covers all platforms. The main integration methods are:
- Client SDK: Collects device information and user behavior data that doesn't involve communication with the server; simple and easy to use
- Server-side SDK: Collects more accurate data; suitable for collecting core business data
- Data import tools: Usually used to import historical data; a server-side SDK combined with LogBus is also a common server-side data collection solution
For general app and web development, we provide:
- Native SDKs: Android SDK, iOS SDK
- Third-party frameworks: Flutter, React Native
- H5 development: JavaScript SDK, Connect H5 with the native SDK
- Major Mini program and quick app platforms: Mini program SDK
For Mini game development, we provide:
- Support for major game engines: LayaBox, Egret Engine, Cocos Creator
- Support for major Mini game and quick game platforms: Mini game SDK
For mobile game development, we provide:
For server-side data collection, we recommend the server-side SDK + Logbus solution. It performs well in terms of stability, timeliness, and efficiency of data import.
If you have heterogeneous historical data to import, or need to backfill some data into the AE system, consider importing it with DataX. Unlike Logbus, DataX is not a resident service. It cannot monitor new data and import it promptly, so it cannot guarantee real-time data. The advantage of DataX is that it supports importing heterogeneous data from multiple data sources and is easy to use.
If you use Filebeat and Logstash to collect logs and want to import the log data into the AE system, you can use the Filebeat +Logstash solution.
When designing a data collection plan, choose the solution that fits your product's technical architecture and business requirements. If you have questions about the collection plan, ask our analysts or technical support staff in the support group.
2. Basics
2.1 AE data model
Before integrating data, you first need to understand what data in AE is.
Designing a data collection plan is essentially deciding which user behavior events to collect based on your business analysis goals. For example, to analyze user top-ups, you might need to collect user payment behavior data. User behavior data can be broken down into: who (WHO) topped up (WHAT), when (WHEN), where (WHERE), and how (HOW), as shown in the figure below:
In AE, user behavior data is organized into user-related data and event-related data, which are stored in the user table and the event table respectively. User data mainly describes the user's status and properties that don't change often. Event data describes information related to specific behavior events.
In the data collection plan, you need to decide when to report user data and when to report events.
All of our data integration guides describe how to report event data and user data separately.
2.2 User identification rules
For each piece of user data or event data, you need to specify which user the data belongs to. In scenarios without an account system, you can use a device-related ID to uniquely identify a user. In scenarios with an account system, however, a user may generate data on multiple devices, and the analysis needs to combine the user's data from multiple platforms, so a device-related unique ID no longer works.
To handle both scenarios, data integration uses a combination of two user IDs to identify users:
- Distinct ID (#distinct_id): By default, the client generates a random distinct ID to identify the user, and also provides APIs to read and modify the default distinct ID.
- Account ID (#account_id): You can set the account ID when the user logs in. The account ID links data from multiple devices.
Each piece of data must contain a distinct ID or an account ID. By default, client SDKs generate a random distinct ID. After you call the login API to set the account ID, all data is reported with both the distinct ID and the account ID. When reporting from the server side, you need to pass in at least one of the two IDs.
In the AE backend, the unique ID that identifies a user is the AE user ID (the #user_id field). When data is received, we create a new user or bind the data to an existing user according to the specified user identification rules.
The user identification rules are very important. If the user IDs are not set correctly, data may be bound to the wrong users, which affects your analysis. Make sure you fully understand these rules before integration, and define your user identification approach in the data collection plan.
2.3 Data format
No matter which method you use to integrate data, the data sent to the data receiver uses a unified data format and the same data limits. The Data rules chapter describes the data format and the corresponding data limits in detail.
If you integrate data through an SDK, you only need to call the corresponding APIs, and the SDK formats the data as required and reports it. If you integrate data through a data import tool or the Restful API, you need to format the data as described in Data rules before reporting it.
For the data format, pay special attention to the naming rules and data types:
- Naming rules: Event names and property names can contain only lowercase letters, digits, and underscores _, must start with a letter, and cannot exceed 50 characters
- Property value data types:
| AE data type | Sample value | Description | Data type |
|---|---|---|---|
| Numeric | 123,1.23 | The data range is -9E15 to 9E15 | Number |
| Text | "ABC","Shanghai" | The default character limit is 2KB | String |
Time | "2019-01-01 00:00:00","2019-01-01 00:00:00.000" | "yyyy-MM-dd HH:mm:ss.SSS" or "yyyy-MM-dd HH:mm:ss". To represent a date, you can use "yyyy-MM-dd 00:00:00" | String |
| Boolean | true,false | - | Boolean |
| List | ["a","1","true"] | All elements in a list are converted to strings A list can contain up to 500 elements | Array(String) |
| Object | {hero_name:"Liu Bei",hero_level:22,hero_equipment: ["Twin Swords","Dilu"],hero_if_support:False} | Each child property (Key) in an object has its own data type. For value descriptions, see the regular property of the corresponding type above An object can contain up to 100 child properties | Object |
| Object group | [{hero_name:"Liu Bei",hero_level:22,hero_equipment: ["Twin Swords","Dilu"],hero_if_support:False}, {hero_name:"Liu Bei",hero_level:22,hero_equipment: ["Twin Swords","Dilu"],hero_if_support:False}] | Each child property (Key) in an object group has its own data type. For value descriptions, see the regular property of the corresponding type above An object group can contain up to 500 objects | Array(Object) |
Note: In the AE backend, the type of a property value is determined by the type of the first value received for that property. If the type of a property's value in later data doesn't match the previously determined type, the property is discarded.
In the AE backend, you may notice that some property names start with #. These are preset properties. Preset properties don't need any special setup; the SDK collects them by default. For details, see Preset properties and system fields.
Note in particular that when the data format or data type is set incorrectly, the data cannot be stored. Therefore, during and after integration, you may need to use the Tracking Plan module to check or monitor whether data is reported correctly, and fix any problems promptly.
Note that a single piece of data cannot exceed 1MB; otherwise, it cannot be stored.
3. Essential integration information
Before your developers start the data integration, make sure the following information is ready:
-
Project APP ID: The APP ID is generated when you create a project in the AE backend. You can also view it on the Project Settings page
-
Determine the data receiver URL
- If you use the cloud service, the receiver URL is: https://global-receiver-ta.thinkingdata.cn
- For on-premises deployments, you need to bind a domain name to the private cluster (or access point) and configure an SSL certificate
-
Verify the receiver URL: Visit https://YOUR_RECEIVER_URL/health-check in a browser. If the page returns ok, the URL is correct
-
The data collection plan, which should include:
- Integration method: client SDK, server-side SDK, data import tool, or a combination of these
- The content of the data to be integrated and when it is triggered
Congratulations, you have finished reading the pre-integration preparation guide. Next, you can start integrating data by following the integration guide for the method you have chosen.

