Configure a COS data source
Tencent Cloud Object Storage (COS) stores object data such as files. By configuring a COS data source, you can read COS files in the DataOps Platform and write the data to AE Built-in Warehouse workspace tables through an offline sync plan for further processing and analysis.
Limitations
- COS currently supports offline read, with AE Built-in Warehouse workspace tables as the Data Target.
- COS can't be used as a Data Target yet, and syncing directly from COS to other external databases isn't supported.
- Workspace tables stored in COS can't take part in data integration yet. When configuring the Data Target, select a Built-in Warehouse workspace table that can take part in integration.
Before you start, prepare the region of the target bucket, the full bucket name, and a Secret Id and Secret Key that can access the bucket and read the target files. Also make sure that the platform's runtime environment can access the corresponding COS service.
Supported file types and configuration rules
COS supports reading the following five file types. When creating the data source, you only need to fill in the connection parameters; the file path, file type, and reading rules are configured in the offline sync plan.
| File type | Description |
|---|---|
| csv | You can configure the column delimiter and whether to skip the header. The default delimiter is a comma. A single valid character is supported, and you can also enter \t for Tab. Except for \t, multi-character delimiters aren't supported, and double quotes, carriage returns, and line breaks can't be used as delimiters. |
| text | You can configure the column delimiter and whether to skip the header. The default delimiter is a comma. The column delimiter is parsed as a regular expression; for example, \t means Tab and \s+ means consecutive whitespace. To split columns by a vertical bar or a period, enter \| or \. respectively. |
| json | Select the json file type and check the recognized fields and the data preview. |
| orc | Select the orc file type and check the recognized fields and the data preview. |
| parquet | Select the parquet file type and check the recognized fields and the data preview. |
The file type must match the actual file format. CSV and Text handle delimiters differently, so after you switch the file type or change the reading rules, check the data preview and field mapping again.
Create a COS data source
- Go to the target workspace and add a data source on the DataOps Platform → Integration → Data Sources page.
- Under the Semi-structured category, select COS and click Next step.
- Fill in the data source name, remarks, and connection parameters. Depending on how you use the data source, choose whether the production and development environments share one configuration or are configured separately.
- We recommend clicking Test in the connectivity test area to check the connection result. With Independent setting, test the production and development environments separately.
- After confirming the configuration, click Complete. After the data source is created, its details page opens, where you can view the connection configuration and connection status.
Data source configuration parameters
| Field name | Description |
|---|---|
| * Datasource Name | Must be unique within the current workspace; 1–80 characters. The name can't be changed after creation. |
| Remarks | Optional; up to 200 characters. You can describe what the data is used for. |
| Environment configuration | not distinguish environment by default: the production and development environments share one configuration. If you select Independent setting, fill in two sets of connection parameters. |
| * Region | Short name of the region where the bucket is located, such as ap-nanjing. Use the region of the actual bucket, not the Chinese region name. |
| * Bucket | Full bucket name including the APPID suffix, such as examplebucket-1250000000. Enter only the name, without the access domain, file path, or a protocol prefix such as cos://. |
| * Secret Id | Identifier of the Tencent Cloud access key, used together with Secret Key. |
| * Secret Key | Access key that corresponds to Secret Id. Enter credentials that have access to the target bucket. |
Parameters marked with * are required. The COS connection configuration doesn't require an Endpoint. For how to fill in the region and bucket, see Tencent Cloud's Preparations and Regions and Access Endpoints.
The following image shows the complete connection configuration area of an existing COS data source, including the environment description, region, bucket, Secret Id, and Secret Key. Sensitive values are masked.
Figure 1: Complete connection configuration of a COS data source. The production and development environments share one configuration; the gray masked areas are redacted parameter values.
If the connectivity test fails, follow the prompts on the page and check in turn whether the region matches the bucket, whether the Secret Id and Secret Key are correct, whether the account has the required access permissions, and whether the network from the platform to COS is reachable. After you change connection parameters, we recommend testing again.
Create an offline sync task
After configuring the data source, go to Integration → Offline Sync and add an integration plan. Select COS as the Data Source, then select the COS data source you created.
Configure data source
| Config items | Description |
|---|---|
| * Data source | Select a configured COS data source. |
| * File path | Relative file path within the bucket, such as incoming/events/data.parquet. The page shows the preset prefix, so you don't need to enter the protocol and bucket again. You can insert space parameters or use the * wildcard as needed, such as incoming/events/*. |
| * File type | Select csv, text, json, orc, or parquet to match the actual file. |
| File Reading Rules | If you select csv or text, you can configure the column delimiter and the skip header option. For details, see "Supported file types and configuration rules." |
| Empty directory marked succeeded | Located under Advanced and off by default. When it's on, the task can be marked as successful if the file directory is empty. Turn it on only when it's acceptable to have no input files. |
After entering the path, use path viewing and data preview to check the actual read scope and parsing results. Once the path, file type, and reading rules are correct, configure the Data Target and field mapping.
Configure the Data Target and field mapping
For the Data Target, select an AE Built-in Warehouse workspace table and configure the write mode for the target table. Overwriting affects existing data in the target table, so choose based on your sync needs.
After configuring the source and target, map source fields to target fields. You can specify the mappings manually or use Name Mapping or Line Mapping. Each target field can correspond to only one source field. Check the meaning and type of each field to avoid incorrect mappings based only on the same name or position.
The following image shows a field mapping example from an existing COS sync plan. The field names and mappings are for illustration only; follow your own file content and target table structure.
Figure 2: Mapping between source fields and target fields in a COS offline sync plan.
Save and check the plan
Fill in the basic information, such as the plan name, owner, synchronization rate, and remarks, then click Save. After saving, you can check Field Mapping and Solution Configuration on the details page. To verify the sync results, use Manual Execution and check the execution records and the target table data.
Mount an offline sync plan on a Flow
To run COS data sync through a Flow, mount the saved offline sync plan on a Flow.
- On the integration plan details page, click Mount on task flow.
- Select the target Flow. To create a new Flow, use the create option in the selection box, or go to the Dev module to create one.
- Enter the sync node name, select an execution mode supported on the page, and set the owner and remarks as needed. The node is linked to the current integration plan.
- Click Create node and mount on it. After mounting succeeds, click Go to flow page to view the node configuration.
After mounting, the new node is in the development state. Check the Flow's dependencies and scheduling configuration and release the Flow. The plan then runs on the Flow's schedule.
To unmount, delete the corresponding offline sync node in the Flow's Dev Mode. If the Flow has already been released, release it again so that the mount relationship in the production environment is updated as well.

