Configure an OSS data source
OSS (Object Storage Service) is a massive, secure, and highly reliable cloud storage service suitable for storing unstructured data.
Limitations
Currently, only the following configuration is supported:
- Reading from OSS (offline read) and writing to AE Built-in Warehouse workspace tables
- Writing from AE Built-in Warehouse workspace tables to OSS (offline write)
Note:
- Writing directly from OSS to databases other than the AE Built-in Warehouse isn't supported yet
Supported field types
| File type | Field type | Offline read (Reader) | Offline write (Writer) |
|---|---|---|---|
orc parquet | boolean | Supported | Supported |
| tinyint | Supported | Supported | |
| smallint | Supported | Supported | |
| integer | Supported | Supported | |
| bigint | Supported | Supported | |
| real | Supported | Supported | |
| double | Supported | Supported | |
| decimal | Supported | Supported | |
| varchar | Supported | Supported | |
| char | Supported | Supported | |
| varbinary | Not supported | Not supported | |
| date | Supported | Supported | |
| timestamp | Supported | Supported | |
| array | Supported | Supported | |
| row | Supported | Supported | |
| map | Supported | Supported | |
| csv | string | Supported. You can configure a single-character column delimiter or \t (Tab); the default is a comma. For details, see "CSV and Text file reading rules." | Writing supported. The column delimiter defaults to a comma; for the options, see "CSV and Text file writing rules." |
| text | string | Supported. The column delimiter is parsed as a regular expression; the default is a comma. For details, see "CSV and Text file reading rules." | Writing supported. The column delimiter defaults to a comma; for the options, see "CSV and Text file writing rules." |
| json | string | Limited support; only a comma is supported as the file delimiter | Limited support; only a comma is supported as the file delimiter |
Create an OSS data source
In the DataOps Platform - Integration module, you can choose to add an OSS data source.
Data source configuration parameters
Fill in the configuration required by the data source and pass the connectivity test to create the OSS data source.
| Field name | Description |
|---|---|
| Basic Information | |
| *Datasource Name | Must be unique within the DataOps Platform space. Can contain only letters, digits, and underscores, and can't start with a digit or an underscore |
| Remarks | Optional |
| Data source configuration | |
| not distinguish environment / Independent setting | Choose one: not distinguish environment means the production and development environments share one configuration; Independent setting means the two environments are configured independently |
| *Endpoint | Access address of OSS |
| *Region | For example, oss-cn-shanghai |
| *Bucket | Address of the OSS bucket, for example: oss://seatunnel-test |
| *Access Key | Access Key used for access |
| *Access Secret | Access Secret used for access |
| Advanced | Other advanced parameters required to connect to the database; you can customize them |
Parameters whose names start with * are required; parameters without * are optional.
Create an offline sync task
After you create the OSS data source and pass the connectivity test as described above, you can configure an OSS offline read task for your scenario.
OSS as the Data Source or Data Target
Select OSS as the Data Source and configure the following parameters:
| Field name | Description |
|---|---|
| *Source Type | Select OSS as the type of the Data Source |
| *Datasource Name | An OSS data source registered on the data source management page; select it from the drop-down list. If you haven't created the data source yet, click the Data sources management button to create an OSS data source. |
| *File path | Reads files stored in OSS object storage through the file path
|
| *File type | Supports the orc, csv, text, parquet, and json types |
| Advanced | Empty directory marked succeeded: When the switch is on, the task is considered successful if the file directory is empty |
Parameters whose names start with * are required; parameters without * are optional.
Select OSS as the Data Target and configure the following parameters:
| Field name | Description |
|---|---|
| Data Source | Select OSS as the type of the Data Target, and select an OSS data source registered on the data source management page from the drop-down list |
| Written Directory | The OSS directory to write data to. Enter it after the preset prefix. It can't start with /
|
| File type | Supports the orc, csv, text, parquet, and json types. If you select csv or text, you can configure file writing rules. See "CSV and Text file writing rules" below |
| Write mode | Defaults to Insert into directory. Options:
|
CSV and Text file reading rules
When OSS is the Data Source and you select the csv or text file type, expand File Reading Rules and configure Column delimiter and Skip header. The default column delimiter is a comma, and Skip header defaults to No. If the file contains a header that shouldn't be read as data, select Yes.
- csv: The column delimiter can be a single valid character, such as a comma, semicolon, or vertical bar; you can also enter
\tfor Tab. Except for\t, multi-character delimiters aren't supported, and double quotes, carriage returns, and line breaks can't be used as delimiters. For example, to split columns by a vertical bar, enter|, not\|. - text: The column delimiter is parsed as a regular expression; for example,
\tmeans Tab,\s+means consecutive whitespace, and::means two consecutive colons. To split columns by a vertical bar or a period, enter\|or\.respectively, so that they aren't interpreted as special characters in the regular expression.
CSV and Text handle delimiters differently. After you switch the file type or change the reading rules, check the data preview and field mapping again.
CSV and Text file writing rules
When OSS is the Data Target and you select the csv or text file type, expand File Written Rule. The page provides the Column delimiter and Skip header options, which default to a comma and No respectively. The column delimiter box accepts up to 8 characters.
Write settings are independent of read settings, so don't apply the regular expression rules of the read side directly to write settings. After you use a custom delimiter or adjust the header option, first check the output file through manual execution to confirm that the delimiter and header are as expected, and only then use the settings for production sync.
Field mapping
After configuring the data source and the target, create field mappings. The system automatically syncs data from source fields to target fields based on the mappings. You can configure field mappings in three ways:
- Method 1: Custom selection. Select a source table field, then select the target field in the target table
- Method 2: Name Mapping. The system automatically maps fields with the same name in the source and target tables
- Method 3: Line Mapping. The system automatically maps fields in the same row
Note that each target field can correspond to only one source field
Field type conversion
During data synchronization, the system uses the CAST() function to try to convert source table fields to the types of the target fields. If the conversion fails, the value of the target field is set to NULL
Notes on data type conversion
- √ means the type conversion is supported
- √, null means the type conversion is supported; if the conversion fails, the value is set to NULL
| Source type/Target type | ROW | ARRAY | MAP | STRING | BOOLEAN | TINYINT | SMALLINT | INT | BIGINT | FLOAT | DOUBLE | DECIMAL | BYTES | DATE | TIMESTAMP | TIME |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ROW | √, null | |||||||||||||||
| ARRAY | √, null | |||||||||||||||
| MAP | √, null | |||||||||||||||
| STRING | √ | √, null | √, null | √, null | √, null | √, null | √, null | √, null | √, null | √, null | √, null | √, null | ||||
| BOOLEAN | √, null | √ | ||||||||||||||
| TINYINT | √, null | √ | √, null | √, null | √, null | |||||||||||
| SMALLINT | √, null | √, null | √ | √, null | √, null | |||||||||||
| INT | √, null | √, null | √, null | √ | √, null | |||||||||||
| BIGINT | √, null | √, null | √, null | √, null | √ | |||||||||||
| FLOAT | √, null | √ | √, null | √, null | ||||||||||||
| DOUBLE | √, null | √, null | √ | √, null | ||||||||||||
| DECIMAL | √, null | √, null | √, null | √ | ||||||||||||
| BYTES | √, null | √ | ||||||||||||||
| DATE | √, null | √ | √, null | |||||||||||||
| TIMESTAMP | √, null | √, null | √ | |||||||||||||
| TIME | √, null | √ |
Basic information settings
Finally, set the basic information of the integration plan, including the plan name, owner, synchronization rate, and remarks.
When you're done, click Save to create the integration plan.
Note: The plan name can't be changed after it's saved
The details page looks like this:
Mount an offline sync plan on a Flow
Mount on Flow
1. Start mounting
- On the integration plan details page, click Mount on task flow in the upper-right corner
2. Select a Flow
-
Select the target Flow from the drop-down menu
-
To create a new Flow:
- Click the New Flow shortcut button below the drop-down menu
- Or go to the Dev module to create one
-
Tip: If the target Flow isn't shown, click the ↻ refresh button on the right
3. Create a sync node
- In the Flow, create a node of the Offline sync plan type
- The node is linked to the current integration plan. When the node runs, it triggers a run of that integration plan
4. Complete mounting
- Click Create node and mount on it to complete the configuration
- After mounting succeeds, click Go to flow page to view the result right away
Note: The task node created by mounting is in the unreleased state. We recommend going to the Flow and releasing the node.
Unmount from a Flow
To unmount an offline sync plan from a Flow:
- If the Flow hasn't been released yet, go to the Flow that the plan is mounted on and delete the Offline sync plan task node in Dev Mode.
- If the Flow has already been released, after deleting the Offline sync plan task node in Dev Mode, release the Flow again. This also unmounts the node from the Production environment.

