Skip to main content

Configure an S3 data source

Last updated 10/05/2026

Amazon S3 (Simple Storage Service) is an object storage service provided by Amazon Web Services (AWS) and one of the most popular cloud storage solutions in the world.

Limitations​

Currently, only the following configuration is supported:

  • Reading from S3 data sources (offline read) and writing to AE Built-in Warehouse workspace tables
  • Writing from AE Built-in Warehouse workspace tables to S3 (offline write)

Note:

  • Writing directly from S3 to databases other than the AE Built-in Warehouse isn't supported yet

Supported field types​

File typeField typeOffline read (Reader)Offline write (Writer)

orc

parquet

booleanSupportedSupported
tinyintSupportedSupported
smallintSupportedSupported
integerSupportedSupported
bigintSupportedSupported
realSupportedSupported
doubleSupportedSupported
decimalSupportedSupported
varcharSupportedSupported
charSupportedSupported
varbinaryNot supportedNot supported
dateSupportedSupported
timestampSupportedSupported
arraySupportedSupported
rowSupportedSupported
mapSupportedSupported
csvstringSupported. You can configure a single-character column delimiter or \t (Tab); the default is a comma. For details, see "CSV and Text file reading rules."Writing supported. The column delimiter defaults to a comma; for the options, see "CSV and Text file writing rules."
textstringSupported. The column delimiter is parsed as a regular expression; the default is a comma. For details, see "CSV and Text file reading rules."Writing supported. The column delimiter defaults to a comma; for the options, see "CSV and Text file writing rules."
jsonstringLimited support; only a comma is supported as the file delimiterLimited support; only a comma is supported as the file delimiter

Create an S3 data source​

In the DataOps Platform - Integration module, you can choose to add an S3 data source.

Data source configuration parameters

Fill in the configuration required by the data source and pass the connectivity test to create the S3 data source.

Parameters whose names start with * are required; parameters without * are optional.

Field nameDescription
Basic Information
*Datasource NameMust be unique within the DataOps Platform space. Can contain only letters, digits, and underscores, and can't start with a digit or an underscore
RemarksOptional
Data source configuration
not distinguish environment / Independent settingChoose one: not distinguish environment means the production and development environments share one configuration; Independent setting means the two environments are configured independently
*EndpointAccess address of S3
*RegionFor example, ap-northeast-1
*BucketAddress of the S3 bucket, for example: S3://seatunnel-test
*Access KeyAccess Key used for access
*Access SecretAccess Secret used for access
AdvancedOther advanced parameters required to connect to the database; you can customize them

Parameters whose names start with * are required; parameters without * are optional.

Create an offline sync task​

After you create the S3 data source and pass the connectivity test as described above, you can configure an S3 offline read task for your scenario.

S3 as the Data Source or Data Target​

Select S3 as the Data Source and configure the following parameters:

Field nameDescription
*Source TypeSelect S3 as the type of the Data Source
*Datasource NameAn S3 data source registered on the data source management page; select it from the drop-down list.
If you haven't created the data source yet, click the Data sources management button to create an S3 data source.
*File path

Reads files stored in S3 object storage through the file path

  • The file path supports space parameters
  • The file path supports the wildcard * (not supported in Data Target paths yet)
  • You can view the full path information and preview data
*File typeSupports the orc, csv, text, parquet, and json types
AdvancedEmpty directory marked succeeded: When the switch is on, the task is considered successful if the file directory is empty

Parameters whose names start with * are required; parameters without * are optional.

Select S3 as the Data Target and configure the following parameters:

Field nameDescription
Data SourceSelect S3 as the type of the Data Target, and select an S3 data source registered on the data source management page from the drop-down list
Written Directory

The S3 directory to write data to. Enter it after the preset prefix. It can't start with /

  • The written directory supports space parameters (click Parameter)
  • The written directory doesn't support wildcards
  • Click Full Path preview to view the full path information
File typeSupports the orc, csv, text, parquet, and json types. If you select csv or text, you can configure file writing rules. See "CSV and Text file writing rules" below
Write mode

Defaults to Insert into directory. Options:

  • Insert into directory (Insert Into)
  • Overwrite directory (Insert Overwrite): Deletes all files and subdirectories in the written directory first, then writes new data files
  • Error if exist (Error If Exist)

CSV and Text file reading rules​

When OSS is the Data Source and you select the csv or text file type, expand File Reading Rules and configure Column delimiter and Skip header. The default column delimiter is a comma, and Skip header defaults to No. If the file contains a header that shouldn't be read as data, select Yes.

  • csv: The column delimiter can be a single valid character, such as a comma, semicolon, or vertical bar; you can also enter \t for Tab. Except for \t, multi-character delimiters aren't supported, and double quotes, carriage returns, and line breaks can't be used as delimiters. For example, to split columns by a vertical bar, enter |, not \|.
  • text: The column delimiter is parsed as a regular expression; for example, \t means Tab, \s+ means consecutive whitespace, and :: means two consecutive colons. To split columns by a vertical bar or a period, enter \| or \. respectively, so that they aren't interpreted as special characters in the regular expression.

CSV and Text handle delimiters differently. After you switch the file type or change the reading rules, check the data preview and field mapping again.

CSV and Text file writing rules​

When OSS is the Data Target and you select the csv or text file type, expand File Written Rule. The page provides the Column delimiter and Skip header options, which default to a comma and No respectively. The column delimiter box accepts up to 8 characters.

Write settings are independent of read settings, so don't apply the regular expression rules of the read side directly to write settings. After you use a custom delimiter or adjust the header option, first check the output file through manual execution to confirm that the delimiter and header are as expected, and only then use the settings for production sync.

Field mapping​

After configuring the data source and the target, create field mappings. The system automatically syncs data from source fields to target fields based on the mappings. You can configure field mappings in three ways:

  • Method 1: Custom selection. Select a source table field, then select the target field in the target table
  • Method 2: Name Mapping. The system automatically maps fields with the same name in the source and target tables
  • Method 3: Line Mapping. The system automatically maps fields in the same row

Note that each target field can correspond to only one source field

Field type conversion

tip

During data synchronization, the system uses the CAST() function to try to convert source table fields to the types of the target fields. If the conversion fails, the value of the target field is set to NULL

Notes on data type conversion

  • √ means the type conversion is supported
  • √, null means the type conversion is supported; if the conversion fails, the value is set to NULL
Source type/Target typeROWARRAYMAPSTRINGBOOLEANTINYINTSMALLINTINTBIGINTFLOATDOUBLEDECIMALBYTESDATETIMESTAMPTIME
ROW√, null
ARRAY√, null
MAP√, null
STRING√√, null√, null√, null√, null√, null√, null√, null√, null√, null√, null√, null
BOOLEAN√, null√
TINYINT√, null√√, null√, null√, null
SMALLINT√, null√, null√√, null√, null
INT√, null√, null√, null√√, null
BIGINT√, null√, null√, null√, null√
FLOAT√, null√√, null√, null
DOUBLE√, null√, null√√, null
DECIMAL√, null√, null√, null√
BYTES√, null√
DATE√, null√√, null
TIMESTAMP√, null√, null√
TIME√, null√

Basic information settings​

Finally, set the basic information of the integration plan, including the plan name, owner, synchronization rate, and remarks.

When you're done, click Save to create the integration plan.

Note: The plan name can't be changed after it's saved

The details page looks like this:

Mount an offline sync plan on a Flow​

Mount on Flow​

1. Start mounting

  • On the integration plan details page, click Mount on task flow in the upper-right corner

2. Select a Flow

  • Select the target Flow from the drop-down menu

  • To create a new Flow:

    • Click the New Flow shortcut button below the drop-down menu
    • Or go to the Dev module to create one
  • Tip: If the target Flow isn't shown, click the ↻ refresh button on the right

3. Create a sync node

  • In the Flow, create a node of the Offline sync plan type
  • The node is linked to the current integration plan. When the node runs, it triggers a run of that integration plan

4. Complete mounting

  • Click Create node and mount on it to complete the configuration
  • After mounting succeeds, click Go to flow page to view the result right away
tip

Note: The task node created by mounting is in the unreleased state. We recommend going to the Flow and releasing the node.

Unmount from a Flow​

To unmount an offline sync plan from a Flow:

  • If the Flow hasn't been released yet, go to the Flow that the plan is mounted on and delete the Offline sync plan task node in Dev Mode.
  • If the Flow has already been released, after deleting the Offline sync plan task node in Dev Mode, release the Flow again. This also unmounts the node from the Production environment.
Was this page helpful?