Skip to main content

Configure a StarRocks data source

Last updated 10/05/2026

StarRocks is a high-performance analytical data warehouse that uses technologies such as vectorization, an MPP architecture, a CBO, intelligent materialized views, and a real-time updatable columnar storage engine to deliver multi-dimensional, real-time, high-concurrency data analysis.

Supported versions​

StarRocks 3.3+

Limitations​

Currently, only the following configuration is supported:

  • Writing from AE Built-in Warehouse workspace tables to StarRocks (offline write)

Note:

  • Reading from StarRocks data sources (offline read) isn't supported yet

Supported field types​

The StarRocks engine determines whether a source field can be written to the target field correctly. If it can't, the value is Null.

Integration plans in the AE DataOps Platform don't force data type conversion; the engine handles it.

StarRocks typeOffline write (Writer)Remarks
Numeric types
TINYINTSupported1-byte signed integer, range [-128, 127]
SMALLINTSupported2-byte signed integer, range [-32768, 32767]
INTSupported4-byte signed integer, range [-2147483648, 2147483647]
BIGINTSupported8-byte signed integer, range [-9223372036854775808, 9223372036854775807]

DECIMAL

Supported

DECIMAL(P [, S])
High-precision fixed-point number. P is the total number of significant digits (precision), and S is the maximum number of digits after the decimal point (scale).
In version 1.19.0 and later, the (P, S) of the decimal type has a default value of decimal(10, 0)
DOUBLESupported8-byte floating-point number
BOOLEANSupportedBOOL, BOOLEAN
Same as TINYINT: 0 means false and 1 means true
LARGEINTSupported16-byte signed integer, range [-2^127 + 1 ~ 2^127 - 1]
FLOATSupported4-byte floating-point number.
String types
CHARSupportedCHAR(M)
Fixed-length string. M is the length of the fixed-length string, and its range is 1~255.

VARCHAR

SupportedVARCHAR(M)
Variable-length string. M is the length of the variable-length string, in bytes. The default value is 1.
STRINGSupportedString with a maximum length of 65533 bytes

BINARY/VARBINARY

Not supported yet

Starting from version 3.0, StarRocks supports the BINARY(M) / VARBINARY(M) data types for storing binary data, in bytes.

The maximum supported length is the same as that of the VARCHAR type, and the value range of M is [1,1048576]. If M isn't specified, it defaults to the maximum value 1048576.

BINARY is an alias of VARBINARY and is used in the same way as VARBINARY.

Time types
DATESupportedDate type. The current value range is ['0000-01-01', '9999-12-31'].
The default print format is 'YYYY-MM-DD'.
DATETIMESupportedDatetime type. The value range is ['0000-01-01 00:00:00', '9999-12-31 23:59:59'].
The print format is 'YYYY-MM-DD HH: MM: SS'
Semi-structured types
JSONSupported
ARRAYSupported
MAPSupported
STRUCTSupported
Aggregation types
HLLNot supported yetHyperLogLog type, used for approximate deduplication.
BITMAPNot supported yetSimilar to HLL (HyperLogLog), BITMAP is commonly used to speed up count distinct deduplication.

Table types​

  • Duplicate key table is simple and easy to use. The data in the table has no constraints, and identical rows can exist repeatedly. It's suitable for storing raw data that needs no constraints or pre-aggregation, such as logs.
  • Primary Key table is powerful and has uniqueness and non-null constraints. It supports scenarios such as real-time updates and partial column updates while ensuring query performance, which makes it suitable for real-time queries.
  • Aggregate table is suitable for storing pre-aggregated data. It reduces the amount of data that aggregate queries need to scan and compute, greatly improving the efficiency of aggregate queries.
  • Unique Key table is suitable for real-time update scenarios and is gradually being replaced by the Primary Key table.
Primary Key tableDuplicate Key tableAggregate tableUnique Key table
Unique constraintThe primary key PRIMARY KEY has unique and non-null constraints.DUPLICATE KEY has no unique constraint.The aggregate key AGGREGATE KEY has a unique constraint.The unique key UNIQUE KEY has a unique constraint.
Logical relationship of key columnsIf the primary key value of new data is the same as that of existing data in the table, a unique constraint conflict occurs, and the new data replaces the existing data.
Compared with the Unique Key table, the Primary Key table has an enhanced underlying storage engine and can now replace the Unique Key table.
Duplicate Key has no unique constraint, so if the Duplicate Key of new data is the same as that of existing data in the table, both the new and old data are kept in the table.If new data has a unique constraint conflict with existing data in the table, the new and old data are aggregated based on the aggregate key and the aggregate functions of the Value columns.If new data has a unique constraint conflict with existing data in the table, the new data replaces the existing data.
A Unique Key table can actually be regarded as an Aggregate table whose aggregate function is replace.

Data types supported

by sort keys

Numeric (including integer and boolean), string, and date/time.Numeric (including integer, boolean, and Decimal), string, and date/time.
Partition/bucket columns

Partition columns and bucket columns must be in the primary key.

None

Partition columns and bucket columns must be in the aggregate key.

Partition columns and bucket columns must be in the unique key.

Aggregation types​

Aggregation typePurposeApplicable column typesTypical scenarios
SUMSums the values for the same dimension columnsNumeric types (INT/BIGINT/DECIMAL)Cumulative metrics such as sales and visits
MINTakes the minimum value for the same dimension columnsNumeric types/date typesLowest values, earliest login time
MAXTakes the maximum value for the same dimension columnsNumeric types/date typesHighest values, latest login time
REPLACEData written later completely overwrites the previous value (whether or not it's NULL)Any typeA user's latest address, an order's final status

REPLACE_IF_NOT_NULL

Only non-NULL values overwrite the previous value (NULL values keep the old value)

Requires the field's default value to be NULL (if the default value is an empty string '', it's treated as a valid value and overwrites the previous value).

Any type

Incremental updates to user information (existing fields are kept)

Write mode​

In AE integration plans, data is written to StarRocks using the Insert +Files write mode

Supported source data source types
  • Data sources:
    • HDFS
    • S3

Supported file types

  • Parquet
  • ORC

CSV/Avro/Json/Text aren't supported yet

Syntax example
INSERT INTO user_behavior_declared
SELECT * FROM FILES
(
"path" = "hdfs://<hdfs_ip>:<hdfs_port>/user/amber/user_behavior_ten_million_rows.parquet",
"format" = "parquet",
"hadoop.security.authentication" = "simple",
"username" = "<hdfs_username>",
"password" = "<hdfs_password>"
);

Create a StarRocks data source​

In the DataOps Platform - Integration module, you can choose to add a StarRocks data source.

Data source configuration parameters

Fill in the configuration required by the data source and pass the connectivity test to create the StarRocks data source.

Field nameDescription
Basic Information
*Datasource NameMust be unique within the DataOps Platform space. Can contain only letters, digits, and underscores, and can't start with a digit or an underscore
RemarksOptional
Data source configuration
not distinguish environment / Independent settingChoose one: not distinguish environment means the production and development environments share one configuration; Independent setting means the two environments are configured independently
*Server Address/IPAddress of the server where the StarRocks database runs. Separate multiple addresses with commas
*PortPort used to access StarRocks
*DatabaseName of a database already created in StarRocks
*UsernameUsername with permission to access the database
*PasswordPassword of the username
AdvancedOther advanced parameters required to connect to the database; you can customize them
Note: Cluster deployment mode by default

Parameters whose names start with * are required; parameters without * are optional.

Create an offline sync task​

After you create the StarRocks data source and pass the connectivity test as described above, you can configure a StarRocks offline write task for your scenario.

StarRocks as the Data Target​

Select StarRocks as the Data Target and configure the following parameters:

Field nameDescription
*Data source (type)Select StarRocks as the target type of the Data Target. This drop-down lists only the types of data sources that already exist in the current space. If you haven't created a StarRocks data source yet, StarRocks isn't listed. Create the data source first through + Data Sources at the bottom of the drop-down list or Data sources management
*Data source (data source)A StarRocks data source registered on the data source management page; select it from the drop-down list.
If you haven't created the data source yet, click the Data sources management button to create a StarRocks data source.
*Fully qualified name (catalog)default_catalog (that is, the internal catalog of StarRocks). External catalogs aren't supported yet
*Fully qualified name (database)Name of the database to write to
*Target TableThe table to write to; select it from the drop-down list. There's a Create Table shortcut next to the drop-down
*Partition Field Value

You can define the partition field value through the Input Method. Suppose the partition field is days:

  • Constant is a custom value: days = 2025-01-01
  • Parameter is a space parameter: days =${bd}, which takes the value of the base date at run time
  • Field is a source table field: days = "dt1"
*Write Mode
  • Insert into: insert new data
  • Overwrite: overwrite the entire table

Additional notes:

Write modePrimary Key tableDuplicate Key tableAggregate tableUnique Key table
Insert into
  • Supported

The actual operation is Upsert

  • Supported

Update isn't supported

  • Supported
  • Supported

Field mapping​

After configuring the data source and the target, create field mappings. The system automatically syncs data from source fields to target fields based on the mappings. You can configure field mappings in three ways:

  • Method 1: Custom selection. Select a source table field, then select the target field in the target table
  • Method 2: Name Mapping. The system automatically maps fields with the same name in the source and target tables
  • Method 3: Line Mapping. The system automatically maps fields in the same row

Note that each target field can correspond to only one source field

Basic information settings​

Finally, set the basic information of the integration plan, including the plan name, owner, synchronization rate, and remarks.

When you're done, click Save to create the integration plan.

Note: The plan name can't be changed after it's saved

Mount an offline sync plan on a Flow​

Mount on Flow​

1. Start mounting

  • On the integration plan details page, click Mount on task flow in the upper-right corner

2. Select a Flow

  • Select the target Flow from the drop-down menu

  • To create a new Flow:

    • Click the New Flow shortcut button below the drop-down menu
    • Or go to the Dev module to create one
  • Tip: If the target Flow isn't shown, click the ↻ refresh button on the right

3. Create a sync node

  • In the Flow, create a node of the Offline sync plan type
  • The node is linked to the current integration plan. When the node runs, it triggers a run of that integration plan

4. Complete mounting

  • Click Create node and mount on it to complete the configuration
  • After mounting succeeds, click Go to flow page to view the result right away
tip

Note: The task node created by mounting is in the unreleased state. We recommend going to the Flow and releasing the node.

Unmount from a Flow​

To unmount an offline sync plan from a Flow:

  • If the Flow hasn't been released yet, go to the Flow that the plan is mounted on and delete the Offline sync plan task node in Dev Mode.
  • If the Flow has already been released, after deleting the Offline sync plan task node in Dev Mode, release the Flow again. This also unmounts the node from the Production environment.
Was this page helpful?