Skip to main content

Configure a MongoDB data source

Last updated 10/03/2026

Offline Sync plans in the AE DataOps Platform support writing (Writer) to MongoDB data sources. This article describes the data synchronization capabilities supported for MongoDB.

MongoDB is a document database (with JSON as its data model).

Among non-relational databases, MongoDB is the most feature-rich and the most similar to relational databases. It supports BSON, a JSON-like binary storage format short for Binary JSON. Like JSON, it supports embedded document objects and array objects, so it can store relatively complex data types.

SQL conceptMongoDB conceptDescription
DatabaseDatabaseA logical namespace. A database contains multiple collections with different names
TableCollectionEquivalent to a table in SQL. A collection can store multiple documents with different fields (types)
RowDocumentA document is equivalent to a row in a table and consists of multiple different fields
ColumnFieldAn attribute in a document, equivalent to a column
IndexIndexSame as the SQL concept
Primary key_id (field)Every document has a unique id field
ViewViewA virtual (not physically existing) collection
Table joinsAggregation operation ($lookup)An aggregation operator "similar to" a table join

MongoDB's most notable feature is its very powerful query language. Its syntax is somewhat like an object-oriented query language, and it can implement almost all the functionality of single-table queries in relational databases. It also supports indexing data.

Gaming scenario: Use MongoDB to store game player information. Players' equipment, points, and so on are stored directly as embedded documents, which makes them easy to query and update.

Supported versions​

MongoDB 4.0+

Limitations​

Currently, only the following configuration is supported:

  • Writing from AE Built-in Warehouse workspace tables to MongoDB (offline write)

Note:

  • Reading from MongoDB data sources (offline read) isn't supported yet

Supported field types​

Field typeOffline write (MongoDB Writer)Description
INTSupported32-bit integer
LONGSupported64-bit integer
STRINGSupported
OBJECTIDSupported
DECIMALSupportedDecimal128, up to 34 digits of precision
NULLSupported
DOUBLESupported
DATESupportedSecond-level precision
TIMESTAMPSupportedMillisecond-level precision
BOOLEANSupported
BINDATANot supported
REGEXNot supported
JAVASCRIPTNot supported
MINKEYNot supported
MAXKEYNot supported

Create a MongoDB data source​

In the DataOps Platform - Integration module, you can choose to add a MongoDB data source.

Data source configuration parameters

Fill in the configuration required by the data source and pass the connectivity test to create the MongoDB data source.

Parameters whose names start with * are required; parameters without * are optional.

Field nameDescription
Basic Information
*Datasource NameMust be unique within the DataOps Platform space. Can contain only letters, digits, and underscores, and can't start with a digit or an underscore
RemarksOptional
Data source configuration
not distinguish environment / Independent settingChoose one: not distinguish environment means the production and development environments share one configuration; Independent setting means the two environments are configured independently
*Mode

The following deployment modes are supported:

  • Cluster mode
  • Replica cluster
  • Single node mode
*Server Address/IP

Address of the server where the database runs

For cluster deployment, click Add to enter information for multiple nodes

*PortPort used to access MongoDB
*DatabaseName of the MongoDB database to connect to
*UsernameUsername with permission to access the database
*PasswordPassword of the username
Auth DatabaseName of the database used when the user was created, such as admin
replicaSetWhen the deployment mode is Replica cluster, enter the replica set name
Advanced

Other advanced parameters required to connect to the database; you can customize them

For example, to connect with an SSL certificate, use the ssl=true parameter

Note: The parameters above are combined into a connection string, which is used to establish the connection to the MongoDB database

For example: mongodb://<username>:<password>@<instance_ip>:<instance_port>/<database_name>?authSource=admin&ssl=true

Create an offline sync task​

After you create the MongoDB data source and pass the connectivity test as described above, you can configure a MongoDB offline write sync task for your scenario.

MongoDB as the Data Target​

Select MongoDB as the Data Target and configure the following parameters:

Field nameDescription
*Source TypeSelect MongoDB as the target type of the Data Target
*Datasource NameA MongoDB data source registered on the data source management page; select it from the drop-down list.
If you haven't created the data source yet, click the Data sources management button to create a MongoDB data source.
*DatabaseName of the database to read
*Collection name

Select the collection name from the drop-down list

Note: If the collection doesn't exist yet, the insert operation creates the collection.

*Write Mode
  • Insert new data: Insert (the task fails if a unique key conflict occurs)
  • Upsert: If the unique key is the same, the new data overwrites the old data; if the unique key is different, the new data is inserted directly. Specify the unique key with the As UPSERT Key switch in the Add Collection Field dialog of field mapping

Note: A document involved in an insert/update command can't exceed 16MB.

Pre statement

Purpose: Performs a pre-operation on the existing data in MongoDB before data is written to MongoDB

Make sure the pre statement you enter follows JSON syntax.

The pre-condition (PreSQL) configuration contains two properties: type and json.

  • type: Required. Valid values: remove, drop
  • json:
    • When type is remove, it's required, and its syntax is a standard MongoDB Query.
    • When type is drop, you don't need to fill it in.
{
"type":"remove",
"item": [
{
"_id": ObjectId("..."),
"student_id": 1001,
"name": "xiaoming",
"grade": "grade_1"
}
]
}
AdvancedBatch size: Amount of data written to MongoDB in a single batch
Batch Write Interval: Maximum time interval for flushing data from the memory buffer to MongoDB, in milliseconds (ms)

Parameters whose names start with * are required; parameters without * are optional.

Field mapping​

After configuring the data source and the target, create field mappings. Based on the mappings, the system automatically creates the fields in the target collection from the data in the source fields.

  • After parsing succeeds, click Generate Name Mapping to generate the Mongo Document structure based on the data types of the source table
  • You can further edit the Collection Field Name and Collection Field Type of the Document to be written.

Note that each target field can correspond to only one source field

Field type mapping (workspace table to Mongo):

Source file typeWorkspace tableMongoDB BSON typeDescription

orc、parquet

tinyintInt
smallintInt
integerInt32-bit integer
bigintLong64-bit integer
varcharString or ObjectID
booleanBoolean
doubleDouble
dateDate
timestampTimestampMillisecond-level precision
decimalDecimal(10)Decimal128
arrayArrayArray type
rowObject
mapObject
csv、textvarcharString or ObjectID

Preview results​

Click Preview Results to sample the source data and generate three sample write results.

Note ⚠️: If a source field is Null, the field isn't shown in the write preview, but it's written to Mongo as "key" :null.

Basic information settings​

Finally, set the basic information of the integration plan, including the plan name, owner, synchronization rate, and remarks.

When you're done, click Save to create the integration plan.

Note: The plan name can't be changed after it's saved

The details page looks like this:

Mount an offline sync plan on a Flow​

Mount on Flow​

1. Start mounting

  • On the integration plan details page, click Mount on task flow in the upper-right corner

2. Select a Flow

  • Select the target Flow from the drop-down menu

  • To create a new Flow:

    • Click the New Flow shortcut button below the drop-down menu
    • Or go to the Dev module to create one
  • Tip: If the target Flow isn't shown, click the ↻ refresh button on the right

3. Create a sync node

  • In the Flow, create a node of the Offline sync plan type
  • The node is linked to the current integration plan. When the node runs, it triggers a run of that integration plan

4. Complete mounting

  • Click Create node and mount on it to complete the configuration
  • After mounting succeeds, click Go to flow page to view the result right away
tip

Note: The task node created by mounting is in the unreleased state. We recommend going to the Flow and releasing the node.

Unmount from a Flow​

To unmount an offline sync plan from a Flow:

  • If the Flow hasn't been released yet, go to the Flow that the plan is mounted on and delete the Offline sync plan task node in Dev Mode.
  • If the Flow has already been released, after deleting the Offline sync plan task node in Dev Mode, release the Flow again. This also unmounts the node from the Production environment.
Was this page helpful?