Workbench
The Dashboard is an Agent-driven collaboration space for DataOps tasks. Around a clear task goal, the Agent proactively plans the handling steps, calls tools to collect information, and keeps moving forward based on the execution results. You can also put related tasks in the same workspace to follow up on and review them together.
The Dashboard is currently in the trial stage. It first opens up the ops troubleshooting scenario: starting from a failed data task, it automatically collects evidence, locates the cause, and gives handling suggestions.
1. Understand the Dashboard, workspaces, and troubleshooting tasks
The trial version currently supports troubleshooting two types of alerts: Flow run failed and Task run failed. Each failure corresponds to one troubleshooting task, which continuously records the original alert, troubleshooting plan, execution process, evidence, conclusion, and handling suggestions.
Three core objects
| Object | Description | How to use it |
|---|---|---|
| Abnormal event | Generated by failure alert rules; the starting point of troubleshooting | View the failed object, time of occurrence, and alert context |
| Troubleshooting task | One complete diagnosis run for a single abnormal event | View the plan, steps, evidence, conclusion, and suggestions |
| Workspace | A container that organizes related troubleshooting tasks; not the same as a DataOps space | Summarize the progress of the same issue or business pipeline and generate a workspace briefing |
How to divide workspaces
Put related tasks for the same issue, the same business pipeline, or anything that needs continuous follow-up in the same workspace, so you can use the briefing to spot common causes and overall progress. Create separate workspaces for unrelated issues. After an issue is resolved, you can archive the workspace so that historical tasks don't interfere with current troubleshooting.
2. Start an ops troubleshooting task
Before you start
Make sure at least one failure alert rule is enabled in the space. The Dashboard only shows anomalies found by alert rules. Without alert rules, there are no abnormal events to troubleshoot.
Select an anomaly and a workspace
When you create a troubleshooting task, select the anomaly to troubleshoot and decide whether to put it in an Existing workspace or a New workspace. If the same issue already has a troubleshooting task, add it to that workspace first. If the issue is independent, create a new workspace.
Add diagnostic context. If you've recently changed code, dependencies, or configuration, or have a clear expectation of the normal result, describe it in Additional notes. The Dashboard includes this information in the troubleshooting instead of judging only by the alert text.
View details. Before creating the task, check the failed object, alert time, and related information to avoid troubleshooting the wrong anomaly.
Confirm the troubleshooting plan
A manually created task first generates a troubleshooting plan. Focus on whether the troubleshooting goal, scope, and known context are accurate. If anything is missing, add background or adjust the goal before you confirm execution. After the plan is confirmed, the Dashboard starts querying evidence and moving the diagnosis forward.
3. How the Dashboard completes troubleshooting
Troubleshooting mechanism
- Generates a troubleshooting plan based on the abnormal event and the additional information.
- Queries task instances, run logs, metrics, change records, and other information according to the plan.
- Reviews the existing evidence after each step. When clues are insufficient, adds troubleshooting steps or requests information from the relevant people.
- When the evidence is sufficient, compiles the failure cause, scope of impact, and handling suggestions.
Each node in the troubleshooting graph corresponds to a diagnosis step. Select a node to view the query purpose, execution result, and evidence source, so you can judge how the conclusion was reached.
How to collaborate on troubleshooting
Add information based on evidence. When you receive a "needs more information" reminder, provide the business background, expected result, recent changes, or known scope of impact. The troubleshooting continues based on the new information.
Go back to the original object to check. When you need to check the full logs or task configuration, you can jump from the troubleshooting task to the related task instance or asset.
Handling boundaries
The Dashboard diagnoses issues and gives suggestions, but it doesn't run fixes automatically. Operations such as rerunning tasks, modifying code, adjusting dependencies, or changing configuration still need to be confirmed and completed by people with the required permissions.
4. Use workspace briefings to track overall progress
The workspace briefing aggregates information from the unarchived tasks in the current workspace. It helps you quickly tell whether an issue already has a conclusion, what needs to be handled next, and whether multiple troubleshooting tasks point to a common cause.
| Area | Question it answers | Suggested action |
|---|---|---|
| Identified | Which issues already have a clear cause or conclusion | Check the conclusion and confirm whether handling is complete |
| Handling suggestions | Which follow-up actions can be taken now | Assign an owner and open the corresponding task to view the basis |
| Key points | Which clues, symptoms, or patterns are worth watching | Judge whether multiple anomalies share a common cause |
After tasks are added, change status, or finish troubleshooting, you can update the briefing to get the latest summary. The briefing only counts unarchived tasks. "Handling suggestions" focuses on suggestions from recently created tasks.
Go from the briefing back to tasks
The briefing is an overview and doesn't replace the full evidence of individual tasks. When you need to confirm the basis of a suggestion, use the related task count or the corresponding entry to open the specific troubleshooting task and view its execution process, evidence, and context.
5. Recommended: turn on automatic creation
When to turn it on
For production environments that need continuous operations and should go into diagnosis right after a failure, we recommend turning on automatic creation. It reduces the time spent waiting for someone to notice anomalies and create tasks repeatedly.
| Setting | Purpose | Recommendation |
|---|---|---|
| Automatically create tasks | Automatically creates a troubleshooting task after a new failure alert occurs | Recommended for production environments |
| Auto-Confirm Plan | Starts executing the diagnosis plan directly after the task is created | Reduces waiting, but doesn't run fixes automatically |
| Workspace policy | The system creates workspaces automatically, or tasks always join a specified workspace | Specify a fixed workspace when you're continuously focusing on the same business |
Effective rules
After it's turned on, only new failure alerts generated afterward that haven't been handled yet are processed. No tasks are backfilled for historical alerts. Automatically created tasks are assigned to the asset owner first. If the specified workspace has been archived, it's restored and used. If it has been deleted, the system falls back to system-created workspaces.
We recommend turning on both automatic creation and external notifications: the system continuously finds and diagnoses issues, and the relevant people get timely reminders when they need to add information or can take action.
6. Handle tasks promptly through @Me and external notifications
@Me collects the troubleshooting reminders you need to follow up on. It distinguishes pending and read status, and you can open the corresponding task directly.
| Reminder type | When it appears | What to do |
|---|---|---|
| Needs more information | The existing evidence is insufficient and troubleshooting needs more business context | Add the expected result, recent changes, or scope of impact |
| Handling suggestions ready | Troubleshooting already has actionable suggestions | Review the basis of the suggestions and arrange follow-up handling |
When a manually created task produces handling suggestions, the task creator is reminded by default. For automatically created tasks, the asset owner is reminded first.
Get reminders even when you're away from the Dashboard
To receive messages through external channels such as Feishu, a space admin needs to enable the Workbench task diagnosis alerts notification rule, allow personal subscriptions, and bind a notification channel to the DataOps space. Users then turn on personal subscriptions as needed.
If you don't receive external reminders
- Confirm that the Workbench task diagnosis alerts rule is enabled in the space.
- Confirm that the admin allows personal subscriptions and that you've turned on your subscription.
- Confirm that an available notification channel is bound to the current DataOps space.
If you don't have management permissions, contact a space admin to complete the configuration above.

