Skip to content
R Roesli.
Go back
ai-agents

The agent runbook for native Fabric dbt, Airflow, Lakehouse, and Direct Lake

An evidence-driven agent runbook for deploying and validating Fabric dbt Jobs, Lakehouse, Airflow, workspace identity, Direct Lake, and DAX.

The companion walkthrough explains how to build Jaffle Shop as a native Fabric solution. This runbook answers a harder question: what must an engineering agent know to recreate it without confusing “item exists” with “the system works”?

The agent needs the item model, adapter boundary, identity chain, definition format, long-running operations, and a validation contract. The target path is:

Jaffle Shop seeds
       |
       v
Fabric dbt Job
dbt Core + dbt-fabricspark
       |
       v
Fabric Lakehouse / OneLake
raw -> staging -> gold
       |
       +-----------------------+
       |                       |
       v                       v
Fabric Apache Airflow     Direct Lake semantic model
workspace identity        relationships + DAX

The tested run used a Fabric trial workspace. The Airflow DAG started the dbt Job, the dbt build completed, Gold Delta tables were populated, and DAX queries returned the expected results.

Handoff before the agent starts

Skills for Microsoft Fabric are an agent prerequisite, not a runtime prerequisite. They provide workflow instructions and tool routing. They do not provide capacity, credentials, tenant permissions, or workspace access.

The machine needs:

Optional integrations include Power BI Modeling MCP, FabricIQ, and a Fabric SQL endpoint tool or ODBC client. If Power BI Modeling MCP is used, accept its EULA.

Authenticate with:

az login

Fabric control-plane operations use:

https://api.fabric.microsoft.com

Do not reuse that token audience for managed Airflow or Power BI without checking the service. Use delegated sign-in, workspace identity, managed identity, or an approved service principal. Never pass a password to the agent.

Fabric preparation

A Fabric administrator must confirm:

The deployment identity needs Contributor or higher access. Creating a workspace identity, changing tenant settings, and assigning capacity may need administrator rights.

Before the DAG can call dbt:

  1. enable workspace identity;
  2. add it to the workspace as Contributor or Member;
  3. create a connection using Workspace identity;
  4. enable Allow Code-First Artifacts;
  5. enable Fabric Connections on Airflow;
  6. attach the connection; and
  7. record its GUID.

Model the build as dependencies

Capacity
  -> Workspace
      -> Workspace identity
      -> Lakehouse
          -> dbt Job
              -> Gold Delta tables
                  -> Direct Lake semantic model
          -> Fabric connection
              -> Airflow item
                  -> Airflow DAG
                      -> dbt Job execution

Keep a manifest with:

Logical nameFabric typeIdentifier
WorkspaceWorkspaceWorkspace ID
LakehouseLakehouseItem ID and SQL endpoint ID
dbt JobDataBuildToolJobItem ID
Airflow JobApacheAirflowJobItem ID and web URL
Semantic modelSemanticModelItem ID
Workspace identityService principalObject and application ID
Fabric connectionConnectionConnection GUID

Names are for people. APIs use GUIDs. List resources and match exact names; never copy identifiers from another environment.

Know the runtime

TargetAdapterSQL
Fabric Data Warehousedbt-fabric 1.10.0T-SQL
Fabric Lakehousedbt-fabricspark 1.12.2Spark SQL

Fabric dbt Job runtime V1.0 supplies dbt Core 1.11, Python 3.12, and dbt-fabricspark 1.12.2 for a Lakehouse profile. The Lakehouse SQL analytics endpoint is read-only. An agent should reject a plan that points dbt-fabric at that endpoint and tries to write models.

Phase 0: preflight

Verify:

  1. authentication for https://api.fabric.microsoft.com;
  2. a capacity supporting dbt and Airflow;
  3. dbt Jobs enabled;
  4. Contributor or higher access;
  5. workspace identity and code-first connections; and
  6. exact item names.

Discover workspaces:

GET https://api.fabric.microsoft.com/v1/workspaces

List items in the selected workspace and filter by exact type and name.

Phase 1: create the Lakehouse

Create:

jaffle_shop_lakehouse

Use logical schemas:

raw
staging
gold

Record both the Lakehouse item ID and SQL endpoint ID. They are different resources.

Phase 2: deploy the dbt Job

Create:

jaffle_shop_native_dbt

A Lakehouse definition has this conceptual shape:

{
  "project": {
    "projectType": "OneLake",
    "folderPath": "dbt"
  },
  "profile": {
    "profileType": "Lakehouse",
    "schema": "gold",
    "connectionSettings": {
      "name": "jaffle_shop_lakehouse",
      "properties": {
        "type": "Lakehouse",
        "typeProperties": {
          "workspaceId": "<workspace-id>",
          "artifactId": "<lakehouse-id>",
          "endPoint": "<lakehouse-sql-endpoint>"
        }
      }
    }
  },
  "command": {
    "operation": "build",
    "arguments": {
      "failFast": true,
      "threads": 4
    }
  }
}

Project configuration:

name: jaffle_shop
version: 1.0.0
config-version: 2
profile: jaffle_shop

model-paths: ["models"]
seed-paths: ["seeds"]
macro-paths: ["macros"]

models:
  jaffle_shop:
    staging:
      +schema: staging
      +materialized: view
    marts:
      +schema: gold
      +materialized: table

seeds:
  jaffle_shop:
    +schema: raw

Use data_tests for primary-key uniqueness, non-nullability, customer relationships, and date relationships.

Fabric definitions are multipart. The safe update loop is:

  1. export the live definition;
  2. decode every part;
  3. change the intended files;
  4. re-encode all required parts;
  5. submit one updateDefinition; and
  6. export again and compare.

Start the job:

POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances?jobType=Execute

Poll:

GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances

202 Accepted or item creation is not success. Follow the specific instance to Completed or Failed.

Phase 3: validate data before BI

select count(*) as order_count
from gold.fct_orders;
select
    sum(amount) as total_revenue,
    min(order_date) as first_order_date,
    max(order_date) as last_order_date
from gold.fct_orders;
select count(*) as orphan_customers
from gold.fct_orders f
left join gold.dim_customers c
    on f.customer_id = c.customer_id
where c.customer_id is null;

The validated test oracle was:

CheckExpected
Orders12
Revenue175.00
Customers10
Dates11
Orphan customer keys0
Orphan date keys0

Phase 4: build Direct Lake

Create:

jaffle_shop_direct_lake

Use a star schema:

dim_customers  1 ---- *  fct_orders
dim_dates      1 ---- *  fct_orders

Relationships:

FromToDirection
fct_orders[customer_id]dim_customers[customer_id]Single
fct_orders[order_date]dim_dates[date_day]Single

Measures:

Total Orders =
DISTINCTCOUNT(fct_orders[order_id])
Total Revenue =
SUM(fct_orders[amount])
Average Order Value =
DIVIDE([Total Revenue], [Total Orders])

Point the model at Gold Lakehouse entities, not the SQL endpoint as DirectQuery. After deployment, confirm the model, frame it or refresh as appropriate, and run:

EVALUATE
ROW(
    "Orders", [Total Orders],
    "Revenue", [Total Revenue],
    "Average Order Value", [Average Order Value],
    "Customers", COUNTROWS('dim_customers'),
    "Dates", COUNTROWS('dim_dates')
)

Then test date slicing:

EVALUATE
SUMMARIZECOLUMNS(
    dim_dates[calendar_year],
    dim_dates[month_number],
    "Orders", [Total Orders],
    "Revenue", [Total Revenue]
)
ORDER BY
    dim_dates[calendar_year],
    dim_dates[month_number]

Phase 5: configure workspace identity

The identity chain is:

Workspace identity
  -> Contributor or Member
  -> WorkspaceIdentity Fabric connection
  -> Allow code-first artifacts
  -> Connection attached to Airflow
  -> Connection GUID in DAG

Create the identity, grant workspace access, create the connection, enable code-first artifacts, and record the GUID. The display name is not enough.

Phase 6: deploy Airflow

Create:

jaffle_shop_airflow_orchestrator

Definition shape:

{
  "properties": {
    "type": "Airflow",
    "typeProperties": {
      "airflowProperties": {
        "airflowEnvironment": "FabricAirflowJob-1.0.0",
        "airflowVersion": "2.10.5",
        "pythonVersion": "3.12",
        "enableAADIntegration": true,
        "enableFabricConnections": true,
        "enableTriggerers": true
      },
      "fabricConnections": ["<fabric-connection-guid>"]
    }
  }
}

DAG:

from datetime import datetime, timedelta

from airflow import DAG
from airflow.providers.microsoft.fabric.operators.run_item import (
    MSFabricRunJobOperator,
)

FABRIC_CONN_ID = "<fabric-connection-guid>"
WORKSPACE_ID = "<workspace-id>"
DBT_JOB_ID = "<dbt-job-id>"

with DAG(
    dag_id="orchestrate_jaffle_shop_dbt",
    schedule=None,
    start_date=datetime(2026, 1, 1),
    catchup=False,
    default_args={
        "owner": "fabric",
        "retries": 1,
        "retry_delay": timedelta(minutes=5),
    },
) as dag:
    run_jaffle_shop_dbt = MSFabricRunJobOperator(
        task_id="run_jaffle_shop_dbt",
        fabric_conn_id=FABRIC_CONN_ID,
        workspace_id=WORKSPACE_ID,
        item_id=DBT_JOB_ID,
        job_type="DataBuildToolJob",
        wait_for_termination=True,
        deferrable=True,
    )

fabric_conn_id is the GUID. job_type="Execute" is wrong; use DataBuildToolJob or DBT.

Phase 7: start Airflow

POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment/start?beta=true

Status:

GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment?beta=true

Expected shape:

{
  "status": "Started",
  "airflowWebUrl": "https://<managed-airflow-host>/login/"
}

Starting an already started environment is an idempotency signal. Follow the status instead of treating it as a failure.

At the tested CLI version, fab job start did not support ApacheAirflowJob. Use the documented REST environment API or portal.

Phase 8: run the DAG

The documented route is the Airflow UI:

  1. open the Airflow item;
  2. enable Fabric Connections;
  3. attach the workspace-identity connection;
  4. open or create the DAG; and
  5. select Run DAG.

The managed endpoint also exposes:

POST {airflowWebUrl}/api/v1/dags/orchestrate_jaffle_shop_dbt/dagRuns
Content-Type: application/json

{
  "conf": {},
  "note": "Triggered by deployment agent"
}

The Airflow endpoint can require a different Entra audience. A Fabric control-plane token may redirect to interactive sign-in. Check the challenge before requesting a token.

Phase 9: prove the chain

Airflow DAG run
  -> Airflow task instance
      -> Fabric dbt Job instance
          -> Gold Lakehouse output
              -> Direct Lake DAX result

Capture the DAG ID and state, task state and try, dbt instance and state, dbt counts, Lakehouse checks, and DAX results. The validated deployment returned:

LayerEvidence
AirflowDAG success
Airflow taskMSFabricRunJobOperator success
dbtPASS=23 WARN=0 ERROR=0 SKIP=0
Lakehouse12 orders and 175.00 revenue
Semantic model12 orders, 175.00 revenue, 14.5833 average order value

Failure atlas

SymptomLikely causeAction
InvalidDefinitionFormatUnsupported format or malformed definitionExport without guessing and use the documented envelope
Invalid LinkedServiceWarehouse source used for LakehouseUse the Lakehouse profile
Project directory missingfolderPath mismatchAlign folder and uploaded parts
dbt seed-key failureSource differs or test assumptions wrongExport live files and compare
conn_id ... isn't definedDisplay name usedUse connection GUID
Airflow operator rejects run typejob_type="Execute"Use DataBuildToolJob or DBT
UserAccessTokenExceptionToken acquisition failedVerify identity and connection
Airflow redirects to sign-inWrong audienceUse the managed Airflow audience
Environment says StartedStart called twiceValidate status and continue
fab job start unsupportedCLI gapUse REST or portal
Duplicate semantic modelsBlind retryList duplicates and deploy once
Unexpected DAX totalsContract or relationship errorRevalidate rows, relationships, and measures

What the agent must never do

  1. Commit tokens, secrets, or credentials.
  2. Use a display name where a GUID is required.
  3. Invent an endpoint after a 404.
  4. Infer success from item existence.
  5. Stop at dbt when the requested outcome includes Airflow and BI.
  6. Deploy a partial semantic-model definition.
  7. Retry model creation blindly.
  8. Treat Warehouse and Lakehouse adapters as interchangeable.
  9. Trust a local project after live deployment without exporting and reconciling.
  10. Call an untested path working.

Final lesson

The hard part is not generating SQL, a DAG, or a DAX measure. It is preserving the contracts between workloads:

Once those boundaries are explicit, the native stack is repeatable. Without them, an agent can leave behind a workspace that looks complete and still fails at the first real orchestration run.

References


Share this post:

Continue exploring

Previous Post
Build your first planning solution with Plan in Microsoft Fabric
Next Post
A native Fabric analytics stack with dbt, Airflow, Lakehouse, and Direct Lake
Community

Join the conversation

Sign in with GitHub to leave a comment.

GitHub

Loading comments…

Sign in with GitHub to comment