The companion walkthrough explains how to build Jaffle Shop as a native Fabric solution. This runbook answers a harder question: what must an engineering agent know to recreate it without confusing “item exists” with “the system works”?
The agent needs the item model, adapter boundary, identity chain, definition format, long-running operations, and a validation contract. The target path is:
Jaffle Shop seeds
|
v
Fabric dbt Job
dbt Core + dbt-fabricspark
|
v
Fabric Lakehouse / OneLake
raw -> staging -> gold
|
+-----------------------+
| |
v v
Fabric Apache Airflow Direct Lake semantic model
workspace identity relationships + DAX
The tested run used a Fabric trial workspace. The Airflow DAG started the dbt Job, the dbt build completed, Gold Delta tables were populated, and DAX queries returned the expected results.
Handoff before the agent starts
Skills for Microsoft Fabric are an agent prerequisite, not a runtime prerequisite. They provide workflow instructions and tool routing. They do not provide capacity, credentials, tenant permissions, or workspace access.
The machine needs:
- GitHub Copilot CLI or another skill-capable agent host;
- Skills for Fabric installed and enabled;
- Azure CLI;
- Git and a writable project directory; and
- network access to Fabric, Entra ID, OneLake, and managed Airflow.
Optional integrations include Power BI Modeling MCP, FabricIQ, and a Fabric SQL endpoint tool or ODBC client. If Power BI Modeling MCP is used, accept its EULA.
Authenticate with:
az login
Fabric control-plane operations use:
https://api.fabric.microsoft.com
Do not reuse that token audience for managed Airflow or Power BI without checking the service. Use delegated sign-in, workspace identity, managed identity, or an approved service principal. Never pass a password to the agent.
Fabric preparation
A Fabric administrator must confirm:
- capacity or trial availability;
- dbt Jobs (preview);
- Apache Airflow Jobs;
- service-principal access where workspace identity is used; and
- required code-first and workspace-identity preview switches.
The deployment identity needs Contributor or higher access. Creating a workspace identity, changing tenant settings, and assigning capacity may need administrator rights.
Before the DAG can call dbt:
- enable workspace identity;
- add it to the workspace as Contributor or Member;
- create a connection using Workspace identity;
- enable Allow Code-First Artifacts;
- enable Fabric Connections on Airflow;
- attach the connection; and
- record its GUID.
Model the build as dependencies
Capacity
-> Workspace
-> Workspace identity
-> Lakehouse
-> dbt Job
-> Gold Delta tables
-> Direct Lake semantic model
-> Fabric connection
-> Airflow item
-> Airflow DAG
-> dbt Job execution
Keep a manifest with:
| Logical name | Fabric type | Identifier |
|---|---|---|
| Workspace | Workspace | Workspace ID |
| Lakehouse | Lakehouse | Item ID and SQL endpoint ID |
| dbt Job | DataBuildToolJob | Item ID |
| Airflow Job | ApacheAirflowJob | Item ID and web URL |
| Semantic model | SemanticModel | Item ID |
| Workspace identity | Service principal | Object and application ID |
| Fabric connection | Connection | Connection GUID |
Names are for people. APIs use GUIDs. List resources and match exact names; never copy identifiers from another environment.
Know the runtime
| Target | Adapter | SQL |
|---|---|---|
| Fabric Data Warehouse | dbt-fabric 1.10.0 | T-SQL |
| Fabric Lakehouse | dbt-fabricspark 1.12.2 | Spark SQL |
Fabric dbt Job runtime V1.0 supplies dbt Core 1.11, Python 3.12, and
dbt-fabricspark 1.12.2 for a Lakehouse profile. The Lakehouse SQL analytics
endpoint is read-only. An agent should reject a plan that points
dbt-fabric at that endpoint and tries to write models.
Phase 0: preflight
Verify:
- authentication for
https://api.fabric.microsoft.com; - a capacity supporting dbt and Airflow;
- dbt Jobs enabled;
- Contributor or higher access;
- workspace identity and code-first connections; and
- exact item names.
Discover workspaces:
GET https://api.fabric.microsoft.com/v1/workspaces
List items in the selected workspace and filter by exact type and name.
Phase 1: create the Lakehouse
Create:
jaffle_shop_lakehouse
Use logical schemas:
raw
staging
gold
Record both the Lakehouse item ID and SQL endpoint ID. They are different resources.
Phase 2: deploy the dbt Job
Create:
jaffle_shop_native_dbt
A Lakehouse definition has this conceptual shape:
{
"project": {
"projectType": "OneLake",
"folderPath": "dbt"
},
"profile": {
"profileType": "Lakehouse",
"schema": "gold",
"connectionSettings": {
"name": "jaffle_shop_lakehouse",
"properties": {
"type": "Lakehouse",
"typeProperties": {
"workspaceId": "<workspace-id>",
"artifactId": "<lakehouse-id>",
"endPoint": "<lakehouse-sql-endpoint>"
}
}
}
},
"command": {
"operation": "build",
"arguments": {
"failFast": true,
"threads": 4
}
}
}
Project configuration:
name: jaffle_shop
version: 1.0.0
config-version: 2
profile: jaffle_shop
model-paths: ["models"]
seed-paths: ["seeds"]
macro-paths: ["macros"]
models:
jaffle_shop:
staging:
+schema: staging
+materialized: view
marts:
+schema: gold
+materialized: table
seeds:
jaffle_shop:
+schema: raw
Use data_tests for primary-key uniqueness, non-nullability, customer
relationships, and date relationships.
Fabric definitions are multipart. The safe update loop is:
- export the live definition;
- decode every part;
- change the intended files;
- re-encode all required parts;
- submit one
updateDefinition; and - export again and compare.
Start the job:
POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances?jobType=Execute
Poll:
GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances
202 Accepted or item creation is not success. Follow the specific instance
to Completed or Failed.
Phase 3: validate data before BI
select count(*) as order_count
from gold.fct_orders;
select
sum(amount) as total_revenue,
min(order_date) as first_order_date,
max(order_date) as last_order_date
from gold.fct_orders;
select count(*) as orphan_customers
from gold.fct_orders f
left join gold.dim_customers c
on f.customer_id = c.customer_id
where c.customer_id is null;
The validated test oracle was:
| Check | Expected |
|---|---|
| Orders | 12 |
| Revenue | 175.00 |
| Customers | 10 |
| Dates | 11 |
| Orphan customer keys | 0 |
| Orphan date keys | 0 |
Phase 4: build Direct Lake
Create:
jaffle_shop_direct_lake
Use a star schema:
dim_customers 1 ---- * fct_orders
dim_dates 1 ---- * fct_orders
Relationships:
| From | To | Direction |
|---|---|---|
fct_orders[customer_id] | dim_customers[customer_id] | Single |
fct_orders[order_date] | dim_dates[date_day] | Single |
Measures:
Total Orders =
DISTINCTCOUNT(fct_orders[order_id])
Total Revenue =
SUM(fct_orders[amount])
Average Order Value =
DIVIDE([Total Revenue], [Total Orders])
Point the model at Gold Lakehouse entities, not the SQL endpoint as DirectQuery. After deployment, confirm the model, frame it or refresh as appropriate, and run:
EVALUATE
ROW(
"Orders", [Total Orders],
"Revenue", [Total Revenue],
"Average Order Value", [Average Order Value],
"Customers", COUNTROWS('dim_customers'),
"Dates", COUNTROWS('dim_dates')
)
Then test date slicing:
EVALUATE
SUMMARIZECOLUMNS(
dim_dates[calendar_year],
dim_dates[month_number],
"Orders", [Total Orders],
"Revenue", [Total Revenue]
)
ORDER BY
dim_dates[calendar_year],
dim_dates[month_number]
Phase 5: configure workspace identity
The identity chain is:
Workspace identity
-> Contributor or Member
-> WorkspaceIdentity Fabric connection
-> Allow code-first artifacts
-> Connection attached to Airflow
-> Connection GUID in DAG
Create the identity, grant workspace access, create the connection, enable code-first artifacts, and record the GUID. The display name is not enough.
Phase 6: deploy Airflow
Create:
jaffle_shop_airflow_orchestrator
Definition shape:
{
"properties": {
"type": "Airflow",
"typeProperties": {
"airflowProperties": {
"airflowEnvironment": "FabricAirflowJob-1.0.0",
"airflowVersion": "2.10.5",
"pythonVersion": "3.12",
"enableAADIntegration": true,
"enableFabricConnections": true,
"enableTriggerers": true
},
"fabricConnections": ["<fabric-connection-guid>"]
}
}
}
DAG:
from datetime import datetime, timedelta
from airflow import DAG
from airflow.providers.microsoft.fabric.operators.run_item import (
MSFabricRunJobOperator,
)
FABRIC_CONN_ID = "<fabric-connection-guid>"
WORKSPACE_ID = "<workspace-id>"
DBT_JOB_ID = "<dbt-job-id>"
with DAG(
dag_id="orchestrate_jaffle_shop_dbt",
schedule=None,
start_date=datetime(2026, 1, 1),
catchup=False,
default_args={
"owner": "fabric",
"retries": 1,
"retry_delay": timedelta(minutes=5),
},
) as dag:
run_jaffle_shop_dbt = MSFabricRunJobOperator(
task_id="run_jaffle_shop_dbt",
fabric_conn_id=FABRIC_CONN_ID,
workspace_id=WORKSPACE_ID,
item_id=DBT_JOB_ID,
job_type="DataBuildToolJob",
wait_for_termination=True,
deferrable=True,
)
fabric_conn_id is the GUID. job_type="Execute" is wrong; use
DataBuildToolJob or DBT.
Phase 7: start Airflow
POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment/start?beta=true
Status:
GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment?beta=true
Expected shape:
{
"status": "Started",
"airflowWebUrl": "https://<managed-airflow-host>/login/"
}
Starting an already started environment is an idempotency signal. Follow the status instead of treating it as a failure.
At the tested CLI version, fab job start did not support ApacheAirflowJob.
Use the documented REST environment API or portal.
Phase 8: run the DAG
The documented route is the Airflow UI:
- open the Airflow item;
- enable Fabric Connections;
- attach the workspace-identity connection;
- open or create the DAG; and
- select Run DAG.
The managed endpoint also exposes:
POST {airflowWebUrl}/api/v1/dags/orchestrate_jaffle_shop_dbt/dagRuns
Content-Type: application/json
{
"conf": {},
"note": "Triggered by deployment agent"
}
The Airflow endpoint can require a different Entra audience. A Fabric control-plane token may redirect to interactive sign-in. Check the challenge before requesting a token.
Phase 9: prove the chain
Airflow DAG run
-> Airflow task instance
-> Fabric dbt Job instance
-> Gold Lakehouse output
-> Direct Lake DAX result
Capture the DAG ID and state, task state and try, dbt instance and state, dbt counts, Lakehouse checks, and DAX results. The validated deployment returned:
| Layer | Evidence |
|---|---|
| Airflow | DAG success |
| Airflow task | MSFabricRunJobOperator success |
| dbt | PASS=23 WARN=0 ERROR=0 SKIP=0 |
| Lakehouse | 12 orders and 175.00 revenue |
| Semantic model | 12 orders, 175.00 revenue, 14.5833 average order value |
Failure atlas
| Symptom | Likely cause | Action |
|---|---|---|
InvalidDefinitionFormat | Unsupported format or malformed definition | Export without guessing and use the documented envelope |
Invalid LinkedService | Warehouse source used for Lakehouse | Use the Lakehouse profile |
| Project directory missing | folderPath mismatch | Align folder and uploaded parts |
| dbt seed-key failure | Source differs or test assumptions wrong | Export live files and compare |
conn_id ... isn't defined | Display name used | Use connection GUID |
| Airflow operator rejects run type | job_type="Execute" | Use DataBuildToolJob or DBT |
UserAccessTokenException | Token acquisition failed | Verify identity and connection |
| Airflow redirects to sign-in | Wrong audience | Use the managed Airflow audience |
Environment says Started | Start called twice | Validate status and continue |
fab job start unsupported | CLI gap | Use REST or portal |
| Duplicate semantic models | Blind retry | List duplicates and deploy once |
| Unexpected DAX totals | Contract or relationship error | Revalidate rows, relationships, and measures |
What the agent must never do
- Commit tokens, secrets, or credentials.
- Use a display name where a GUID is required.
- Invent an endpoint after a
404. - Infer success from item existence.
- Stop at dbt when the requested outcome includes Airflow and BI.
- Deploy a partial semantic-model definition.
- Retry model creation blindly.
- Treat Warehouse and Lakehouse adapters as interchangeable.
- Trust a local project after live deployment without exporting and reconciling.
- Call an untested path working.
Final lesson
The hard part is not generating SQL, a DAG, or a DAX measure. It is preserving the contracts between workloads:
- dbt uses the Lakehouse adapter and Spark SQL;
- Lakehouse publishes stable Delta contracts;
- Direct Lake consumes Gold;
- Airflow authenticates through workspace identity and a GUID connection;
- the Airflow operator expects an item type, not
Execute; and - asynchronous operations must reach a terminal state.
Once those boundaries are explicit, the native stack is repeatable. Without them, an agent can leave behind a workspace that looks complete and still fails at the first real orchestration run.
References
- dbt job in Microsoft Fabric
- Configure a dbt job
- Workspace identity in Apache Airflow Jobs
- Run a Fabric item using Apache Airflow DAGs
- Apache Airflow environment REST API
- Develop Direct Lake semantic models
- Fabric REST API documentation
- Apache Airflow provider for Microsoft Fabric
- Native Fabric dbt + Airflow + Direct Lake walkthrough
Loading comments…