De bijbehorende walkthrough legt uit hoe je Jaffle Shop als native Fabric-oplossing bouwt. Dit runbook beantwoordt een moeilijkere vraag: wat moet een engineering agent weten om dit opnieuw te maken zonder “item bestaat” te verwarren met “het systeem werkt”?
De agent heeft het item model, de adaptergrens, identity chain, definition format, long-running operations en een validatiecontract nodig. Het doelpad:
Jaffle Shop seeds
|
v
Fabric dbt Job
dbt Core + dbt-fabricspark
|
v
Fabric Lakehouse / OneLake
raw -> staging -> gold
|
+-----------------------+
| |
v v
Fabric Apache Airflow Direct Lake semantic model
workspace identity relationships + DAX
De geteste run gebruikte een Fabric trial-workspace. De Airflow DAG startte de dbt Job, de dbt build werd voltooid, Gold Delta tables werden gevuld en DAX-queries gaven de verwachte resultaten.
Handoff voordat de agent begint
Skills for Microsoft Fabric zijn een vereiste voor de agent, niet voor de runtime. Ze bieden workflowinstructies en toolrouting. Ze leveren geen capacity, credentials, tenant permissions of workspacetoegang.
De machine heeft nodig:
- GitHub Copilot CLI of een andere agenthost met skillondersteuning;
- Skills for Fabric geïnstalleerd en ingeschakeld;
- Azure CLI;
- Git en een schrijfbare projectdirectory; en
- netwerktoegang tot Fabric, Entra ID, OneLake en managed Airflow.
Optionele integraties zijn Power BI Modeling MCP, FabricIQ en een Fabric SQL endpoint-tool of ODBC-client. Accepteer de EULA als Power BI Modeling MCP wordt gebruikt.
Meld aan met:
az login
Fabric control-planeoperations gebruiken:
https://api.fabric.microsoft.com
Hergebruik die token audience niet voor managed Airflow of Power BI zonder de service te controleren. Gebruik delegated sign-in, workspace identity, managed identity of een goedgekeurde service principal. Geef nooit een password aan de agent.
Fabric voorbereiden
Een Fabric administrator moet bevestigen:
- beschikbare capacity of trial;
- dbt Jobs (preview);
- Apache Airflow Jobs;
- service-principal access waar workspace identity wordt gebruikt; en
- vereiste previewswitches voor code-first en workspace identity.
De deployment identity heeft Contributor- of hogere toegang nodig. Een workspace identity maken, tenant settings wijzigen en capacity toewijzen kan administratorrechten vereisen.
Voordat de DAG dbt kan aanroepen:
- schakel workspace identity in;
- voeg die als Contributor of Member aan de workspace toe;
- maak een connection met Workspace identity;
- schakel Allow Code-First Artifacts in;
- schakel Fabric Connections op Airflow in;
- koppel de connection; en
- leg de GUID vast.
Modelleer de build als dependencies
Capacity
-> Workspace
-> Workspace identity
-> Lakehouse
-> dbt Job
-> Gold Delta tables
-> Direct Lake semantic model
-> Fabric connection
-> Airflow item
-> Airflow DAG
-> dbt Job execution
Houd een manifest bij:
| Logische naam | Fabric-type | Identifier |
|---|---|---|
| Workspace | Workspace | Workspace ID |
| Lakehouse | Lakehouse | Item ID and SQL endpoint ID |
| dbt Job | DataBuildToolJob | Item ID |
| Airflow Job | ApacheAirflowJob | Item ID and web URL |
| Semantic model | SemanticModel | Item ID |
| Workspace identity | Service principal | Object and application ID |
| Fabric connection | Connection | Connection GUID |
Namen zijn voor mensen. API’s gebruiken GUIDs. List resources en match exacte namen; kopieer nooit identifiers uit een andere omgeving.
Ken de runtime
| Target | Adapter | SQL |
|---|---|---|
| Fabric Data Warehouse | dbt-fabric 1.10.0 | T-SQL |
| Fabric Lakehouse | dbt-fabricspark 1.12.2 | Spark SQL |
Fabric dbt Job runtime V1.0 levert dbt Core 1.11, Python 3.12 en
dbt-fabricspark 1.12.2 voor een Lakehouse-profile. Het Lakehouse SQL
analytics endpoint is read-only. Een agent moet een plan afwijzen dat
dbt-fabric op dat endpoint richt en modellen probeert te schrijven.
Fase 0: preflight
Controleer:
- authentication voor
https://api.fabric.microsoft.com; - een capacity die dbt en Airflow ondersteunt;
- dbt Jobs ingeschakeld;
- Contributor- of hogere toegang;
- workspace identity en code-first connections; en
- exacte itemnamen.
Ontdek workspaces:
GET https://api.fabric.microsoft.com/v1/workspaces
List items in de geselecteerde workspace en filter op exact type en naam.
Fase 1: maak het Lakehouse
Maak:
jaffle_shop_lakehouse
Gebruik logische schema’s:
raw
staging
gold
Leg zowel de Lakehouse item ID als SQL endpoint ID vast. Het zijn verschillende resources.
Fase 2: deploy de dbt Job
Maak:
jaffle_shop_native_dbt
Een Lakehouse-definition heeft conceptueel deze vorm:
{
"project": {
"projectType": "OneLake",
"folderPath": "dbt"
},
"profile": {
"profileType": "Lakehouse",
"schema": "gold",
"connectionSettings": {
"name": "jaffle_shop_lakehouse",
"properties": {
"type": "Lakehouse",
"typeProperties": {
"workspaceId": "<workspace-id>",
"artifactId": "<lakehouse-id>",
"endPoint": "<lakehouse-sql-endpoint>"
}
}
}
},
"command": {
"operation": "build",
"arguments": {
"failFast": true,
"threads": 4
}
}
}
Projectconfiguratie:
name: jaffle_shop
version: 1.0.0
config-version: 2
profile: jaffle_shop
model-paths: ["models"]
seed-paths: ["seeds"]
macro-paths: ["macros"]
models:
jaffle_shop:
staging:
+schema: staging
+materialized: view
marts:
+schema: gold
+materialized: table
seeds:
jaffle_shop:
+schema: raw
Gebruik data_tests voor primary-key uniqueness, non-nullability,
customerrelationships en daterelationships.
Fabric-definitions zijn multipart. De veilige updatecyclus:
- exporteer de live definition;
- decodeer ieder part;
- wijzig de bedoelde bestanden;
- encodeer alle vereiste parts opnieuw;
- verstuur één
updateDefinition; en - exporteer opnieuw en vergelijk.
Start de job:
POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances?jobType=Execute
Poll:
GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances
202 Accepted of item creation is geen succes. Volg de specifieke instance tot
Completed of Failed.
Fase 3: valideer data vóór BI
select count(*) as order_count
from gold.fct_orders;
select
sum(amount) as total_revenue,
min(order_date) as first_order_date,
max(order_date) as last_order_date
from gold.fct_orders;
select count(*) as orphan_customers
from gold.fct_orders f
left join gold.dim_customers c
on f.customer_id = c.customer_id
where c.customer_id is null;
De gevalideerde test oracle:
| Controle | Verwacht |
|---|---|
| Orders | 12 |
| Revenue | 175.00 |
| Customers | 10 |
| Dates | 11 |
| Orphan customer keys | 0 |
| Orphan date keys | 0 |
Fase 4: bouw Direct Lake
Maak:
jaffle_shop_direct_lake
Gebruik een star schema:
dim_customers 1 ---- * fct_orders
dim_dates 1 ---- * fct_orders
Relationships:
| Van | Naar | Direction |
|---|---|---|
fct_orders[customer_id] | dim_customers[customer_id] | Single |
fct_orders[order_date] | dim_dates[date_day] | Single |
Measures:
Total Orders =
DISTINCTCOUNT(fct_orders[order_id])
Total Revenue =
SUM(fct_orders[amount])
Average Order Value =
DIVIDE([Total Revenue], [Total Orders])
Richt het model op Gold Lakehouse-entities, niet op het SQL endpoint als DirectQuery. Bevestig het model na deployment, frame of refresh waar nodig en voer uit:
EVALUATE
ROW(
"Orders", [Total Orders],
"Revenue", [Total Revenue],
"Average Order Value", [Average Order Value],
"Customers", COUNTROWS('dim_customers'),
"Dates", COUNTROWS('dim_dates')
)
Test daarna dateslicing:
EVALUATE
SUMMARIZECOLUMNS(
dim_dates[calendar_year],
dim_dates[month_number],
"Orders", [Total Orders],
"Revenue", [Total Revenue]
)
ORDER BY
dim_dates[calendar_year],
dim_dates[month_number]
Fase 5: configureer workspace identity
De identity chain:
Workspace identity
-> Contributor or Member
-> WorkspaceIdentity Fabric connection
-> Allow code-first artifacts
-> Connection attached to Airflow
-> Connection GUID in DAG
Maak de identity, verleen workspacetoegang, maak de connection, schakel code-first artifacts in en leg de GUID vast. De display name is niet genoeg.
Fase 6: deploy Airflow
Maak:
jaffle_shop_airflow_orchestrator
Definition shape:
{
"properties": {
"type": "Airflow",
"typeProperties": {
"airflowProperties": {
"airflowEnvironment": "FabricAirflowJob-1.0.0",
"airflowVersion": "2.10.5",
"pythonVersion": "3.12",
"enableAADIntegration": true,
"enableFabricConnections": true,
"enableTriggerers": true
},
"fabricConnections": ["<fabric-connection-guid>"]
}
}
}
DAG:
from datetime import datetime, timedelta
from airflow import DAG
from airflow.providers.microsoft.fabric.operators.run_item import (
MSFabricRunJobOperator,
)
FABRIC_CONN_ID = "<fabric-connection-guid>"
WORKSPACE_ID = "<workspace-id>"
DBT_JOB_ID = "<dbt-job-id>"
with DAG(
dag_id="orchestrate_jaffle_shop_dbt",
schedule=None,
start_date=datetime(2026, 1, 1),
catchup=False,
default_args={
"owner": "fabric",
"retries": 1,
"retry_delay": timedelta(minutes=5),
},
) as dag:
run_jaffle_shop_dbt = MSFabricRunJobOperator(
task_id="run_jaffle_shop_dbt",
fabric_conn_id=FABRIC_CONN_ID,
workspace_id=WORKSPACE_ID,
item_id=DBT_JOB_ID,
job_type="DataBuildToolJob",
wait_for_termination=True,
deferrable=True,
)
fabric_conn_id is de GUID. job_type="Execute" is fout; gebruik
DataBuildToolJob of DBT.
Fase 7: start Airflow
POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment/start?beta=true
Status:
GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment?beta=true
Verwachte vorm:
{
"status": "Started",
"airflowWebUrl": "https://<managed-airflow-host>/login/"
}
Een al gestarte omgeving opnieuw starten is een idempotencysignaal. Volg de status in plaats van dit als failure te behandelen.
Bij de geteste CLI-versie ondersteunde fab job start geen
ApacheAirflowJob. Gebruik de gedocumenteerde REST environment API of portal.
Fase 8: draai de DAG
De gedocumenteerde route is de Airflow UI:
- open het Airflow-item;
- schakel Fabric Connections in;
- koppel de workspace-identity connection;
- open of maak de DAG; en
- selecteer Run DAG.
Het managed endpoint biedt ook:
POST {airflowWebUrl}/api/v1/dags/orchestrate_jaffle_shop_dbt/dagRuns
Content-Type: application/json
{
"conf": {},
"note": "Triggered by deployment agent"
}
Het Airflow-endpoint kan een andere Entra audience vereisen. Een Fabric control-plane token kan naar interactieve sign-in redirecten. Controleer de challenge voordat je een token aanvraagt.
Fase 9: bewijs de keten
Airflow DAG run
-> Airflow task instance
-> Fabric dbt Job instance
-> Gold Lakehouse output
-> Direct Lake DAX result
Leg DAG ID en state, task state en try, dbt instance en state, dbt counts, Lakehouse-checks en DAX-results vast. De gevalideerde deployment gaf:
| Laag | Evidence |
|---|---|
| Airflow | DAG success |
| Airflow task | MSFabricRunJobOperator success |
| dbt | PASS=23 WARN=0 ERROR=0 SKIP=0 |
| Lakehouse | 12 orders and 175.00 revenue |
| Semantic model | 12 orders, 175.00 revenue, 14.5833 average order value |
Failure atlas
| Symptoom | Waarschijnlijke oorzaak | Actie |
|---|---|---|
InvalidDefinitionFormat | Unsupported format of malformed definition | Export zonder te gokken en gebruik de documented envelope |
Invalid LinkedService | Warehouse source gebruikt voor Lakehouse | Gebruik het Lakehouse-profile |
| Project directory missing | folderPath mismatch | Lijn folder en uploaded parts uit |
| dbt seed-key failure | Source verschilt of test assumptions fout | Export live files en vergelijk |
conn_id ... isn't defined | Display name gebruikt | Gebruik connection GUID |
| Airflow operator rejects run type | job_type="Execute" | Gebruik DataBuildToolJob of DBT |
UserAccessTokenException | Token acquisition failed | Verifieer identity en connection |
| Airflow redirects to sign-in | Wrong audience | Gebruik de managed Airflow audience |
Environment says Started | Start tweemaal aangeroepen | Valideer status en ga door |
fab job start unsupported | CLI-gap | Gebruik REST of portal |
| Duplicate semantic models | Blind retry | List duplicates en deploy één keer |
| Unexpected DAX totals | Contract- of relationshipfout | Valideer rows, relationships en measures opnieuw |
Wat de agent nooit mag doen
- Tokens, secrets of credentials committen.
- Een display name gebruiken waar een GUID nodig is.
- Na een
404een endpoint verzinnen. - Succes afleiden uit het bestaan van een item.
- Stoppen bij dbt wanneer het gevraagde resultaat Airflow en BI bevat.
- Een gedeeltelijke semantic-modeldefinition deployen.
- Model creation blind opnieuw proberen.
- Warehouse- en Lakehouse-adapters als uitwisselbaar behandelen.
- Een lokaal project na live deployment vertrouwen zonder export en reconciliation.
- Een ongetest pad werkend noemen.
Laatste les
Het moeilijke deel is niet SQL, een DAG of een DAX-measure genereren. Het is de contracten tussen workloads bewaren:
- dbt gebruikt de Lakehouse-adapter en Spark SQL;
- Lakehouse publiceert stabiele Delta-contracten;
- Direct Lake consumeert Gold;
- Airflow authenticeert via workspace identity en een GUID-connection;
- de Airflow-operator verwacht een item type, niet
Execute; en - asynchronous operations moeten een terminal state bereiken.
Met expliciete grenzen is de native stack herhaalbaar. Zonder die grenzen kan een agent een workspace achterlaten die compleet lijkt en toch bij de eerste echte orchestration run faalt.
Referenties
- dbt job in Microsoft Fabric
- Een dbt job configureren
- Workspace identity in Apache Airflow Jobs
- Een Fabric-item draaien met Apache Airflow DAGs
- Apache Airflow environment REST API
- Direct Lake semantic models ontwikkelen
- Fabric REST API-documentatie
- Apache Airflow-provider voor Microsoft Fabric
- Native Fabric dbt + Airflow + Direct Lake-walkthrough
Reacties laden…