Ga naar de inhoud
R Roesli.
Ga terug
ai-agents

Het agent-runbook voor native Fabric dbt, Airflow, Lakehouse en Direct Lake

Een evidence-driven agent-runbook voor het deployen en valideren van Fabric dbt Jobs, Lakehouse, Airflow, workspace identity, Direct Lake en DAX.

De bijbehorende walkthrough legt uit hoe je Jaffle Shop als native Fabric-oplossing bouwt. Dit runbook beantwoordt een moeilijkere vraag: wat moet een engineering agent weten om dit opnieuw te maken zonder “item bestaat” te verwarren met “het systeem werkt”?

De agent heeft het item model, de adaptergrens, identity chain, definition format, long-running operations en een validatiecontract nodig. Het doelpad:

Jaffle Shop seeds
       |
       v
Fabric dbt Job
dbt Core + dbt-fabricspark
       |
       v
Fabric Lakehouse / OneLake
raw -> staging -> gold
       |
       +-----------------------+
       |                       |
       v                       v
Fabric Apache Airflow     Direct Lake semantic model
workspace identity        relationships + DAX

De geteste run gebruikte een Fabric trial-workspace. De Airflow DAG startte de dbt Job, de dbt build werd voltooid, Gold Delta tables werden gevuld en DAX-queries gaven de verwachte resultaten.

Handoff voordat de agent begint

Skills for Microsoft Fabric zijn een vereiste voor de agent, niet voor de runtime. Ze bieden workflowinstructies en toolrouting. Ze leveren geen capacity, credentials, tenant permissions of workspacetoegang.

De machine heeft nodig:

Optionele integraties zijn Power BI Modeling MCP, FabricIQ en een Fabric SQL endpoint-tool of ODBC-client. Accepteer de EULA als Power BI Modeling MCP wordt gebruikt.

Meld aan met:

az login

Fabric control-planeoperations gebruiken:

https://api.fabric.microsoft.com

Hergebruik die token audience niet voor managed Airflow of Power BI zonder de service te controleren. Gebruik delegated sign-in, workspace identity, managed identity of een goedgekeurde service principal. Geef nooit een password aan de agent.

Fabric voorbereiden

Een Fabric administrator moet bevestigen:

De deployment identity heeft Contributor- of hogere toegang nodig. Een workspace identity maken, tenant settings wijzigen en capacity toewijzen kan administratorrechten vereisen.

Voordat de DAG dbt kan aanroepen:

  1. schakel workspace identity in;
  2. voeg die als Contributor of Member aan de workspace toe;
  3. maak een connection met Workspace identity;
  4. schakel Allow Code-First Artifacts in;
  5. schakel Fabric Connections op Airflow in;
  6. koppel de connection; en
  7. leg de GUID vast.

Modelleer de build als dependencies

Capacity
  -> Workspace
      -> Workspace identity
      -> Lakehouse
          -> dbt Job
              -> Gold Delta tables
                  -> Direct Lake semantic model
          -> Fabric connection
              -> Airflow item
                  -> Airflow DAG
                      -> dbt Job execution

Houd een manifest bij:

Logische naamFabric-typeIdentifier
WorkspaceWorkspaceWorkspace ID
LakehouseLakehouseItem ID and SQL endpoint ID
dbt JobDataBuildToolJobItem ID
Airflow JobApacheAirflowJobItem ID and web URL
Semantic modelSemanticModelItem ID
Workspace identityService principalObject and application ID
Fabric connectionConnectionConnection GUID

Namen zijn voor mensen. API’s gebruiken GUIDs. List resources en match exacte namen; kopieer nooit identifiers uit een andere omgeving.

Ken de runtime

TargetAdapterSQL
Fabric Data Warehousedbt-fabric 1.10.0T-SQL
Fabric Lakehousedbt-fabricspark 1.12.2Spark SQL

Fabric dbt Job runtime V1.0 levert dbt Core 1.11, Python 3.12 en dbt-fabricspark 1.12.2 voor een Lakehouse-profile. Het Lakehouse SQL analytics endpoint is read-only. Een agent moet een plan afwijzen dat dbt-fabric op dat endpoint richt en modellen probeert te schrijven.

Fase 0: preflight

Controleer:

  1. authentication voor https://api.fabric.microsoft.com;
  2. een capacity die dbt en Airflow ondersteunt;
  3. dbt Jobs ingeschakeld;
  4. Contributor- of hogere toegang;
  5. workspace identity en code-first connections; en
  6. exacte itemnamen.

Ontdek workspaces:

GET https://api.fabric.microsoft.com/v1/workspaces

List items in de geselecteerde workspace en filter op exact type en naam.

Fase 1: maak het Lakehouse

Maak:

jaffle_shop_lakehouse

Gebruik logische schema’s:

raw
staging
gold

Leg zowel de Lakehouse item ID als SQL endpoint ID vast. Het zijn verschillende resources.

Fase 2: deploy de dbt Job

Maak:

jaffle_shop_native_dbt

Een Lakehouse-definition heeft conceptueel deze vorm:

{
  "project": {
    "projectType": "OneLake",
    "folderPath": "dbt"
  },
  "profile": {
    "profileType": "Lakehouse",
    "schema": "gold",
    "connectionSettings": {
      "name": "jaffle_shop_lakehouse",
      "properties": {
        "type": "Lakehouse",
        "typeProperties": {
          "workspaceId": "<workspace-id>",
          "artifactId": "<lakehouse-id>",
          "endPoint": "<lakehouse-sql-endpoint>"
        }
      }
    }
  },
  "command": {
    "operation": "build",
    "arguments": {
      "failFast": true,
      "threads": 4
    }
  }
}

Projectconfiguratie:

name: jaffle_shop
version: 1.0.0
config-version: 2
profile: jaffle_shop

model-paths: ["models"]
seed-paths: ["seeds"]
macro-paths: ["macros"]

models:
  jaffle_shop:
    staging:
      +schema: staging
      +materialized: view
    marts:
      +schema: gold
      +materialized: table

seeds:
  jaffle_shop:
    +schema: raw

Gebruik data_tests voor primary-key uniqueness, non-nullability, customerrelationships en daterelationships.

Fabric-definitions zijn multipart. De veilige updatecyclus:

  1. exporteer de live definition;
  2. decodeer ieder part;
  3. wijzig de bedoelde bestanden;
  4. encodeer alle vereiste parts opnieuw;
  5. verstuur één updateDefinition; en
  6. exporteer opnieuw en vergelijk.

Start de job:

POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances?jobType=Execute

Poll:

GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/items/{dbtJobId}/jobs/instances

202 Accepted of item creation is geen succes. Volg de specifieke instance tot Completed of Failed.

Fase 3: valideer data vóór BI

select count(*) as order_count
from gold.fct_orders;
select
    sum(amount) as total_revenue,
    min(order_date) as first_order_date,
    max(order_date) as last_order_date
from gold.fct_orders;
select count(*) as orphan_customers
from gold.fct_orders f
left join gold.dim_customers c
    on f.customer_id = c.customer_id
where c.customer_id is null;

De gevalideerde test oracle:

ControleVerwacht
Orders12
Revenue175.00
Customers10
Dates11
Orphan customer keys0
Orphan date keys0

Fase 4: bouw Direct Lake

Maak:

jaffle_shop_direct_lake

Gebruik een star schema:

dim_customers  1 ---- *  fct_orders
dim_dates      1 ---- *  fct_orders

Relationships:

VanNaarDirection
fct_orders[customer_id]dim_customers[customer_id]Single
fct_orders[order_date]dim_dates[date_day]Single

Measures:

Total Orders =
DISTINCTCOUNT(fct_orders[order_id])
Total Revenue =
SUM(fct_orders[amount])
Average Order Value =
DIVIDE([Total Revenue], [Total Orders])

Richt het model op Gold Lakehouse-entities, niet op het SQL endpoint als DirectQuery. Bevestig het model na deployment, frame of refresh waar nodig en voer uit:

EVALUATE
ROW(
    "Orders", [Total Orders],
    "Revenue", [Total Revenue],
    "Average Order Value", [Average Order Value],
    "Customers", COUNTROWS('dim_customers'),
    "Dates", COUNTROWS('dim_dates')
)

Test daarna dateslicing:

EVALUATE
SUMMARIZECOLUMNS(
    dim_dates[calendar_year],
    dim_dates[month_number],
    "Orders", [Total Orders],
    "Revenue", [Total Revenue]
)
ORDER BY
    dim_dates[calendar_year],
    dim_dates[month_number]

Fase 5: configureer workspace identity

De identity chain:

Workspace identity
  -> Contributor or Member
  -> WorkspaceIdentity Fabric connection
  -> Allow code-first artifacts
  -> Connection attached to Airflow
  -> Connection GUID in DAG

Maak de identity, verleen workspacetoegang, maak de connection, schakel code-first artifacts in en leg de GUID vast. De display name is niet genoeg.

Fase 6: deploy Airflow

Maak:

jaffle_shop_airflow_orchestrator

Definition shape:

{
  "properties": {
    "type": "Airflow",
    "typeProperties": {
      "airflowProperties": {
        "airflowEnvironment": "FabricAirflowJob-1.0.0",
        "airflowVersion": "2.10.5",
        "pythonVersion": "3.12",
        "enableAADIntegration": true,
        "enableFabricConnections": true,
        "enableTriggerers": true
      },
      "fabricConnections": ["<fabric-connection-guid>"]
    }
  }
}

DAG:

from datetime import datetime, timedelta

from airflow import DAG
from airflow.providers.microsoft.fabric.operators.run_item import (
    MSFabricRunJobOperator,
)

FABRIC_CONN_ID = "<fabric-connection-guid>"
WORKSPACE_ID = "<workspace-id>"
DBT_JOB_ID = "<dbt-job-id>"

with DAG(
    dag_id="orchestrate_jaffle_shop_dbt",
    schedule=None,
    start_date=datetime(2026, 1, 1),
    catchup=False,
    default_args={
        "owner": "fabric",
        "retries": 1,
        "retry_delay": timedelta(minutes=5),
    },
) as dag:
    run_jaffle_shop_dbt = MSFabricRunJobOperator(
        task_id="run_jaffle_shop_dbt",
        fabric_conn_id=FABRIC_CONN_ID,
        workspace_id=WORKSPACE_ID,
        item_id=DBT_JOB_ID,
        job_type="DataBuildToolJob",
        wait_for_termination=True,
        deferrable=True,
    )

fabric_conn_id is de GUID. job_type="Execute" is fout; gebruik DataBuildToolJob of DBT.

Fase 7: start Airflow

POST https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment/start?beta=true

Status:

GET https://api.fabric.microsoft.com/v1/workspaces/{workspaceId}/apacheAirflowJobs/{airflowJobId}/environment?beta=true

Verwachte vorm:

{
  "status": "Started",
  "airflowWebUrl": "https://<managed-airflow-host>/login/"
}

Een al gestarte omgeving opnieuw starten is een idempotencysignaal. Volg de status in plaats van dit als failure te behandelen.

Bij de geteste CLI-versie ondersteunde fab job start geen ApacheAirflowJob. Gebruik de gedocumenteerde REST environment API of portal.

Fase 8: draai de DAG

De gedocumenteerde route is de Airflow UI:

  1. open het Airflow-item;
  2. schakel Fabric Connections in;
  3. koppel de workspace-identity connection;
  4. open of maak de DAG; en
  5. selecteer Run DAG.

Het managed endpoint biedt ook:

POST {airflowWebUrl}/api/v1/dags/orchestrate_jaffle_shop_dbt/dagRuns
Content-Type: application/json

{
  "conf": {},
  "note": "Triggered by deployment agent"
}

Het Airflow-endpoint kan een andere Entra audience vereisen. Een Fabric control-plane token kan naar interactieve sign-in redirecten. Controleer de challenge voordat je een token aanvraagt.

Fase 9: bewijs de keten

Airflow DAG run
  -> Airflow task instance
      -> Fabric dbt Job instance
          -> Gold Lakehouse output
              -> Direct Lake DAX result

Leg DAG ID en state, task state en try, dbt instance en state, dbt counts, Lakehouse-checks en DAX-results vast. De gevalideerde deployment gaf:

LaagEvidence
AirflowDAG success
Airflow taskMSFabricRunJobOperator success
dbtPASS=23 WARN=0 ERROR=0 SKIP=0
Lakehouse12 orders and 175.00 revenue
Semantic model12 orders, 175.00 revenue, 14.5833 average order value

Failure atlas

SymptoomWaarschijnlijke oorzaakActie
InvalidDefinitionFormatUnsupported format of malformed definitionExport zonder te gokken en gebruik de documented envelope
Invalid LinkedServiceWarehouse source gebruikt voor LakehouseGebruik het Lakehouse-profile
Project directory missingfolderPath mismatchLijn folder en uploaded parts uit
dbt seed-key failureSource verschilt of test assumptions foutExport live files en vergelijk
conn_id ... isn't definedDisplay name gebruiktGebruik connection GUID
Airflow operator rejects run typejob_type="Execute"Gebruik DataBuildToolJob of DBT
UserAccessTokenExceptionToken acquisition failedVerifieer identity en connection
Airflow redirects to sign-inWrong audienceGebruik de managed Airflow audience
Environment says StartedStart tweemaal aangeroepenValideer status en ga door
fab job start unsupportedCLI-gapGebruik REST of portal
Duplicate semantic modelsBlind retryList duplicates en deploy één keer
Unexpected DAX totalsContract- of relationshipfoutValideer rows, relationships en measures opnieuw

Wat de agent nooit mag doen

  1. Tokens, secrets of credentials committen.
  2. Een display name gebruiken waar een GUID nodig is.
  3. Na een 404 een endpoint verzinnen.
  4. Succes afleiden uit het bestaan van een item.
  5. Stoppen bij dbt wanneer het gevraagde resultaat Airflow en BI bevat.
  6. Een gedeeltelijke semantic-modeldefinition deployen.
  7. Model creation blind opnieuw proberen.
  8. Warehouse- en Lakehouse-adapters als uitwisselbaar behandelen.
  9. Een lokaal project na live deployment vertrouwen zonder export en reconciliation.
  10. Een ongetest pad werkend noemen.

Laatste les

Het moeilijke deel is niet SQL, een DAG of een DAX-measure genereren. Het is de contracten tussen workloads bewaren:

Met expliciete grenzen is de native stack herhaalbaar. Zonder die grenzen kan een agent een workspace achterlaten die compleet lijkt en toch bij de eerste echte orchestration run faalt.

Referenties


Deel deze post:

Lees verder

Vorige post
Bouw je eerste planningsoplossing met Plan in Microsoft Fabric
Volgende post
Een native Fabric analytics stack met dbt, Airflow, Lakehouse en Direct Lake
Community

Praat mee

Meld je aan met GitHub om een reactie achter te laten.

GitHub

Reacties laden…

Meld je aan met GitHub om te reageren