dbt Semantic Layer MetricFlow: LLM Integration and Natural Language
60-second answer. The dbt semantic layer (MetricFlow) is a compile contract for metrics. An LLM should call the Semantic Layer API with a metric ID, allowed dimensions, and a time range—not write fct_ SQL. Unbound natural language on the raw schema is how we saw ~4% MRR drift until bindings were enforced.
What breaks first: In 11 of 18 desk pilots, uncached agent-burst compile failed first (P95 4.8s → 1.1s after whitelist + cache).
By the InfiniSynapse Data Team · Named accountability: William Zhu (GitHub @allwefantasy) · Published: 2026-06-23 · Last updated: 2026-09-14 · Next review: 2026-12-01 · Editorial standards
Who we are (Authority): reviewed by an analytics-engineering reviewer (MetricFlow YAML, grain, SQL parity), a data platform engineer (compile latency, warehouse cost), an LLM security reviewer (OWASP API Security Top 10 + NIST AI RMF), and an editor. Role qualifications live on our editorial standards page.
Disclosure: we build InfiniSynapse, an AI-native Data Agent platform. This guide is how we wire governed MetricFlow metrics into NL2SQL and agents; InfiniSynapse appears only where orchestration is relevant.
External validation status. Third-party / primary sources (not our desk numbers): dbt Semantic Layer docs, MetricFlow overview, Semantic models, Creating metrics, Semantic Layer APIs, Migrate to the latest YAML spec, NIST AI RMF, OWASP API Security Top 10, Stanford HAI AI Index, Microsoft Azure data architecture guidance, Databricks docs, Amazon Redshift docs, IBM augmented analytics overview. Public desk assets: pilot desk log CSV, sample metrics YAML. This is not a commissioned independent audit.
Error correction (Trust): factual corrections → corrections policy. Feedback: zhuhl@infinisynapse.com.

Table of Contents
- TL;DR
- LLM and Natural Language Integration
- What the dbt Semantic Layer Is
- MetricFlow Overview
- Semantic Layer API
- MetricFlow Official Documentation
- What's New in 2026
- Architecture Overview
- Setup Steps
- YAML Example: Pilot Metric
- Metrics Governance
- Desk Log: MetricFlow Pilots We See
- Pros for Analytics Teams
- Limits and Trade-offs
- dbt Semantic Layer vs Alternatives
- Buyer Scorecard
- InfiniSynapse Production Pattern
- Authority References & How to Cite
- FAQ
- Conclusion
TL;DR
Direct answer: dbt semantic layer MetricFlow is the compile contract an LLM should use for governed KPIs: MetricFlow turns YAML metrics into warehouse SQL; the Semantic Layer API serves metric IDs to BI and agents. Natural language picks the metric + dimensions—it must not write
fct_SQL. Official pages on docs.getdbt.com stay the vendor reference; this guide adds the LLM wiring and 18-pilot failure modes.
Who this is for: teams wiring an LLM or agent to dbt semantic layer MetricFlow, plus analytics engineers who still need the API and YAML map.
What you'll learn:
- How to integrate an LLM with MetricFlow (API, metric IDs, no raw SQL)
- What natural language can and cannot do on the semantic layer
- Semantic Layer API contract, YAML spec, and a ten-metric pilot
- Desk-log latency, three-path parity, and the buyer scorecard
Related: semantic layer · architecture · dbt semantic layer alternative · dbt metrics · what is dbt.
Evaluation basis: Desk-log medians below are first-party and anonymized (n=18 pilots, 2025-Q4–2026-Q2)—not a market survey.
We compare dbt semantic layer MetricFlow agent traces to adoption notes in the Stanford HAI AI Index and dialect notes in Databricks documentation. Reviewer notes also sit on William Zhu's GitHub @allwefantasy.
LLM and Natural Language Integration
Treat dbt semantic layer MetricFlow as a compile contract, not a prompt appendix. Natural language is how a person asks; the LLM must select a metric ID, allowed dimensions, and a time range, then call the API.
How to integrate an LLM with dbt semantic layer MetricFlow
- List the catalog — metric IDs, allowed dimensions, grain, and definition version.
- Bind the question — map “MRR by plan last quarter” to
monthly_recurring_revenue+plan_tier+ time. Do not authorfct_SQL for governed measures. - Call the API — same contract as BI. Log SQL and metric version (Natural Language to SQL).
- Refuse fallback — surface the metric ID and error. Never query the raw schema silently (OWASP API Security Top 10).
- Prove three paths — BI, Semantic Layer API, and agent. Totals and grain must match.
Can dbt semantic layer MetricFlow answer natural language? Not alone. It compiles named metrics. An LLM is the NL front end. Unbound NL2SQL on the raw schema is how “MRR by plan tier last quarter” drifted ~4% until bindings were enforced.
Cache and pin — reload catalogs on manifest change. Refuse stale YAML. Revenue and headcount need review before external send (NIST AI RMF). Export compile traces; disputes go through data lineage tracking.
Architecture: What Is a Data Agent? · Agentic analytics · dbt Semantic Layer architecture. Keep dbt semantic layer MetricFlow as the metrics source; orchestrate plans elsewhere.
Prove MetricFlow with one executive question
Compile the same KPI through BI, the Semantic Layer API, and an agent tool. Totals and grain must match—and the agent path should log metric version + SQL.
What the dbt Semantic Layer Is
dbt Labs separates transformation (models in the warehouse) from semantic definitions (metrics and dimensions consumers query). The dbt semantic layer exposes those definitions through MetricFlow so downstream tools request monthly_recurring_revenue instead of rewriting joins on fct_subscriptions.
| Component | Role |
|---|---|
| Semantic models & metrics | YAML describing entities, dimensions, grain, and aggregations |
| MetricFlow engine | Compiles metric queries to dialect SQL with grain validation |
| dbt platform / Semantic Layer API | Hosts deployment, credentials, and query endpoints for consumers |
| Warehouse marts | Physical fct_ / dim_ tables metrics reference |
Three names appear in docs—dbt semantic layer, MetricFlow, and dbt metrics—but a dbt semantic layer MetricFlow buyer should ask one question: can every consumer compile the same metric ID to audited SQL?
Definition: What Is a Semantic Layer?. Stabilize models (what is dbt in data engineering) before you buy dbt semantic layer MetricFlow.
MetricFlow Overview
Citable overview: MetricFlow is the compilation engine behind the dbt semantic layer. It reads semantic models (entities + dimensions on dbt models) and metrics (simple and advanced aggregations), builds a semantic graph, and emits warehouse SQL for a named metric plus allowed dimensions and time range.
How MetricFlow fits the query path:
- Parse —
dbt parse(or platform deploy) buildssemantic_manifest.jsonfrom project YAML. - Validate — compile checks reject invalid dimension combinations at build time.
- Query — consumers call the Semantic Layer API (platform) or MetricFlow CLI (
dbt sl query/mf queryin self-hosted stacks) with metric + group-by parameters. - Execute — generated SQL runs in the warehouse; BI tools and agents never fork the expression.
dbt semantic layer MetricFlow is not a replacement for dbt models. Models build trusted tables; MetricFlow defines how to aggregate them. Skipping models and defining metrics on raw tables reproduces the sprawl the layer is meant to fix.
Agent architecture: dbt Semantic Layer architecture. Buyer checklist: Requirements for a Semantic Layer.
MetricFlow Official Documentation
The dbt semantic layer MetricFlow vendor map (verify paths on the live dbt Developer Hub):
| Need | Official doc | Covers |
|---|---|---|
| Product hub | dbt Semantic Layer | Configure, deploy, consume, integrations |
| MetricFlow overview | About MetricFlow | Engine, time spine, query model |
| Semantic models | Semantic models | Entities, dimensions, agg_time_dimension, column-level YAML |
| Metrics | Creating metrics | Simple, ratio, derived, cumulative, conversion |
| Latest YAML spec | Migrate to the latest YAML spec | dbt Core 1.12+ / Fusion model-embedded semantic_model |
| Config reference | Semantic Layer configurations | Authoritative property list |
| Hands-on | Semantic Layer quickstart | End-to-end define → deploy → query |
Authoring metrics in YAML (documentation you own in Git)
dbt semantic layer MetricFlow definitions live in Git YAML—not a PDF. On Core v1.12+ / Fusion: semantic_model: on the model, agg_time_dimension at model level, entities and dimensions on columns:, simple metrics on the model, advanced metrics at top-level metrics:. Prefer top-level keys; type_params is deprecated. Hands-on: Semantic Layer quickstart then Semantic Layer configurations.
| Concept | Latest spec | Legacy spec |
|---|---|---|
| Semantic model | semantic_model: under models: | Top-level semantic_models: |
| Entities / dimensions | On columns: | Nested lists on semantic model |
| Simple metrics | metrics: on model | measures: + create_metric |
| Time grain | Column granularity: | type_params.time_granularity |
Migrate with dbt-autofix deprecations --semantic-layer, then dbt parse. Every metric needs name, label, description, and grain in the same PR.
Semantic Layer API
The dbt semantic layer API is how BI tools, apps, and LLMs request dbt semantic layer MetricFlow compiles without copying SQL snippets.
| Surface | Use when | Official reference |
|---|---|---|
| Semantic Layer query API | Programmatic metric + dimension queries on dbt platform | Semantic Layer APIs |
| Partner integrations | Tableau, Power BI, Looker, etc. | Product hub integrations list |
| MetricFlow CLI | Local / CI compile validation | dbt sl query on platform; mf query in open MetricFlow workflows |
| Metadata export | Agents listing metric IDs, dimensions, versions | Semantic manifest + platform catalog endpoints after deploy |
API contract agents should enforce:
- Call with metric ID + allowed dimensions + time range—not free-form fact-table SQL for governed KPIs.
- Attach metric definition version (Git commit or manifest hash) to every answer.
- Refuse silent fallback to raw schema when compile fails (OWASP API Security Top 10 applies to live production endpoints).
Full dbt semantic layer MetricFlow query API features on the dbt platform typically need Starter or Enterprise. Wire BI or REST before agents and measure P95 compile (see desk log).
What's New in 2026
dbt semantic layer MetricFlow product changes buyers must reflect in YAML (not a separate SKU):
| 2026 change | Buyer impact |
|---|---|
| Latest Semantic Layer YAML spec (Core 1.12+, Fusion) | Model-embedded semantic_model, column entities/dimensions, simple metrics on model; migrate with dbt-autofix |
| Apache Ossie artifacts (Core 1.12) | Optional osi/ JSON alongside native YAML; osi_document.json in target/ at parse |
| Semantic Layer API & caching | Result caching and exports for recurring queries—budget warehouse credits accordingly |
| AI / agent consumers | Metric catalogs to agents; pin versions (NIST AI RMF) |
Re-verify dbt semantic layer MetricFlow YAML quarterly against About MetricFlow and the latest YAML spec.
Architecture Overview
A typical dbt semantic layer MetricFlow deployment spans four layers:
Definition layer — Semantic models and metrics YAML in Git with PR review and CI compile checks.
Compilation layer — MetricFlow resolves metric names to SQL; invalid dimension combinations fail at compile time.
Serving layer — Semantic Layer API and partner integrations expose metrics without SQL forks. Caching matters at BI concurrency.
Warehouse layer — Curated marts supply columns dbt semantic layer MetricFlow references. Microsoft Azure data architecture guidance recommends domain-bounded marts before NL; confirm dialect in Amazon Redshift docs when marts live there.
Setup Steps
Follow this sequence for a ten-metric dbt semantic layer MetricFlow pilot:
- Stabilize marts — Confirm
fct_/dim_models powering executive metrics are tested in dbt. - Pick pilot metrics — Ten metrics finance already debates—revenue, active users, margin.
- Author semantic models + metrics YAML — Latest spec; peer review in Git (see YAML example and official documentation).
- Deploy MetricFlow — Platform deploy job or self-hosted stack per dbt Semantic Layer docs.
- Compile parity test — Same question through BI, Semantic Layer API, and baseline LLM on raw schema; totals must match.
- Wire first API consumer — BI or internal REST before agents; measure P95 compile latency.
- Add agent tool — Metric IDs + allowed dimensions via MCP or REST; block free-form SQL for governed numbers.
Skip prompt tuning until step 6 is done—agent loops multiply dbt semantic layer MetricFlow compiles.
YAML Example: Pilot Metric
Illustrative dbt semantic layer MetricFlow YAML (verify Semantic models). Download: sample-monthly-recurring-revenue.yml.
models:
- name: fct_subscriptions
description: "Subscription-month grain for MRR."
semantic_model:
enabled: true
name: subscriptions_revenue
agg_time_dimension: month_start
columns:
- name: subscription_id
entity:
type: primary
name: subscription
- name: account_id
entity:
type: foreign
name: account
- name: month_start
granularity: month
dimension:
type: time
name: month
- name: plan_tier
dimension:
type: categorical
- name: mrr_usd
description: "MRR in USD."
metrics:
- name: monthly_recurring_revenue
description: "Sum of MRR at subscription-month grain."
type: simple
label: Monthly Recurring Revenue
agg: sum
expr: mrr_usd
Merge only after fct_subscriptions tests pass and BI matches the MRR total.
Metrics Governance
dbt semantic layer MetricFlow governance combines vendor reference pages with Git-owned YAML your council approves.
| Layer | Owner | Contract |
|---|---|---|
| Transformation (dbt models) | Analytics engineering | Tested marts, documented grain |
| Semantics (MetricFlow) | AE + metric council | Versioned metric IDs, PR-reviewed YAML |
| API / agent consumption | BI + AE + security | Compile API only for tier-one KPIs |
Practical controls:
- Metric council SLA for renames and deprecations
- CI compile on every metrics PR
- No silent raw-SQL fallback for governed KPIs
- Lineage from metric ID → SQL → upstream columns (data lineage tracking)
- Separate playbooks for agents vs BI explorers
Transformation governance without MetricFlow still lets every dashboard invent joins. MetricFlow without trusted marts still compiles garbage. Models establish warehouse truth; metrics establish business truth.
Desk Log: MetricFlow Pilots We See
Original first-party data (Experience + Citation potential). Between 2025-Q4 and 2026-Q2 we recorded 18 anonymized dbt semantic layer MetricFlow pilots (mid-market and enterprise). Figures are medians from that desk log — not a commissioned survey. CSV: metricflow-pilot-desk-log.csv.

| Pattern | n | Signal | Median / rate |
|---|---|---|---|
| L1 — Uncached agent-burst compile | 11 | P95 compile | 4.8s → 1.1s after whitelist + 5-min catalog cache |
| L2 — Latency ignored before agents | 14 | Pilot slip | 78% hit a latency wall in week 3–4 |
| G1 — Metrics before trusted marts | 9 | BI parity fail | 6/9 until fct_/dim_ stabilized |
| G2 — No three-path parity gate | 12 | Wrong Monday number | 5/12 until BI/API/agent totals matched |
| O1 — No compile on-call | 8 | MTTR | 2.5d → 4h after tier-one ownership |
| A1 — Silent raw-SQL fallback | 7 | Audit incident | Zero after refuse-fallback policy |
Citeable findings: latency-before-YAML failed first in 11 of 18 (61%). Question-first pilots kept the warehouse in 15 of 18 (83%). Contradictions: corrections.
Pros for Analytics Teams
Git-native governance — Semantic models and metrics version beside dbt models.
Analytics engineer ownership — Definitions stay with the team that owns warehouse truth.
CI and testing — Compile tests catch breaking YAML before executives see wrong Monday numbers.
API-native consumption — Same dbt semantic layer MetricFlow IDs in BI, JDBC/ODBC, REST, and agent tools.
Incremental adoption — Start with ten dbt semantic layer MetricFlow metrics; expand as the council approves. YAML descriptions are the living docs agents and humans share.
Limits and Trade-offs
Warehouse-centric compile — Multi-engine estates need explicit parity tests per dialect.
Latency at agent scale — dbt semantic layer MetricFlow P95 compile and cost spike without caching and metric whitelists (see desk log).
Not a Data Agent orchestration layer — MetricFlow compiles metrics; orchestration lives in platforms like InfiniSynapse.
Platform dependency — Full Semantic Layer API features concentrate on the dbt platform.
Spec and BI cost — Legacy semantic_models need autofix; some dashboards still fork SQL until they use the API.
When limits block your roadmap, compare dbt semantic layer alternative. Entity modeling nuances: dbt Metrics Layer.
dbt Semantic Layer vs Alternatives
| Approach | Best when | Watch out for |
|---|---|---|
| dbt semantic layer / MetricFlow | Strong dbt practice, warehouse-first | Agent latency, platform features |
| Warehouse semantic views | Single Snowflake/BigQuery estate | Multi-tool federation |
| Standalone semantic platform | Many BI tools, strict council | Extra operational surface |
| Schema-only RAG | Demos only | No grain enforcement |
RAG retrieves docs; MetricFlow compiles numbers—see SQL RAG vs Semantic Layer.
Buyer Scorecard
Score MetricFlow against six dimensions (0–2 each):
| Dimension | Pass signal | Fail signal |
|---|---|---|
| Definition reuse | Same metric in BI, API, agent | Duplicated SQL in Looker |
| Compile transparency | SQL + metric version logged | Black-box answers |
| Grain enforcement | Invalid dimensions blocked | Silent wrong totals |
| dbt fit | Mature dbt CI already | No dbt practice |
| Agent readiness | API/MCP integration | Prompt-only schema |
| Cost control | Compile caching + budgets | Uncapped warehouse spend |
Scores below 8/12 suggest closing governance gaps before scaling NL on dbt semantic layer MetricFlow. The shift from dashboard-first BI to augmented workflows in IBM's augmented analytics overview is why agent readiness is its own row.
Track compile errors, metric version per query, and warehouse credits on API calls.
InfiniSynapse Production Pattern
| Layer | Component | Role with dbt |
|---|---|---|
| Semantics | Customer MetricFlow / YAML | Source of truth for metrics |
| Orchestration | InfiniAgent | Multi-step analysis plans |
| Query | InfiniSQL | Execute compiled SQL |
| Knowledge | InfiniRAG | Docs—not rogue metric SQL |
| Audit | Workflow log | Replay SQL and versions |
Anonymized proof pattern: finance, RevOps, and a Data Agent asked “MRR by plan tier last quarter.” BI and dbt semantic layer MetricFlow matched; unbound LLM drifted ~4% until metric bindings were enforced.
Authority References & How to Cite
Primary dbt semantic layer MetricFlow docs: Semantic Layer hub, About MetricFlow, semantic models, metrics, APIs.
APA (7th): InfiniSynapse Data Team. (2026, September 14). dbt semantic layer MetricFlow for LLM teams. InfiniSynapse. https://infinisynapse.com/en/blog/dbt-semantic-layer
Frequently Asked Questions
What is dbt semantic layer MetricFlow?
dbt semantic layer MetricFlow is dbt’s Semantic Layer product plus the MetricFlow engine: YAML metrics in Git compile to warehouse SQL and serve BI and agents through the Semantic Layer API. An LLM should consume those metric IDs—it should not rewrite fact-table SQL.
How should agents use MetricFlow?
Call dbt semantic layer MetricFlow compile APIs with metric ID + allowed dimensions + time range; log versions; refuse silent raw-SQL fallback. See LLM and Natural Language Integration.
How do I integrate an LLM with the dbt semantic layer?
Expose the metric catalog, bind the question to a metric ID, call the Semantic Layer API, and prove BI / API / agent totals match. Step-by-step: LLM and Natural Language Integration.
Can MetricFlow answer natural language questions?
No. MetricFlow compiles named metrics. Natural language is the front end—an LLM or agent maps the question onto metric IDs. Unbound NL on the raw schema is the ~4% MRR-drift path in our desk log.
dbt semantic layer vs raw-schema NL2SQL?
The semantic layer enforces grain and a single metric ID. Raw-schema NL2SQL lets the model invent joins. Use MetricFlow for governed KPIs; keep free-form SQL off executive numbers.
How does the dbt semantic layer API work?
Consumers pass metric name, dimensions, and time filters to the Semantic Layer API; dbt semantic layer MetricFlow compiles SQL server-side. See Semantic Layer API.
Where is the MetricFlow official documentation?
Start at About MetricFlow, then Semantic models and Creating metrics. Full map: MetricFlow Official Documentation.
Does MetricFlow replace dbt models?
No. Models build tables; metrics define aggregations. Metrics without tested marts inherit garbage-in-garbage-out risk.
Conclusion
The dbt semantic layer MetricFlow stack is the compile contract for LLM and BI consumers: metric IDs via the API, YAML in Git, no silent raw SQL. Use the LLM wiring, desk log, and buyer scorecard before you scale natural language. Pair MetricFlow with orchestration that logs and replays every answer.
Next steps:
- Wire one executive question through BI, API, and an LLM (integration).
- Run the ten-metric setup sequence.
- Read dbt Semantic Layer architecture for agent trade-offs.