Skip to main content

Fabric capacity cost is not set by the rate card. It is set by CU-seconds consumed, smoothed across 30-second timepoints, and by how fast you clear carryforward before throttling starts. EPC Group runs capacity as an operating discipline: named roles, a monitoring cadence with thresholds, and a Capacity Debt Ledger. EPC Group is a Houston-based Microsoft consulting firm operating since 1997, with six Microsoft Solutions Partner designations.

Key Facts

  • Fabric capacity is $0.18 per CU-hour pay-as-you-go and $938.00 per CU per year reserved (US East, July 2026).
  • Consumption is evaluated in 30-second timepoints — 2,880 in 24 hours. An F64 has 1,920 CU-seconds in each.
  • Interactive operations smooth over five to 64 minutes; background operations over 24 hours.
  • The throttling ladder has four stages: overage protection, interactive delay, interactive rejection, background rejection.
  • Recovery is arithmetic: ((% − 100) / 100) × window. A 250% background rejection means at least 36 hours.
  • Capacity overage bills at 3× pay-as-you-go — $0.54 per CU-hour in US East.
  • Autoscale Billing for Spark moves Spark off capacity entirely at 0.5 CU hour, with no bursting or smoothing.

Last updated by Errin O'Connor, Founder & Chief AI Architect, EPC Group

Quick facts

QuestionAnswer
Unit of consumptionCapacity Unit second (CU-second); 1 CU-hour = 3,600 CU-seconds
Evaluation window30 seconds (a “timepoint”)
CU-hours per day, F641,536 (F256 = 6,144; F512 = 12,288)
SmoothingInteractive 5–64 minutes; background 24 hours
First throttling stage20-second delay on new interactive operations
Primary instrumentFabric Capacity Metrics app (Compute 14 days; Storage 30 days)
Fastest way to stop throttlingPause and resume, or scale up
Spark / Warehouse on capacity1 CU = 2 Spark vCores or 0.5 Warehouse vCores
Chargeback sourceFabric Chargeback app (refreshed daily)

What a CU-second actually is

For finance. A Fabric capacity is a metered pipe, not a server. The SKU number is the pipe’s width in Capacity Units: an F64 delivers 64 CUs every second, 1,536 CU-hours a day. Every query, refresh, notebook run and pipeline consumes some of that flow, measured in CU-seconds. You are billed for the pipe, not per query, and the question is whether the work fits through it. When it does not, Fabric does not send a larger bill by default — it slows people down. That is why Fabric cost management is really reliability management with a price tag.

For engineers. Multiply the SKU’s CU count by 30 for the CU-seconds available in one 30-second timepoint: an F64 has 1,920. Two mechanisms sit between raw consumption and that ceiling. Bursting lets an operation temporarily use more compute than the SKU provides, so a big job finishes fast. Smoothing then spreads that operation’s CU cost across future timepoints — five to 64 minutes for interactive, 24 hours for background. A one-CU-hour background job on an F2 contributes 3,600 CU-seconds ÷ 2,880 timepoints = 1.25 CU-seconds per timepoint, about 2.1% of each, even though it consumed six times the compute available in the next ten minutes.

The consequence: peak utilization is not the number that matters — smoothed future consumption is. A capacity showing 300% instantaneous spikes may never throttle, while one showing a flat 95% background line is a single bad pipeline away from a 24-hour outage. If you are still on P SKUs, the P-to-F migration runbook comes first, and the F SKU cost model derives the reserved break-even.

The Four-Desk Capacity Operating Model

Most Fabric cost overruns are ownership failures: nobody is accountable for the capacity as an object, so nobody watches it. EPC Group’s Four-Desk Capacity Operating Model assigns four standing responsibilities. One person can hold two desks; the desks still need naming.

DeskAccountable forFabric permissionCadence
Capacity OwnerSKU decision, reservation, budget, go/no-go on scaling upAzure subscription owner / reservation purchaserMonthly
Capacity StewardUtilization, throttling, surge thresholds, pause schedules, incident responseFabric capacity administratorDaily
Workload OwnerThe CU cost of their workspace — refresh schedules, model design, pipeline efficiencyWorkspace admin (+ capacity contributor)Weekly
FinOps AnalystAllocation, showback, forecast variance, maintaining the Capacity Debt LedgerRead access to Chargeback and Cost ManagementMonthly

Two permission details matter. A capacity contributor can assign workspaces but cannot change capacity settings or delete the capacity — the correct grant for Workload Owners. Only a capacity administrator can enable surge protection, capacity overage or autoscale billing for Spark. Do not hand capacity admin out for convenience; all three settings have direct cost consequences.

The Capacity Debt Ledger

Capacity debt is the accumulated gap between what a workload was sized for and what it actually costs. It is not a metaphor: Fabric has a literal debt mechanic in carryforward CUs, which accrue when smoothed consumption exceeds a timepoint’s allowance and must be paid off through burndown before throttling stops. The Metrics app Health page reports a Cumulative debt sparkline per capacity. The Ledger extends that from the platform’s 24-hour horizon to your budget’s twelve-month one. Keep one row per workload — usually per workspace — and refresh it monthly from the Chargeback app and the Metrics app matrix.

ColumnSourceExample
WorkloadWorkspace or item nameWS-Finance-Reporting
OwnerNamed Workload OwnerJ. Alvarez
Sized CU-hours/monthBudgeted at onboarding180
Actual CU-hours/monthChargeback app, Utilization (CU) by date274
Debt (CU-hours)Actual − Sized+94
Debt ($)Debt × $0.18 (or reserved equivalent)$16.92
Debt trend3-month direction↑ rising
Overloaded minutesMetrics app, matrix by item and operation42
Performance deltaMetrics app, % change vs 7 days ago−18%
Root causeModel design, refresh frequency, concurrency, data growth, new featureHourly refresh on a daily source
Remediation / due / statusAction, date, Open / In progress / Written off2× daily refresh, 2026-09-15, Open

Four rules make it work.

Rule 1 — every workload gets a size at onboarding. A workspace admitted to a shared capacity without a CU-hour budget cannot accrue debt, because there is no baseline. This is the highest-leverage governance change available.

Rule 2 — debt is denominated in CU-hours first, dollars second. CU-hours are stable across regions and billing models; dollars move with your reservation.

Rule 3 — write-offs are explicit. Sometimes a workload is legitimately bigger than it was sized for. Re-baseline it, record the write-off, raise the forecast. Never let unexplained debt roll forward silently — that is how an F64 becomes an F256 with nobody able to say why.

Rule 4 — the ledger drives the SKU conversation, not the reverse. When total ledger debt exceeds roughly 15% of capacity for three consecutive months and remediation has failed, that is the evidence for a scale-up. Before then, scaling up buys your way out of a design problem. Our data governance practice treats it as a governance artifact.

Monitoring cadence

FrequencyMetricWhereAction threshold
DailyHealth status per capacityMetrics app Health pageAnything other than Healthy or Suspended → Steward investigates same day
DailyInteractive delay %Compute → ThrottlingAny 30-second window >100% → open an incident
DailyBlocked workspacesHealth page>0 → confirm the block was intended
Weekly24-hour background %Compute → Throttling → Background rejectionSustained >70% → review refresh schedules
WeeklyPeak and average utilizationCompute page cardsAverage >80% for five consecutive days → begin scale-up analysis
WeeklyOverloaded minutes and performance delta by itemCompute → matrix by item and operation>30 overloaded minutes, or delta worse than −25% week on week → assign to the Workload Owner
MonthlyCU-hours by workspace, item, domainChargeback appVariance >20% against ledger baseline → record as debt
MonthlyBilled overage CU-hoursCompute → Overages (Billed)Recurring overage → compare 3× cost against a scale-up
MonthlyReservation coverage, storage growthCost Management / Storage pageUncovered always-on CU → extend the reservation
QuarterlyTotal ledger debt vs capacityCapacity Debt Ledger>15% for three months → SKU decision

Automate the daily rows. Fabric emits Microsoft.Fabric.Capacity.Summary events every 30 seconds and Microsoft.Fabric.Capacity.State on state change through the Real-Time hub. Build an Activator rule on backgroundRejectionThresholdPercentage, interactiveDelayThresholdPercentage or interactiveRejectionThresholdPercentage that emails the Steward or triggers a user-defined function. Native email alerts at 100% of provisioned CU are also available.

Throttling: the ladder and the recovery maths

Throttling is progressive by design, so refreshes survive longer than dashboards.

Future capacity consumedStageUser impact
≤ 10 minutesOverage protectionNone. Jobs may consume 10 minutes of future capacity freely
10–60 minutesInteractive delayNew interactive operations delayed 20 seconds at submission
60 minutes – 24 hoursInteractive rejectionInteractive operations rejected; background operations still start and run
> 24 hoursBackground rejectionAll requests rejected, interactive and background

Recovery is calculable: minimum time to recover = ((% of the rejection type − 100) ÷ 100) × window duration. At 250%, that is 15 minutes for interactive delay, 90 minutes for interactive rejection, 36 hours for background rejection. Those are minimums — background jobs keep accumulating future consumption, so real incidents run longer.

Three behaviours change per workload. Almost all Warehouse operations are reported as background to take advantage of 24-hour smoothing, so warehouse users hit rejection rather than delay. Real-Time Intelligence skips the 20-second delay stage and throttles only at the 60-minute rejection threshold. Eventstreams are not throttled at all; the CU allocated to keeping streams open is reduced until the capacity recovers. In-flight operations are never throttled — only new submissions.

Users see status code CapacityLimitExceeded with “Your organization’s Fabric compute capacity has exceeded its limits. Try again later”, or “Cannot load model due to reaching capacity limits.” Put both strings in the service-desk knowledge base so tickets route to the Steward, not the report author.

The throttling-incident response runbook

  1. Confirm it is throttling, not design. Slow reports are more often a bad semantic model than an overloaded capacity. Open the Compute page, filter to the incident time, and check whether CU% exceeded 100%. If not, hand it to the Workload Owner as a performance problem.
  2. Identify the stage. Check the Interactive delay, Interactive rejection and Background rejection tabs. Whichever is above 100% tells you what users experience and which recovery window applies.
  3. Calculate recovery time and publish it. An honest “reports will be rejected for roughly 90 more minutes” beats an hour of silence.
  4. Decide whether to intervene. Capacities self-heal. If recovery is under 30 minutes and no executive-visible workload is affected, wait.
  5. If you intervene, choose one lever. Scale up — more idle capacity per timepoint burns down carryforward faster. Pause and resume — pausing bills accumulated smoothed usage immediately and the capacity resumes with zero future consumption, clearing throttling instantly. Capacity overage — pays off the current window at 3×. The trap: enabling overage during heavy throttling charges all cumulative carryforward at the moment you switch it on.
  6. Find the cause. Drill into the timepoint summary and detail pages to rank operation types and items by CU-seconds in the window that tipped over.
  7. Apply a control, not just a fix. If one workspace caused it, set a workspace-level surge limit. If background jobs caused it, set a capacity-level background rejection threshold below 100%.
  8. Post an entry in the Capacity Debt Ledger. Every throttling incident is debt made visible. Record workload, root cause, remediation and owner before closing.

Eviction and prioritization policy for competing workloads

A shared capacity has no built-in notion of importance. Declare it.

Workload classControlConfigurationTrade-off
Executive and regulatory reportingIsolateOwn capacity, never sharedHighest cost, lowest risk; the only real guarantee
Business-critical, shared capacityMission criticalWorkspace state = Mission CriticalExempt from workspace-level surge protection, not from capacity-level throttling
Standard departmentalAvailableDefault state; subject to surge protectionAuto-blocked if it exceeds its CU % limit
Known noisy neighbourWorkspace CU limitRejection threshold as % of capacity over rolling 24 hChecks run every 5 minutes, so limits are soft
Actively misbehavingBlockedManual block, indefinite or for N hoursAll operations rejected; unblocking is a manual reset
Bursty SparkMove off capacityAutoscale Billing for SparkServerless pay-as-you-go; batch jobs queue, interactive throttle
Copilot and data agentsCentralizeDesignate a Fabric Copilot capacityConsolidates AI spend, including Power BI Desktop and PPU workspaces

Capacity-level surge protection is the blunt instrument: set a background operations rejection threshold and a recovery threshold, and the capacity rejects new background operations before entering deep throttling. Read its limits honestly. It does not stop in-progress jobs, guarantee interactive requests are spared, block autoscale-billed operations, or affect OneLake activity. The rejection threshold is not an upper bound either, because in-flight jobs keep reporting usage. Microsoft’s own conclusion: to fully protect a critical solution, isolate it in its own capacity.

Workspace-level surge protection has sharper edges. It does not apply to Dataflows Gen1, paginated reports, scorecards or Activator; autoscale compute is excluded; and once a workspace is auto-blocked, raising the limit or deleting the rule will not unblock it.

The dev/test pause schedule

Pause and resume is the largest structural saving on F SKUs, and it does not exist on P SKUs — one of the better reasons to finish the P-to-F migration. See also Power BI Premium and Premium vs PPU.

Worked example, US East rates, July 2026. A dev/test F16 costs 16 × $0.18 = $2.88 per hour, so 730 × $2.88 = $2,102.40 per month running continuously. Run it 12 hours a day, five days a week — about 260 hours — and it costs $748.80. Saving: $1,353.60 per month, roughly 64%, about $16,200 a year on one non-production capacity.

Chargeback and showback allocation methods

MethodHow it worksBest forWeakness
Direct capacityOne capacity per business unit; the Azure resource is the cost centreLarge, funded unitsPoor utilization; you pay for everyone’s headroom
Workspace proportionalChargeback app: workspace % of capacity CU × the capacity billShared capacity, most organisationsShared platform services are hard to attribute
Domain rollupChargeback domain/subdomain view aggregates workspaces to a business domainEstates with Fabric domains configuredUnassigned items land in “No domain”
User-attributedUtilization (CU) details matrix, per-user breakdownSelf-service governanceService principals report as “Power BI Service”
Showback onlyPublish consumption, do not billFirst 6–12 months of a shared capacitySlower behaviour change than real chargeback
Azure tagsTag the capacity resource; allocate in Cost ManagementTeams running Azure tag policyResolves to capacity only, never to workspace

Start with showback by workspace for two quarters, then move to workspace-proportional chargeback once the Ledger has baselines to defend the numbers. Two notes: the Chargeback app refreshes daily, so month-end settles a day late; and if the tenant setting Show user data in the Fabric Capacity Metrics app and reports is disabled, every user appears as “Masked user”. Settle that trade-off before promising per-user reporting. For the framing that makes chargeback stick with finance, see our CFO AI governance conversation.

What breaks — failure modes

SymptomRoot causeFix
Capacity looks fine at 95% but users are rejectedPeak utilization is not the throttling signal; smoothed future consumption isRead the Throttling tab, not Utilization
Surge protection is on but the Metrics app shows old thresholdsThrottling visuals do not reflect applied surge settingsCheck actual thresholds in the Admin portal
Raising a workspace CU limit does not unblock the workspaceAuto-blocked workspaces stay blocked until an admin resets themSet the workspace back to Available manually
A “mission critical” workspace is still throttledMission critical exempts from workspace-level surge protection onlyIsolate the workload in its own capacity
Overage charges appear despite a spending limitThe limit is evaluated every 5 minutes, so it can be exceededSet it below one-third of daily CU hours; monitor the overage meter
Costs jumped after enabling autoscale billing for SparkSpark left the capacity and bills separately at 0.5 CU hourDownsize the base SKU; monitor the autoscale meter
Reducing the autoscale Max CU killed running jobsLowering, enabling or disabling Max CU cancels active Spark jobsChange the setting in a maintenance window
A preview workload shows zero cost, then starts costing moneyNon-billable operations do not drain capacity — until the feature goes billableTrack non-billable usage and size for it before GA

What changed in 2026

Frequently asked questions

What is a capacity unit in Microsoft Fabric?

A Capacity Unit (CU) is Fabric's measure of compute. An F64 provides 64 CUs every second, or 1,536 CU-hours per day. Consumption is measured in CU-seconds and evaluated in 30-second timepoints, 2,880 per day. One CU-hour equals 3,600 CU-seconds.

Why does my capacity throttle when utilization is below 100%?

Because throttling is driven by smoothed future consumption, not instantaneous utilization. Background operations smooth over 24 hours and interactive over five to 64 minutes, so committed future consumption can exceed the threshold while the current timepoint looks healthy. Read the Throttling tab, not Utilization.

How long does Fabric throttling last?

At minimum, ((percentage − 100) ÷ 100) × the window duration. A 250% interactive delay clears in about 15 minutes, a 250% interactive rejection in about 90 minutes, and a 250% background rejection in about 36 hours. Background jobs keep accruing consumption, so real incidents usually run longer.

What is the fastest way to stop throttling?

Pause and resume the capacity. Pausing bills the accumulated smoothed usage and returns the capacity to a healthy state immediately; on resume it has zero future consumption. Scaling up also works, by burning down carryforward faster.

Is capacity overage cheaper than scaling up?

Only for rare, short spikes. Overage bills at three times the pay-as-you-go rate, so Microsoft recommends keeping the limit below one-third of daily CU hours — beyond that, the cost matches scaling up the SKU. If you are throttled regularly, scale up instead.

Does surge protection protect my dashboards?

Not reliably. Capacity-level surge protection rejects new background operations earlier than the system default, reducing pressure on interactive users, but it does not guarantee interactive requests avoid delay or rejection. To fully protect a critical solution, isolate it in its own correctly sized capacity.

How much can I save by pausing dev/test capacity?

At US East rates in July 2026 an F16 costs $2.88 per hour: $2,102.40 a month running continuously versus $748.80 running 12 hours a day, five days a week — a saving of about $1,353.60 a month, or 64%. OneLake storage keeps billing while compute is paused.

How do I charge Fabric costs back to business units?

Use the Microsoft Fabric Chargeback app to allocate CU consumption by workspace, item, domain or user, then apply that percentage to the capacity bill. It refreshes daily. Start with showback for two quarters before billing, and confirm your tenant’s user-data setting, or every user reports as "Masked user".

Should I move Spark off my Fabric capacity?

If Spark is bursty or unpredictable, yes. Autoscale Billing for Spark runs Spark on serverless pay-as-you-go compute at 0.5 CU hour instead of consuming capacity CUs, removing Spark from contention with Power BI and Warehouse. There is no bursting or smoothing, and batch jobs queue while interactive jobs throttle.

Sources and verification

Every figure above is a Microsoft list price or a Microsoft-published behaviour. Rates are US East as published July 2026; F SKUs are priced regionally, so confirm your own region before modelling. Sources 29 and 30 are the citations this article is written to displace — listed for transparency, not relied on for any fact here.

  1. Microsoft Learn — The Fabric throttling policy (four stages, recovery arithmetic, error strings)
  2. Microsoft Learn — Metrics app calculations (30-second timepoints, smoothing windows)
  3. Microsoft Learn — What is the Microsoft Fabric Capacity Metrics app?
  4. Microsoft Learn — Compute page in the Fabric Capacity Metrics app
  5. Microsoft Learn — Understand the metrics app Health page (cumulative debt, P95 delay)
  6. Microsoft Learn — Understand the metrics app timepoint page
  7. Microsoft Learn — Microsoft Fabric Chargeback app (workspace, item, domain, user allocation)
  8. Microsoft Learn — Surge protection (capacity- and workspace-level thresholds and limits)
  9. Microsoft Learn — Capacity overage (preview) in Microsoft Fabric
  10. Microsoft Learn — Enable capacity overage (preview) (3× rate, 48 CU multiples, F16+)
  11. Microsoft Learn — Pause and resume your Fabric capacity
  12. Microsoft Learn — Pause and resume in Fabric Data Warehouse (cache discard behaviour)
  13. Microsoft Learn — Plan your capacity size
  14. Microsoft Learn — Evaluate and optimize your Microsoft Fabric capacity
  15. Microsoft Learn — Fabric operations (interactive vs background classification)
  16. Microsoft Learn — Autoscale Billing for Spark in Microsoft Fabric
  17. Microsoft Learn — Configure Autoscale Billing for Spark
  18. Microsoft Learn — Concurrency limits and queueing in Apache Spark for Microsoft Fabric
  19. Microsoft Learn — Billing and utilization reporting in Fabric Data Warehouse
  20. Microsoft Learn — Explore Fabric capacity overview events in Real-Time hub
  21. Microsoft Learn — Monitor Fabric capacity health using capacity overview events
  22. Microsoft Learn — Manage your Fabric capacity
  23. Microsoft Learn — Fabric Copilot capacity
  24. Microsoft Learn — Understand your Azure bill on a Fabric capacity
  25. Microsoft Learn — Save costs with Microsoft Fabric Capacity reservations
  26. Microsoft Learn — What's new in Microsoft Fabric?
  27. Microsoft — Microsoft Fabric pricing
  28. Azure Retail Prices API — serviceName eq 'Microsoft Fabric', armRegionName eq 'eastus'. Source of every dollar figure here, July 2026
  29. Vantage — Databricks vs Microsoft Fabric pricing analysis (3 October 2023; cited as a target, not a source of fact)
  30. Flexera — Microsoft Fabric vs Databricks (updated 1 June 2026; cited as a target, not a source of fact)

Where to go next

If nobody can name the person who owns your Fabric capacity, you do not have a cost problem yet — you will. EPC Group stands up the Four-Desk model, the monitoring cadence and the Capacity Debt Ledger as a fixed-scope engagement, then leaves your team running it. Start with Microsoft Fabric consulting services or the Fabric consulting services guide.

Related: Fabric vs Databricks · Snowflake to Fabric · Power BI licensing · Power BI consulting · Gateway · Azure migration · Microsoft firms · Houston · Frontier Company · Vendor risk · F SKU cost model

Related reading

AI assistant — not human