If Power BI is down right now
Do these while you are contacting us. Each one either narrows the problem or preserves evidence we will need, and none of them makes anything worse.
- Open the Fabric Capacity Metrics app and read the Throttling tab, not Utilization. A capacity can sit at 95% and still reject requests, because throttling is driven by smoothed future consumption.
- Check whether the failure is one report or every report on the workspace. One report is almost always a semantic model or credential problem, not a capacity problem.
- Look at the gateway status before anything else if refreshes are failing. An offline or unreauthorized gateway produces failures that look like data-source outages.
- Check whether anything was published or changed in the last 24 hours. Most "sudden" Power BI failures follow a deployment.
- If users cannot view content at all, confirm the capacity SKU — below F64 every viewer needs a Pro or PPU licence, and losing free-viewer access looks exactly like an outage.
What we triage in the first hour
Power BI emergencies cluster into a small number of failure modes. These are the ones we rule in or out first, in roughly this order:
- Gateway connectivity and credential reauthorization, the most common cause of a refresh that worked yesterday
- Scheduled refresh failures — timeout, source change, or a credential rotated without the gateway being updated
- Capacity throttling and the recovery arithmetic that tells you how long the incident lasts before it self-clears
- Row-level security broken by a model change, where the wrong people see the right data or nobody sees anything
- Tenant-setting lockouts after an admin change that silently removed sharing or export rights
How same-day engagement works
- You call or submit the form. 24×7 intake — it reaches a person at any hour, including weekends.
- A senior architect scopes the incident with you directly. Not an account manager, not a triage tier. The person asking the questions is the person who will work the problem.
- Remote engagement can begin the same day. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
Who is actually on the call
EPC Group has been a Microsoft-only consultancy since 1997 — 11,000+ enterprise engagements, six Microsoft Solutions Partner designations, and 216+ M&A tenant migrations covering 1.83 million users. We staff senior architects only and we are US-based. There is no offshore first line and no junior tier learning on your outage.
We work on environments other firms built, which is most of this work. We will also tell you when the fastest fix is something your existing partner should do — the first hour is about restoring service, not about scope.
After the fire
Restoring service is the first job, not the whole job. Once you are stable you get a written root-cause summary and a short list of the changes that would stop a repeat — usually a governance gap, a permissions structure, or missing alerting. Whether we implement those is entirely your call, and plenty of clients take the list and do it in house. If you would rather it be somebody’s standing job, that is what managed services is for.
Frequently asked questions
Is same-day response real, or is that marketing?
It is real and it is deliberately worded precisely: we operate 24×7 intake and same-day response. That means your call or form submission reaches a person at any hour, and a senior architect engages the same day. We do not publish a callback-minutes number, because a number nobody can guarantee at 3am is worse than no number at all.
My Power BI refresh has been failing since last night. Where do you start?
With the gateway and the credentials, because that is where most overnight refresh failures live — a rotated password, an expired OAuth token, or a gateway that went offline during a server reboot. If the gateway is healthy we move to the source: schema changes, a table rename, or a query that now exceeds its timeout.
Reports are slow but nothing is technically broken. Is that an emergency?
It can be, and the distinction matters for the fix. Slow reports are more often a semantic model problem than a capacity problem. We check whether CU utilization actually exceeded its limit during the window; if it did not, the fix is model design, not a bigger SKU — and buying a bigger SKU would have hidden the real problem at real cost.
Our capacity is throttled. How long until it recovers on its own?
It is calculable, not a guess. The minimum recovery is ((percentage above 100) ÷ 100) × the smoothing window for whichever rejection type is active. A 250% interactive rejection clears in roughly 90 minutes; a 250% background rejection takes about 36 hours. Pausing and resuming the capacity clears it immediately, which is often the right call.
Will you help with an environment another firm built?
Yes, and that is most emergency work. We do not require that we built it, we do not require you to switch providers afterwards, and we will tell you plainly if the fastest fix is something your existing partner should do. The first hour is about restoring service, not about scope.
What should I have ready when I call?
Global Administrator or equivalent access, the approximate time the problem started, what changed in the previous 48 hours, the exact error text or a screenshot, and the number of users affected. That last change is the single most useful fact — most incidents trace to a change, not to a spontaneous failure.
Is this remote or onsite?
Remote by default, because remote starts immediately and almost every Microsoft cloud incident is resolved that way. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
What does emergency support cost?
We scope each incident at intake rather than publishing an emergency rate. Cost depends on severity, how much of the estate is affected, whether after-hours work is needed, and how long remediation runs. You get a scope and a number before work starts — nobody signs a blank cheque during an outage.
Can you help if our Power BI estate was built by someone else?
Yes, and that is the majority of emergency Power BI work. We do not need to have built the semantic model to fix the refresh, and we will hand back a documented root cause whether or not you continue with us.
Power BI down right now?
(888) 381-972524×7 intake · same-day response · or submit an emergency request.
Related
All emergency Microsoft support · Power BI consulting · Fabric capacity FinOps and throttling · Power BI gateway guidance
