If Microsoft 365 is down right now
Do these while you are contacting us. Each one either narrows the problem or preserves evidence we will need, and none of them makes anything worse.
- Check the Microsoft 365 Service Health dashboard and note the incident ID if one exists — it changes both the fix and what you tell your users.
- Establish whether the problem is authentication or a single service. If sign-in itself fails, the investigation starts at Entra ID and Conditional Access, not at the app people are complaining about.
- Review Conditional Access policy changes in the last 48 hours. A policy applied without an exclusion is the most common self-inflicted tenant lockout.
- Check whether a licence assignment or subscription lapsed. Expired licences fail in ways that look like outages.
- Confirm at least one break-glass administrator account can still sign in, and do not lose that access while troubleshooting.
What we triage in the first hour
Microsoft 365 emergencies cluster into a small number of failure modes. These are the ones we rule in or out first, in roughly this order:
- Tenant-wide authentication failures and Conditional Access lockouts, including the self-inflicted policy change
- Whether the incident is Microsoft-side — a live advisory changes the response from fixing to communicating
- Licensing and subscription lapses presenting as service failures
- Directory synchronisation failures between on-premises Active Directory and Entra ID
- Multi-service failures that share one root cause, usually identity, rather than several coincidental ones
How same-day engagement works
- You call or submit the form. 24×7 intake — it reaches a person at any hour, including weekends.
- A senior architect scopes the incident with you directly. Not an account manager, not a triage tier. The person asking the questions is the person who will work the problem.
- Remote engagement can begin the same day. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
Who is actually on the call
EPC Group has been a Microsoft-only consultancy since 1997 — 11,000+ enterprise engagements, six Microsoft Solutions Partner designations, and 216+ M&A tenant migrations covering 1.83 million users. We staff senior architects only and we are US-based. There is no offshore first line and no junior tier learning on your outage.
We work on environments other firms built, which is most of this work. We will also tell you when the fastest fix is something your existing partner should do — the first hour is about restoring service, not about scope.
After the fire
Restoring service is the first job, not the whole job. Once you are stable you get a written root-cause summary and a short list of the changes that would stop a repeat — usually a governance gap, a permissions structure, or missing alerting. Whether we implement those is entirely your call, and plenty of clients take the list and do it in house. If you would rather it be somebody’s standing job, that is what managed services is for.
Frequently asked questions
Is same-day response real, or is that marketing?
It is real and it is deliberately worded precisely: we operate 24×7 intake and same-day response. That means your call or form submission reaches a person at any hour, and a senior architect engages the same day. We do not publish a callback-minutes number, because a number nobody can guarantee at 3am is worse than no number at all.
Nobody can sign in to Microsoft 365. Where do you start?
At identity, not at the application people are reporting. A tenant-wide sign-in failure is almost always Entra ID — a Conditional Access policy applied without an exclusion, an expired federation certificate, or an MFA configuration change. We also confirm immediately that a break-glass account still works, because losing that during troubleshooting turns a bad hour into a much longer one.
We locked ourselves out with a Conditional Access policy. Is that recoverable?
Yes, and this is one of the more common calls. Recovery runs through a break-glass account excluded from Conditional Access, which is exactly why that account should exist before you need it. If no such account exists the path is longer and involves Microsoft support, which is why establishing one is the first hardening recommendation afterwards.
How do we tell whether this is our problem or Microsoft’s?
The Service Health dashboard is the first answer, but it lags real incidents. The faster signal is scope: if a failure crosses unrelated services and multiple tenants you have contact with, it is Microsoft. If it is confined to your tenant and started after a change, it is yours — and it is almost always the change.
Will you help with an environment another firm built?
Yes, and that is most emergency work. We do not require that we built it, we do not require you to switch providers afterwards, and we will tell you plainly if the fastest fix is something your existing partner should do. The first hour is about restoring service, not about scope.
What should I have ready when I call?
Global Administrator or equivalent access, the approximate time the problem started, what changed in the previous 48 hours, the exact error text or a screenshot, and the number of users affected. That last change is the single most useful fact — most incidents trace to a change, not to a spontaneous failure.
Is this remote or onsite?
Remote by default, because remote starts immediately and almost every Microsoft cloud incident is resolved that way. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
What does emergency support cost?
We scope each incident at intake rather than publishing an emergency rate. Cost depends on severity, how much of the estate is affected, whether after-hours work is needed, and how long remediation runs. You get a scope and a number before work starts — nobody signs a blank cheque during an outage.
Do you cover the whole tenant or just one workload?
The whole tenant. Multi-service failures usually share one root cause in identity or policy, and treating them as separate incidents is how organisations spend a day fixing three symptoms of the same problem.
Microsoft 365 down right now?
(888) 381-972524×7 intake · same-day response · or submit an emergency request.
Related
All emergency Microsoft support · Microsoft 365 consulting · Microsoft 365 security hardening checklist · Microsoft Entra ID enterprise guide
