If SharePoint is down right now
Do these while you are contacting us. Each one either narrows the problem or preserves evidence we will need, and none of them makes anything worse.
- Check the Microsoft 365 Service Health dashboard first — if Microsoft has an open SharePoint advisory, the fix is not on your side and the incident number matters more than any change you make.
- Identify whether the failure is one site collection or tenant-wide. Open a second, unrelated site: if it loads, you have a scoped problem, which is far better news.
- Stop making permission changes. Broken inheritance is the most common root cause and further edits make the original state harder to reconstruct.
- Check the site collection storage quota. A site that has hit its quota fails in ways that look like corruption but are not.
- If content vanished rather than erred, check the site Recycle Bin and then the second-stage Recycle Bin before assuming data loss.
What we triage in the first hour
SharePoint emergencies cluster into a small number of failure modes. These are the ones we rule in or out first, in roughly this order:
- Site and site-collection outages — whether the failure is Microsoft-side, tenant-side, or a single site
- Permission inheritance breaks, including the "Everyone except external users" over-share that surfaces during Copilot rollouts
- Search index failures where content exists but cannot be found, which reads to users as data loss
- Storage and quota lockups on site collections that have silently filled
- Workflow and Power Automate failures against SharePoint lists after a credential or connector change
How same-day engagement works
- You call or submit the form. 24×7 intake — it reaches a person at any hour, including weekends.
- A senior architect scopes the incident with you directly. Not an account manager, not a triage tier. The person asking the questions is the person who will work the problem.
- Remote engagement can begin the same day. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
Who is actually on the call
EPC Group has been a Microsoft-only consultancy since 1997 — 11,000+ enterprise engagements, six Microsoft Solutions Partner designations, and 216+ M&A tenant migrations covering 1.83 million users. We staff senior architects only and we are US-based. There is no offshore first line and no junior tier learning on your outage.
We work on environments other firms built, which is most of this work. We will also tell you when the fastest fix is something your existing partner should do — the first hour is about restoring service, not about scope.
After the fire
Restoring service is the first job, not the whole job. Once you are stable you get a written root-cause summary and a short list of the changes that would stop a repeat — usually a governance gap, a permissions structure, or missing alerting. Whether we implement those is entirely your call, and plenty of clients take the list and do it in house. If you would rather it be somebody’s standing job, that is what managed services is for.
Frequently asked questions
Is same-day response real, or is that marketing?
It is real and it is deliberately worded precisely: we operate 24×7 intake and same-day response. That means your call or form submission reaches a person at any hour, and a senior architect engages the same day. We do not publish a callback-minutes number, because a number nobody can guarantee at 3am is worse than no number at all.
My SharePoint site is down — what do you check first?
Whether the outage is yours. The first check is the Microsoft 365 Service Health dashboard, because a live Microsoft advisory changes the entire response: there is nothing to fix on your side and the job becomes communication and workaround. If Microsoft is healthy, we scope the blast radius — one site, one site collection, or the tenant — before touching anything.
Users lost access after a permissions change. Can it be undone?
Usually, and the most important thing is to stop editing permissions immediately. Every further change makes the original state harder to reconstruct. We capture the current state first, identify what inheritance was broken and where, then restore access at the group level rather than re-granting user by user, which is what created the sprawl in the first place.
Content has disappeared from a library. Is it recoverable?
Very often yes. The order is: site Recycle Bin, then the second-stage site collection Recycle Bin, then version history on the parent library, then a restore of the library to a previous point in time. Files deleted by a sync client behave differently from files deleted in the browser, so the recovery path depends on how they were removed.
Will you help with an environment another firm built?
Yes, and that is most emergency work. We do not require that we built it, we do not require you to switch providers afterwards, and we will tell you plainly if the fastest fix is something your existing partner should do. The first hour is about restoring service, not about scope.
What should I have ready when I call?
Global Administrator or equivalent access, the approximate time the problem started, what changed in the previous 48 hours, the exact error text or a screenshot, and the number of users affected. That last change is the single most useful fact — most incidents trace to a change, not to a spontaneous failure.
Is this remote or onsite?
Remote by default, because remote starts immediately and almost every Microsoft cloud incident is resolved that way. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
What does emergency support cost?
We scope each incident at intake rather than publishing an emergency rate. Cost depends on severity, how much of the estate is affected, whether after-hours work is needed, and how long remediation runs. You get a scope and a number before work starts — nobody signs a blank cheque during an outage.
What happens after the incident is resolved?
You get a written root-cause summary and a short list of the hardening changes that would prevent a repeat — usually governance, permissions structure, or an alerting gap. Whether we implement those is your call, and many clients take the list and do it in house.
SharePoint down right now?
(888) 381-972524×7 intake · same-day response · or submit an emergency request.
Related
All emergency Microsoft support · SharePoint consulting · SharePoint Governance Health Check · SharePoint migration best practices
