If Exchange is down right now
Do these while you are contacting us. Each one either narrows the problem or preserves evidence we will need, and none of them makes anything worse.
- Check the Microsoft 365 Service Health dashboard before anything else — mail flow incidents are one of the more common Microsoft-side advisories.
- Run a message trace on a specific failed message. The trace tells you where delivery stopped, which is worth more than any amount of guessing.
- Check whether a transport rule was added or edited recently. A mis-scoped rule can silently reroute or drop mail for an entire domain.
- Look at your outbound spam policy and any recent tenant restriction — outbound sending limits can quietly block a domain that started sending more volume.
- If this followed a migration or cutover, verify MX, SPF, DKIM and DMARC records resolve to what you expect from an external network, not from inside your own.
What we triage in the first hour
Exchange emergencies cluster into a small number of failure modes. These are the ones we rule in or out first, in roughly this order:
- Total mail-flow stoppage — inbound, outbound, or both, and whether the boundary is Microsoft, connector, or DNS
- Transport rule loops and mis-scoped rules that reroute or silently drop mail for a whole domain
- NDR storms and backscatter, including the sending-limit blocks that follow them
- Hybrid and connector failures between on-premises Exchange and Exchange Online
- Stuck or failed migration batches that leave mailboxes split between environments
How same-day engagement works
- You call or submit the form. 24×7 intake — it reaches a person at any hour, including weekends.
- A senior architect scopes the incident with you directly. Not an account manager, not a triage tier. The person asking the questions is the person who will work the problem.
- Remote engagement can begin the same day. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
Who is actually on the call
EPC Group has been a Microsoft-only consultancy since 1997 — 11,000+ enterprise engagements, six Microsoft Solutions Partner designations, and 216+ M&A tenant migrations covering 1.83 million users. We staff senior architects only and we are US-based. There is no offshore first line and no junior tier learning on your outage.
We work on environments other firms built, which is most of this work. We will also tell you when the fastest fix is something your existing partner should do — the first hour is about restoring service, not about scope.
After the fire
Restoring service is the first job, not the whole job. Once you are stable you get a written root-cause summary and a short list of the changes that would stop a repeat — usually a governance gap, a permissions structure, or missing alerting. Whether we implement those is entirely your call, and plenty of clients take the list and do it in house. If you would rather it be somebody’s standing job, that is what managed services is for.
Frequently asked questions
Is same-day response real, or is that marketing?
It is real and it is deliberately worded precisely: we operate 24×7 intake and same-day response. That means your call or form submission reaches a person at any hour, and a senior architect engages the same day. We do not publish a callback-minutes number, because a number nobody can guarantee at 3am is worse than no number at all.
Mail has stopped flowing entirely. What is the first thing you do?
A message trace on a specific failed message, because it tells you exactly where delivery stopped rather than where you assume it did. In parallel we check Service Health for a live Microsoft advisory. Those two facts together usually identify the boundary — Microsoft, a connector, a transport rule, or DNS — within the first fifteen minutes.
We are getting an NDR storm. How do you stop it?
First identify whether the source is a compromised mailbox, a mis-scoped transport rule, or a mailbox loop. The urgency is real: sustained volume triggers outbound sending restrictions, and a restricted tenant is a much longer problem to unwind than the original cause. Containment comes before diagnosis here.
Our migration cutover failed halfway. Can it be rolled back?
It depends where it stopped, and the honest answer is that partial cutovers are usually finished forward rather than rolled back. We establish which mailboxes are where, restore mail flow for everyone regardless of location, and then complete or reverse the batch deliberately rather than under time pressure.
Will you help with an environment another firm built?
Yes, and that is most emergency work. We do not require that we built it, we do not require you to switch providers afterwards, and we will tell you plainly if the fastest fix is something your existing partner should do. The first hour is about restoring service, not about scope.
What should I have ready when I call?
Global Administrator or equivalent access, the approximate time the problem started, what changed in the previous 48 hours, the exact error text or a screenshot, and the number of users affected. That last change is the single most useful fact — most incidents trace to a change, not to a spontaneous failure.
Is this remote or onsite?
Remote by default, because remote starts immediately and almost every Microsoft cloud incident is resolved that way. Onsite is scheduled when the problem is genuinely physical — on-premises hardware, network equipment, or a site that cannot grant remote access.
What does emergency support cost?
We scope each incident at intake rather than publishing an emergency rate. Cost depends on severity, how much of the estate is affected, whether after-hours work is needed, and how long remediation runs. You get a scope and a number before work starts — nobody signs a blank cheque during an outage.
Do you support hybrid Exchange, not just Exchange Online?
Yes. Hybrid is where a large share of emergency Exchange work happens, because the failure modes live in the connectors, certificates and DNS between the two environments rather than in either one alone.
Exchange down right now?
(888) 381-972524×7 intake · same-day response · or submit an emergency request.
Related
All emergency Microsoft support · Exchange support services · Exchange migration · Microsoft 365 consulting
