Reliability

Sales automation reliability: what happens when something breaks

Omnia Lab is a sales operations system that runs on top of your CRM — Bitrix24, amoCRM or any other. Failures resolve in the customer's favor: a broken check still sends the lead to a rep, a dead AI provider fails over to a backup, reminders fire on any budget.

· · Author: Dmitry Marenich

Automation rarely fails loudly. A messenger integration dies, leads stop arriving, and the team finds out a day later from an angry customer. Omnia Lab runs an inquiry from first contact to payment on top of your existing CRM, and that whole path rests on one condition: no message and no call ever falls out of the flow. So a failure surfaces before the customer feels it: every 30 minutes the platform pulls back messages that never arrived, more than three hours of silence during work hours raises an alert, AI provider status is checked every 3 minutes. And when something does break, the failure resolves in the customer's favor — the lead still gets assigned, the conversation stays visible, a check switches off instead of hiding data.

What happens to a lead when a check fails?

Most automation stops on error. It could not verify, so it did not pass anything through. In sales that is the worst possible choice. A real person is waiting while their inquiry sits in quarantine behind a broken check. We do the opposite. If we cannot tell whether an inquiry is qualified, it goes to a rep. If the marketplace boilerplate filter fails, the message counts as a live one. Every failure pushes data forward instead of burying it.

What stands between a failure and a lost lead?

Each of these runs on its own clock and does not wait for anyone to notice trouble.

  1. 01

    Every 30 minutes: message recovery

    The platform re-checks messengers and open channels and picks up whatever never arrived. It also syncs the "replied" status both ways. A counter on the dashboard shows how many messages were recovered.

  2. 02

    Three hours of silence: alert

    A flow watchdog tracks whether customer messages are arriving at all. More than three hours with none during work hours triggers a signal. A broken integration gets caught before the sales floor notices.

  3. 03

    Every 3 minutes: AI provider checks

    Plus platform uptime and an error log. Planned maintenance does not count as downtime, so the number stays honest.

  4. 04

    Two transcription services

    They cover each other: if one is unavailable, the call goes through the other. A call goes unanalyzed only if both die at once.

  5. 05

    A backup AI provider

    Switching is automatic and the reason shows up under Errors. Work keeps going, and afterwards you can see exactly what happened.

  6. 06

    Every 2 minutes: phone system polling

    For PBXs that never send call-end events. A missed call still becomes a Call back task: due in 30 minutes, or — if the call came in overnight — 30 minutes after the workday starts.

  7. 07

    Every minute: stuck deals

    New deals parked on a service account are found and handed to whoever is on shift.

What timings and limits does the platform run on?

Message recovery
every 30 minutes
Flow watchdog threshold
3 hours of silence during work hours
AI provider checks
every 3 minutes
Daily AI budget
$15 by default, alerts at 50% and 100%
Close reviews
20 per hour, the rest is queued
Reminder lifetime
never longer than 48 hours
Analytics refresh
hourly at :25, full reconciliation once a day
Data depth in the panel
90 days
Long-deal memory consolidation
Sunday, 23:30

What happens when the AI budget runs out?

A daily spend cap keeps one strange day from blowing up the bill. But a cap must never gag a live conversation, so processes fall into two classes.

  • The daily AI ceiling is $15 by default and adjustable in settings. Alerts fire at 50% and at 100%.
  • When the ceiling is hit, only deferrable work pauses: nightly deal analysis, weekly analytics, close reviews.
  • Customer-facing loops never pause. Lead routing, response-time escalations, commitment reminders and callbacks run on any budget.
  • Close reviews are capped at 20 per hour. Anything over the cap is queued, not dropped, and processed later.
  • During a funnel cleanup — more than 30 closures in 2 hours — deal returns go quiet. The system does not argue with a deliberate human decision.

Why do leads get lost in the first place?

  • Nobody watches the channel itself. The webhook dies, no messages arrive, and the quiet gets read as a slow day. Hence the three-hour flow watchdog.
  • Everything rides on one vendor. One transcription service, one AI provider, and any outage stops the work. We run two of each with automatic switching.
  • A failed check blocks. The classifier could not decide, so the inquiry goes to quarantine, the customer waits and no rep knows. We do the reverse.
  • Recovery with no trace. Something gets re-fetched somewhere, but nobody knows how much or what. Our recovered-message counter sits on the dashboard.
  • One spend cap stops everything. The AI budget runs out and customer reminders stop with it. We split deferrable work from customer-facing work.
  • Reminders pile up for weeks into a feed nobody reads. Here a new reminder physically deletes the previous one, and only after it has been sent successfully.

The cost of those lost hours is known: replying within an hour produces 7 times more qualified leads than replying an hour later, and 60 times more than a day later (Harvard Business Review, "The Short Life of Online Sales Leads", 2011, across 2,241 companies). A broken channel eats those hours quietly.

What does the system deliberately not do?

  • It does not hide data when a check fails. The filter switches off and the data flows through.
  • It does not delete the old reminder until the new one is delivered. The chat is never left empty.
  • It does not overwrite deal card comments. They are only appended to.
  • It does not count planned maintenance as downtime. There is no point in flattering the uptime number.
  • It does not hit your CRM with live queries for reports. The panel runs on its own data store and leaves the portal alone.
  • It does not roll nighttime escalations over to the morning. They are cancelled, so nobody starts the day under an avalanche of pings.

What results do current clients see?

Deals with a prepayment
from 4—7% to 13%
Median first response time
from 32 to 6.5 minutes
Revenue
+33% in 4 months without adding headcount
Rollout
3—5 days
Free audit
3—5 days

Measured before and after rollout at current clients, one of them a premium car dealer. Reliability is not abstract here: numbers like these hold only while no inquiry ever falls out of the flow.

Частые вопросы

What happens if the messenger integration dies overnight?
Messages that never arrived are picked up at the next recovery pass, which runs every 30 minutes. Nighttime escalations are cancelled rather than stacked, so nobody wakes up to a wall of pings. Overnight leads arrive as a single digest at 10:05. If the channel stays silent for more than three working hours, an alert fires.
How would I even know messages were being lost?
The dashboard has a dedicated counter for recovered messages, so you can see how many the platform pulled back for you. Every failure and provider switch is written to the error log with its reason. Nothing has to be dug out of raw logs by hand.
What happens when an AI provider goes down?
A backup provider takes over automatically and the reason appears under Errors. Transcription works the same way: two services cover each other, and analysis fails only if both go down at once. Lead routing and reminders keep running through all of it: when an AI check cannot complete, fail-open takes over — the lead goes to a rep and the reminder still arrives.
Can the system go quiet because the AI budget ran out?
Customer-facing loops never pause — leads get routed, escalations run, reminders arrive. When the daily ceiling is reached, only heavy analysis and reporting are deferred. You also get warned in advance, at 50% and at 100% of the cap.
Will the analytics panel overload our CRM?
No. The panel reads from a separate data store that refreshes hourly at :25 and fully reconciles with the CRM once a day, with 90 days of depth. Your portal keeps running as usual.
Can we switch off a module that gets in the way?
Yes. 33 modules have their own toggles, so any function can be turned off. There are 85 settings in the panel, they apply within 60 seconds and need no restart. You never have to wait on us for that.

Read next

Start with a free audit

In 3—5 days we map your inquiry flow: where messages break between the messenger and the CRM, how long the first reply actually takes, what quietly disappears in between. You get the findings as numbers, with no obligation, and you decide what to do next. If we move ahead, rollout takes 3—5 days and runs on top of your existing CRM.