Quality scoring

Scoring your reps can check line by line

Omnia Lab is a sales-team control system that runs on top of your CRM — Bitrix24, amoCRM or any other. Each score comes with the reasoning: you can see where the points went, an unqualified buyer is not held against the rep, and any score can be contested.

· · Author: Dmitry Marenich

Omnia Lab is an operating system for a sales team: it runs every lead from arrival to revenue on top of your CRM, and rep scoring is one part of it. Most disputes about a score come down to one thing: the rep cannot see where the number came from. We split facts from arithmetic. The AI only finds what was actually said and backs it with a quote. Code computes the score under fixed rules: the outcome sets the band, process quality sets the position inside it, and a critical error caps the score — and that error counts only when there is a verbatim quote to show. The result is a score you can walk a rep through line by line, instead of pulling rank to defend it.

Why do reps argue with their scores?

Almost every argument grows out of six situations. Name them before you fix them.

  • "The AI decided" — and nobody can explain how. A number without reasoning reads as a verdict, and the first question is always why not an 80.
  • The same mistake costs different amounts on different calls. If points are deducted by feel, a rep finds two near-identical calls with different scores in an afternoon, and the quality conversation is over.
  • A rep gets marked down for a buyer nobody could close. A reseller, a kid, someone whose budget is half of what they asked about.
  • Haggling and "that's expensive" get read as "no money." The rep heard a real buyer; the system filed them as hopeless.
  • The review lands on the wrong person. A colleague picked up the phone, the task went to the deal owner.
  • There is nowhere to appeal. Without a way for a manager to say "I looked, this one is fine," scoring turns into background noise.

How is a call score calculated?

Seven steps, identical on every call. On none of them does the model decide the number.

  1. 01

    The AI finds facts, it does not grade

    The model reads the transcript and records what actually happened: the need was explored, a next step was named, a time was agreed. Every finding carries a quote. The model never outputs a score.

  2. 02

    The outcome sets the band

    First the system looks at how the call ended. The outcome picks the band: 90-100, 80-89, 65-79, 40-64 or 1-39. A smooth call that led nowhere does not reach the top band.

  3. 03

    Process sets the position inside the band

    Within the band, placement comes from how the work was done: discovery, moving toward a next step, closing. The gap between 66 and 78 is a process gap.

  4. 04

    A critical error caps, it does not deduct

    A serious violation does not shave off some arbitrary 15 points. It sets a ceiling the call cannot rise above, however good the rest was. The argument about how many points were taken disappears with the deduction.

  5. 05

    No quote, no error

    A critical error counts only with a verbatim quote from the conversation. If the model senses a violation but cannot show the line, the finding is dropped entirely.

  6. 06

    Compensation: late still counts

    A skipped step that gets closed later, or closed by the outcome itself, earns full credit. The name was asked at the very end but the sale happened — the step is credited, not flagged.

  7. 07

    Company rules sit in their own layer

    Banned phrasing and weak closings are defined by the company. Each rule carries a severity: "critical" caps the score at 64, "signal" is a note in the review.

How do you tell an unqualified buyer from weak selling?

Holding a rep to the same standard on a retail buyer and on a reseller is not fair. So every conversation gets a client profile, with a confidence level and quotes.

  • Eight segments: retail buyer, budget nowhere near what they are asking for, reseller, broker, kid or prank call, scammer, service-only, unclear.
  • Failing to push an objectively unqualified buyer is not counted as a mistake. Politeness is still scored for everyone, every time — a prank call does not excuse rudeness.
  • Skepticism is the default: a mitigating segment counts only at 75% confidence or higher and with a direct quote. Doubt works against the excuse, otherwise every miss gets written off as a difficult buyer.
  • Haggling, "that's expensive," asking for a discount, short replies and silence are never treated as proof that a buyer cannot pay. That is ordinary buyer behavior.
  • A scammer is left out of the stats on how hard reps push a deal, but is always surfaced to the manager and escalated. They do not vanish into a pile of unqualified leads.
  • The low-budget threshold is set by the company, with its own policy. What is small change for one team is a normal deal for another.

What thresholds and defaults does it ship with?

Out-of-the-box values. What you configure belongs to your company: its conversation rules, thresholds and module toggles. The bands and the formula are computed by code the same way for everyone — that is the point.

Score bands
90-100, 80-89, 65-79, 40-64, 1-39 — the band comes from the call outcome
Critical error
caps the score instead of deducting; company rules marked "critical" cap at 64
Mitigating client segment
counts at ≥75% confidence with a direct quote
Daily rep score
0.6 × conversation + 0.4 × process
Rep score vs deal score
separate metrics, separate columns in the dashboard
Call review task
always for new calls; for follow-ups only when the score is below 65
Comment on the deal timeline
when the rep axis is below 50 or a red flag fires, at most one per deal per day
Pause on new leads
after a call rated 50 or lower, 2 working hours; night does not burn the clock (8pm → until 11am the next day); a rep alone on shift keeps receiving leads; off by default
Methodology base
ISO 18295-1, COPC CX, BARS, SPIN, MEDDIC, plus Gong Labs and MIT research

What if a rep or manager disagrees with a score?

A score with no right of reply survives about a week. After that people stop reading it.

  1. 01

    The review shows the reasoning, not just a number

    The task is written in the rep's own terms: what to do now, the score and the logic behind it, what was strong and what was weak, a recommendation, a link to the methodology. You can go through it line by line.

  2. 02

    The score follows the voice on the call

    If someone other than the deal owner picked up, the score, the report and the task all go to the person who actually spoke.

  3. 03

    "Reviewed by the manager"

    A sales manager clears a deal off the radar with one of five reasons, in two classes. "Work has been done" holds until the first new risk appears. "Disagreeing with the system" holds until the risk gets worse, up to 14 days. The mark can be revoked, and everything lands in the audit trail.

  4. 04

    An "exclude" button on response-time measurement

    A technical outlier can be excluded with a button, and the record shows who did it and when. Quietly editing the stats is not an option.

  5. 05

    What is never scored at all

    Terminal and won stages. Reps flagged as do-not-ping, owners outside the directory, pipelines you did not put in scope. No scores appear there, so there is nothing to argue about.

What does the system deliberately not do?

  • It will not let the AI assign the score. The model finds and quotes, the code does the arithmetic — the same way for everyone, every time.
  • It will not deduct for a hunch. Without a verbatim quote, a critical error is not recorded.
  • It will not punish a rep for working out of order. A step closed later, or closed by the outcome itself, gets full credit.
  • It will not count a weak push on an unqualified buyer as a mistake — and it will not excuse rudeness for anyone.
  • It will not call anyone out in public. The daily review, one piece of praise and one growth area, goes to the rep privately; a deal-timeline comment appears at most once a day.
  • It will not hide the awkward parts. Scammers stay visible to the manager, and the "reviewed" mark mutes a signal while staying in the audit trail.

What changes in the numbers after rollout?

Measured before and after rollout at a client we work with today, a premium car dealership. No names.

Deals with a prepayment
from 4—7% to 13%
Median first response to a customer
from 32 minutes to 6.5
Revenue
+33% in 4 months, with no extra headcount
Rollout
3—5 days
Free audit
3—5 days

A score sells nothing by itself. It changes what the rep does on the next call, and that part shows up in the revenue.

Частые вопросы

My reps say the AI grades them however it feels. What do I tell them?
That code computes the score under fixed rules, and the AI only finds facts and backs them with quotes from the call. The outcome sets the band, process quality sets the position inside it. Any score can be broken down step by step and shown.
The buyer was hopeless. Why did the rep score low?
They shouldn't, if the chat shows it. The client profile is assessed separately, and failing to push an objectively unqualified buyer is not a mistake. But a mitigating segment counts only at 75% confidence or higher with a direct quote, and politeness is scored for everyone regardless.
The buyer haggled and said it was expensive. Doesn't that mean no budget?
No. Haggling, "that's expensive," a discount request, short replies and silence are never proof that someone cannot pay. It is normal buying behavior. Treat it as proof and every lost conversation becomes the buyer's fault.
What if the sales manager thinks a score is wrong?
They mark the deal "reviewed" with one of five reasons, in two classes. "Work has been done" holds until a new risk appears; "disagreeing with the system" holds until the risk worsens, up to 14 days. The mark is revocable and the whole thing is audited.
Does a bad score cost the rep anything?
By default, no — they get a review task with a recommendation. There is an optional setting that pauses new leads after a call rated 50 or lower, for 2 working hours, and night time does not burn that clock. If the rep is alone on shift, leads keep coming anyway.
So this is a call-review service?
No. Omnia Lab runs every inbound lead from arrival to revenue: it distributes by shift, watches first response time, remembers commitments on both sides, catches missed calls, checks whether closures were justified, fills in the CRM and shows the manager the real numbers. Call review is one of its signal sources. It runs on top of your CRM rather than replacing it.

Read next

Free sales audit

We take your real calls and chats and show how each one would be scored, and where that score would part ways with your own read. It takes 3—5 days and nothing in your CRM changes. You get the breakdown, not a pitch deck.