Meridian sells fleet maintenance software to small trucking outfits. The program I built this for is their inbound demo booking desk, fourteen bilingual reps whose whole job is one call. Qualify the lead, book a demo with an Account Executive. They don't close and they don't quote.
The number leadership watches is demos booked. You can game that number in about four seconds. Book everyone. Let the AE find out on the call that the prospect runs nine vehicles against a fifteen vehicle floor and has no budget. The rep's dashboard looks great, the pipeline fills up with junk, and nobody notices for two weeks.
So the first thing I decided was that the grid had to be allowed to disagree with the dashboard. A rep can book a demo and still score badly. A rep can turn a prospect away and score a 96. If the grid can't produce that result, you've just measured booking rate twice and given it a fancier name.
Five pillars, eighteen criteria, everything scored 1 to 5. Discovery carries the most weight, and that's the part doing the pushing back:
| Pillar | Weight | Why it sits there |
|---|---|---|
| A. Discovery & Qualification | 30% | The thing the headline metric misses, so it gets the biggest share. |
| B. Sales Execution | 25% | Value framing, honest product talk, actually asking for the demo. |
| C. Compliance & Documentation | 20% | Disclosure, consent, and CRM notes the AE can work from. |
| D. Communication & Rapport | 15% | Matters, but it's the easiest thing to score generously and the least predictive. |
| E. Next-Step Integrity | 10% | Scored on its own, separately from whether a demo got booked. |
Communication sits low for a reason. It's the pillar evaluators inflate, because a warm confident rep sounds like a good rep. There's a rep in the sample data built to expose that. Charming, high D scores, discovery so thin his qualified leads keep falling apart later.
Six triggers drop the total to zero. People argue with this one, so here's my defence. Most calls on this desk carry no compliance risk at all, and a handful carry a lot of it. Average those together and you end up with a nice looking 84 that quietly contains a consent violation. I would rather the number be ugly and honest.
| Code | Trigger |
|---|---|
| AF-01 | Substantive discussion before the recording disclosure |
| AF-02 | Quoting a rate, term or discount the rep isn't authorised to offer |
| AF-03 | Presenting a roadmap feature as if it ships today |
| AF-04 | Continuing after an opt-out, or not logging a consent withdrawal |
| AF-05 | Falsified disposition, so inflating booked demos or burying a refusal |
| AF-06 | Disclosing another client's data, volumes or pricing |
Two of those exist only because the desk isn't allowed to close. A rep who invents a discount is doing the AE's job badly, and it's the most tempting thing in the world to do when you can feel a deal slipping away from you.
live_weight is the interesting part. Objection handling doesn't come up
on every call, sometimes the prospect just says yes. Scoring that a zero punishes
the rep for something the prospect did, so criteria marked
NA fall out completely and the remaining pillars rescale
to 100. Small thing, and it removes a huge amount of arguing from the feedback
conversation.
Two calls from the sample month. I picked these two because telling them apart is the entire reason the framework exists.
The first one booked a demo.
| Pillar | Raw | Weighted |
|---|---|---|
| A. Discovery & Qualification | 16/25 | 19.2 |
| B. Sales Execution | 10/20 | 12.5 |
| C. Compliance & Documentation | 12/20 | 12.0 |
| D. Communication & Rapport | 12/15 | 12.0 |
| E. Next-Step Integrity | 5/10 | 5.0 |
| Total | 0, AUTO-FAIL |
On paper that's a 60.7. It scores zero, on two triggers, and they feed each other. At 04:12 the rep answered price pressure with "for a fleet your size we can probably do around fifteen percent off", which is a discount the desk has zero authority to hand out AF-02. Then dispositioned the lead as Demo Booked, Qualified, when the prospect had eleven vehicles against a fifteen vehicle floor and had mentioned twice that budget sat with a parent company nobody had talked to yet AF-05.
The coaching note on that one doesn't open with the discount. It opens with the qualification, because the discount got invented to rescue a deal that never should have been qualified in the first place. Fix the qualification and the discount problem mostly goes away on its own.
The second one turned the prospect away.
| Pillar | Raw | Weighted |
|---|---|---|
| A. Discovery & Qualification | 25/25 | 30.0 |
| B. Sales Execution | 14/15 | 23.3 |
| C. Compliance & Documentation | 20/20 | 20.0 |
| D. Communication & Rapport | 14/15 | 14.0 |
| E. Next-Step Integrity | 10/10 | 10.0 |
| Total | 97.3, Exemplary |
Outcome on that call: disqualified. The rep ran full discovery, got the pain down to an actual number (fourteen unplanned roadside events last quarter, roughly $2,300 each once you count the customer penalty), found a real deadline in a failed CVOR audit with an October re-inspection, asked for the demo, and then took it back once the fleet turned out to sit under the floor. Logged the October re-inspection as a nurture trigger on the way out the door.
B is 14/15 instead of 15/15 only because the booking slot never ended up being needed. Watch the denominator there too. Objection handling was NA, so the pillar is out of 15 and the weights rescaled around it. Best call on the desk that month, and it booked nothing.
Scoring calls one at a time is the input. The output is what happens when you roll twenty evaluations across six reps together:
| Rep | Avg | Auto-fails | Weakest pillar |
|---|---|---|---|
| L. Fontaine | 42.7 | 1 | E, Next-Step Integrity (52%) |
| J. Aliyev | 61.5 | – | C, Compliance (47%) |
| M. Tremblay | 76.7 | – | A, Discovery (59%) |
| K. Persaud | 85.8 | – | D, Communication (76%) |
| R. Okonkwo | 90.8 | – | C, Compliance (85%) |
| T. Nakamura | 96.8 | – | D, Communication (95%) |
Tremblay is the charmer. 76.7 overall, discovery sitting at 59%. Listen to one of his calls and he sounds fine, which is the whole argument for putting the weight where I put it.
First time I ran the numbers, the outlier detection flagged Fontaine and nobody else. Aliyev at 61.5, compliance at 47%, clearly struggling, went straight through unflagged.
The auto-fail did it. That zero pulled Fontaine's average down to 42.7 and pushed the standard deviation out to 20.3, and once the spread is that wide, a threshold of one deviation below the mean catches basically nobody. The failing rep was hiding the weak one.
Fix was to run the outlier baseline on scored calls only. The zeros stay in the reported average where the client should absolutely see them, they just come out of the distribution, since they measure a compliance event rather than how the rep is performing. Deviation drops to 16.1, both reps get flagged, and they get two different conversations.
Monthly, three calls, every evaluator scores blind. Tolerance is 5 points on the total and 1 point on any single criterion. The sample set has a call where four evaluators landed 10.1 points apart, well out of tolerance, and the split traces back to two criteria:
| Criterion | Range | The disagreement |
|---|---|---|
| A4 Authority & decision process | 2–4 | Prospect volunteered "I'd loop in ops". Does that count as authority surfaced, or does it only count once the rep gets a name and maps the approval path? |
| B1 Value framing tied to stated pain | 3–5 | Three capabilities named, one of them tied back to something the prospect said. Two evaluators scored the energy, two scored the tie-back. |
Nobody was being careless there. Both criteria were worded loosely enough to support two honest readings, so my working rule is that a split past tolerance counts as a defect in the grid until proven otherwise. Rewrite the wording before you correct the evaluator. A4 now reads "mapped who else signs off and what the internal approval path looks like", which makes "I'd loop in ops" a clean 2.
Get this part wrong and scores stop being comparable between evaluators, which quietly poisons every trend you build on top of them, including that outlier analysis further up the page.
One Python file, standard library, no dependencies. The grid lives in JSON as the single source of truth and the readable grid document gets generated from it, so the document can't drift away from the math the way a Word file always does eventually.
$ python3 scoring.py scored 24 evaluations (20 live, 4 calibration) out/grid.md, out/summary.md, out/roster.csv, out/evaluations/*.mdFour outputs pointed at four different readers. A rendered grid for the client, a per-call feedback report written so it can go straight to the rep, a flat CSV with every criterion as its own column for whoever wants to pivot it, and a summary carrying rep trends, desk-wide pillar weakness, the five weakest criteria and the calibration spread.
That last table is the one that changes what anybody does on Monday. Across the sample month the weakest criteria are CRM notes accuracy at 3.60 and authority mapping at 3.70, both of them desk-wide, both sitting in the process rather than in any one person. Coaching six reps individually won't move either number. Training and workflow will, and pointing at the right one is most of this job.