Your AI graded the call. That still isn't coaching.
Conversation intelligence made something real possible: every call can be scored. Talk ratios, objections, competitor mentions, methodology coverage. The library that nobody reopened became a dashboard somebody can open.
Then a lot of teams stopped there.
A grade landed. A summary landed. A manager nodded in pipeline review. The next call looked like the last one. Feedback arrived; behavior did not move. That gap is not a tooling bug. It is what happens when the workflow ends at information.
Coaching is what happens after the grade.
Feedback is not a coaching loop
Elay Cohen put the blunt version into circulation: AI can give sellers immediate feedback on a practice pitch, a role play, or a customer conversation. That is useful. Feedback alone does not change behavior.
The frontline manager is still the person who decides what gets reinforced this week, which skill gap matters most for this seller, what good looks like inside a real deal, how the team will practice it, and whether the behavior shows up again with buyers. The best model is not AI or manager coaching. It is the combination: AI creates a fast feedback loop; managers turn that feedback into accountability and improvement in the flow of work.
When the only new artifact after a call is a summary, you did not coach. You filed.
Gartner’s own framing of the manager problem makes the same point from the other side. Sales managers can be force multipliers, with up to 6x impact on seller performance, yet only about 18% of managers report leading high-performing teams. The role is often misaligned with what matters. Handing that role another stack of call scores without a loop does not close the gap. It fills a calendar with review theater.
The math that breaks human-only coaching
Jonathan Jones named the capacity problem out loud: a sales leader can face hundreds of calls a month. A great coach can truly work maybe 10 to 15 in a week — really listen, spot the moment, give feedback specific enough to land. The rest of the library sits unread, which is how teams ended up buying conversation intelligence in the first place.
So the industry automated listening. That fixed coverage. It did not fix coaching. Coverage without a loop produces managers who are drowning in evidence and still starving for time to act on it. Reps get a grade. Managers get a queue. Nobody gets a next practice.
If your AI grades 100% of calls and your managers coach 5% of them, you did not scale coaching. You scaled grading.
What has to happen after the score
A coaching loop has five moves. Miss one and you are back to filing.
Pick the gap that matters this week. Not every low score earns airtime. A weak Paper Process in early discovery is noise. A named Economic Buyer who has never attended a session with a proposal outstanding is the week’s work.
Define what good looks like on this deal. Generic advice (“ask better discovery questions”) dies on contact with a real opportunity. The standard has to name the buyer, the missing evidence, and the conversation that would produce it.
Practice. Role play, talk track, or a dry run of the opening. Feedback without rehearsal is a lecture.
Put the move on the next call. Coaching that does not change the next customer conversation was a meeting about coaching.
Check whether it showed up. The next transcript is the test. If the gap is still empty, the loop continues. If the evidence arrived, the score should be allowed to move — including down when evidence decays.
That fifth step is where most AI scorecards quietly fail. They report. They do not close the loop into the deal the rep is actually running.
A worked example of filing vs coaching
Filing: After a discovery call, the AI returns: talk ratio 42%, competitor mentioned once, MEDDPICC incomplete, overall call score 67. The manager skims it Friday. Pipeline review notes “needs coaching on discovery.” Nothing is scheduled. Next week’s call opens the same way.
Coaching: The same call produces a deal-level gap: Economic Buyer scored 3/15 — identified by name, never in a room. Recommended action: get them into the technical review this week. The manager’s 1:1 spends twelve minutes on that move only: how the champion introduces the EB, what question unlocks budget ownership, what “good” sounds like if the EB deflects. The rep practices the open once. The next meeting invite includes the EB. The following analysis either raises the score with evidence or keeps the gap flagged with a harder next step.
Same call. Different job for the software.
Where Ruby sits in the loop
Ruby’s MEDDPICC score is built as an action card, not a report card. Every element with no verifiable evidence scores 0 and names the gap. Every gap carries a recommended action. The analysis closes with a short prioritized list of what to do next — so “54/100” is never the deliverable. “54, because the economic buyer has never attended a session, so get them into the technical review this week” is.
That is the bridge from grading into coaching. The score exists so a manager can choose this week’s gap without watching 40 minutes of tape. The recommended action exists so the 1:1 has a concrete standard. The next meeting’s transcript is the accountability check: Ruby re-scores from evidence, trends each element against the last analysis, and flags declines even when the total went up.
Conversation intelligence tells you what happened on the call. Deal coaching changes what happens on the next one. If your stack only does the first job, you will keep mistaking a filled library for a coached team.
Run this on your next five deals
Open five open opportunities that had a call in the last two weeks. For each, answer three questions in under two minutes:
What is the single highest-severity gap the evidence supports right now?
What does “good” look like on the next call for that gap — named person, named ask?
Who owns checking whether that evidence showed up after the call?
If you cannot answer (1) without scrubbing a recording, your AI is grading and not routing. If you can answer (1) but (2) and (3) are blank, you have feedback without a loop. Fix the loop before you buy another scorecard.
Frequently Asked Questions
What is the difference between AI feedback and coaching?
Feedback tells a seller what happened or what was missing. Coaching turns that signal into a prioritized gap, a standard for the next conversation, practice, and a check that the behavior showed up with the buyer. AI can supply the signal at scale. Managers (and deal-level systems that attach next actions) close the loop.
Why don’t call summaries improve win rates on their own?
Summaries optimize for recall of the last meeting. Win rates move when the next meeting changes. A summary that never becomes a named gap, a practiced move, and a verified follow-up is documentation.
How many calls can a manager really coach well?
Far fewer than a modern CI library produces. Leaders regularly face hundreds of calls a month; deep coaching that lands is closer to a handful per week per manager. That is why automated scoring helps — and why scoring without prioritization and next actions fails.
How does Ruby connect scoring to coaching?
Ruby scores MEDDPICC from call and CRM evidence, attaches a recommended action to every gap, and trends each element over time so managers coach the move that would change the deal — not a generic call grade. Details live in how Ruby writes deal intelligence into HubSpot and in the MEDDPICC scorecard design.