Four Outcomes, not Two.
Week one on Harbourgate Logistics. A rep picks up the account and runs the research routine everyone runs: the website, the pricing page, the last two funding announcements, the LinkedIn profiles of everyone on the invite, a competitor's G2 reviews, a 10-K because the account is public, and a Slack message to whoever sold into freight last.
Two hours in, there are twelve tabs open and no thesis.
The research produced facts, and several of the facts contradict each other. Harbourgate is expanding into two new regions while running a cost-reduction program. They hired a VP of Data in March, then posted a job in June for the role that hire was meant to replace. Every one of those is a real signal. None of them tells the rep what to lead with at 11am tomorrow.
So the rep opens the standard deck and runs discovery as a fishing expedition, hoping the buyer names a problem the product happens to solve. Sometimes the buyer does. More often the buyer describes a priority the rep cannot connect to anything specific, and the meeting ends with a polite second call that never gets booked.
Nothing in that morning was a research shortage. The account simply had no hypothesis attached to it — nothing the 11am call could confirm or kill.
And the teams who do write one down mostly count two ways it can end: the hypothesis survived, or it died. There are four. The two that get missed are where deals quietly stall.
What is a use case hypothesis?
A use case hypothesis is a written, falsifiable claim about which specific problem your product solves for a specific prospect, made before you have spoken to them. It names the pain you believe exists in that account, the capability that addresses it, and the proof point that makes the claim credible to that buyer. It has to be specific enough that one discovery call can confirm it or kill it.
A pitch asserts value and then gets defended. A hypothesis proposes value and invites correction, which is why a good one makes discovery sharper rather than redundant. You are not walking in to present a conclusion. You are walking in with the best available theory and the questions that would prove it wrong.
What makes a hypothesis qualified
Three parts, and it needs all three. Miss one and you have a talking point.
A named problem in the buyer's world. Not a capability gap inferred from their tech stack. A specific operational or commercial problem, attributable to something observable about that account: how they go to market, what they have publicly committed to, who they have just hired. "They probably need better reporting" is not a problem. "Their support team tripled in twelve months and their published SLA has not changed" is.
A capability that addresses it. Which part of your product, specifically. A category name will not survive the buyer asking how in the first four minutes.
Proof that lands for this buyer. The customer story or ROI model closest to their situation. A logo from their industry at their scale beats a bigger logo from a different one, and both beat a generic percentage.
One more test sits on top of those three. The hypothesis has to be specific enough to be wrong. If a discovery call cannot disprove it, what you are holding is positioning.
Account research alone never produces a thesis
Any AI can research an account. Deciding which use case matters is a different problem, because that decision requires knowing your side of the equation as well as theirs.
Ask a general-purpose AI to research Harbourgate and it will tell you accurate things about Harbourgate. It cannot tell you which of your seven use cases fits them, which of your customer stories they will find credible, or where you tend to lose to the incumbent they are probably already running. It has none of that, so the output is generic by construction.
Ruby holds that second half in two forms.
The first is a retrievable knowledge base. Your material is typed rather than dumped into one pile: customer stories, sales plays, feature and benefit detail, battle cards, competitor material, ROI models, reference decks, ICP definitions. Typing it is what lets an agent retrieve the three documents that bear on this account instead of summarizing everything you have.
The second is structured organizational context. Your ICP, competitive positioning, sales process, pricing, and win/loss patterns, held as context rather than as documents.
Agents draw on both, selectively. The Account Research agent reads your ICP and competitive context. The Hypothesis Generation agent reads your ICP, sales process, competitive positioning and win/loss patterns, and retrieves against your sales plays, customer stories and feature detail. The Deal Health agent reads your sales process and meeting-quality standards. Each agent sees only the organizational context its job requires, which is why the output reads as your company's argument rather than as a summary of your whole drive.
What the hypothesis stage produces
Adding an account does not start any of this. A new account sits in a pre-hypothesis state where no stage agents run. The work begins when the deal enters the hypothesis stage — the second of Ruby's five journey stages, and the first that runs any agents at all. That is still well before the first call. Preparation is a stage in the process rather than a scramble the night before.
The chain runs in order.
Account research is a hard prerequisite for everything after it. What the company does, how it makes money, its size and structure, market position, recent developments and hiring signals, and the technology context that shapes what buying your product would involve. Plus per-attendee background for the people on the invite.
Hypothesis generation matches that account picture against your organizational memory and produces a packet of eight sections: Our Story, Top Use Cases, Why Now, Challenger Insights, Similar Customers, Objection Handling, Potential Value, Ideal Meeting.
Two parts of that packet do the real work. Top Use Cases carries two to three candidate use cases, each with its own fit reasoning and, where an analogous customer exists, the concrete outcome that customer saw. Writing down alternatives is deliberate, because the use case you would have picked is frequently not the one that survives discovery. Alongside each hypothesis sits the specific question that would test it.
The first meeting kit is built from both. A six-section brief, and a first meeting deck built around the hypothesis rather than around your standard corporate flow. The deck is generated against a fixed slide structure at temperature zero, then re-validated for structure and fabrication after the model has written it.
The detail that matters most is easy to miss. The discovery questions in the brief are generated from the hypothesis's own validation tests, trimmed but not reworded. What the rep walks in to ask is what would prove the rep wrong.
Nobody is handed twelve tabs to synthesize. They get a thesis with its reasoning shown, which they can accept, adjust, or reject on the basis of something they know that Ruby does not.
The brief labels it a hypothesis, on purpose
There is a failure mode worse than having no pre-call thesis, and that is a speculative thesis mistaken for a validated one. A rep who believes an unvalidated assumption is a confirmed requirement will stop asking about it, build a business case on it, and forecast against it.
So Ruby labels pre-validation thinking as exactly that, and the labels are literal. On a first meeting against an existing account, section 5 of the six-section brief reads Deal Recap (Hypothesis — To Be Validated), and its bullets are labeled in kind: Why Do Anything (hypothesis), Why Us (hypothesis), Why Now (hypothesis). Before a deal exists at all, that section reads Account Angle (Working Hypothesis — Pre-Deal). Where there is no account on file it reads Working Hypothesis (To Be Validated), sitting above a Company Snapshot instead of an Account Recap. A first meeting brief carries no Last Meeting section at all, because there is no earlier call to stand on.
The discipline runs down to the cell level. In the attendee table, the role column names a pre-meeting role hypothesis and renders one hedged label, never a confirmed role. Only Ruby's champion identification step can call anyone a champion. Everywhere else, a promising contact stays a potential champion until the advocacy shows up in the record. When Ruby has nothing on an attendee, the brief says so in one line and tells the rep to confirm the role at the start of the meeting rather than speculating about seniority or department.
The rule is the same one that governs everything Ruby writes. No evidence, no entry. A hypothesis is allowed to be speculative. It is not allowed to arrive looking like a finding.
The discovery call is the test
Ruby ingests meeting transcripts from the tools your team already runs, including Google Meet, Microsoft Teams, Zoom, Fathom, Granola, Gong and Fireflies. Nobody uploads a recording and nobody has to remember to.
Four outcomes, not two
Most teams run a two-state model without ever saying so: the hypothesis survived, or it died. That binary is why assumptions go unchecked for a quarter. It has nowhere to put a hypothesis the call half-supported, and nowhere at all to put one the call never touched.
Ruby re-reads every pending hypothesis against what the buyer actually said and resolves each into one of four states. Because the packet carries two or three hypotheses, a single call usually produces more than one.
Harbourgate's packet carried three: margin leakage from manual exception handling, inconsistent reporting definitions across the regional teams, and rate leakage at carrier contract renewal. One discovery call resolved them differently.
Validated. Clear evidence supports it: a direct mention, quantified pain, or stakeholder confirmation. The operations director raised the reporting problem unprompted — three regional teams reporting on incompatible definitions, so consolidated reporting takes a week — and put a number on it. That hypothesis is now sourced to their words, so it carries into the business case with their framing and their figure instead of an industry benchmark.
Partially validated. A related signal, a tangential mention, the right pain in a different framing. Had the director described the reporting mess without ever quantifying it, this is where it would have landed: confidence up, hypothesis still open, the number still to get.
Invalidated. Evidence contradicts it. Exception handling — the hypothesis the rep would have opened with — does not survive contact. Annoying, the director says, but nobody measures it. Ruby marks it invalidated, cites the moment in the call that did it, and drops the confidence score.
No evidence. Carrier rate leakage never came up. Most teams do not track this state, and it is the one that costs them. The hypothesis stays pending at unchanged confidence rather than quietly decaying into an assumption nobody rechecks.
One call, three hypotheses, three different answers. The rep leaves with a use case sourced to the buyer's own words rather than the one they walked in planning to lead with, which is the whole argument for writing down two or three before the meeting instead of one.
Confidence moves in proportion to the evidence. A direct quote carrying a number moves it hardest, an explicit denial moves it hardest the other way, and a call that did not address the hypothesis moves it not at all.
Ruby also tracks what fits none of the four. Surprises are findings that contradict every existing hypothesis or introduce something nobody predicted: an unforecast pain point, a stakeholder priority that maps to nothing you sell, a competitive or budget shift. A buyer priority that matches none of your use cases is the most valuable outcome a first call can produce, because it is a qualification signal arriving in week one instead of in the quarter you forecast the deal into.
Every subsequent meeting is another test. The use case that wins a deal is often not the one that opened it.
Where the hypothesis goes to work
Implicated Pain. Pain is the heaviest element in Ruby's MEDDPICC scoring, carrying 20 of 100 points against 15 each for Metrics, Economic Buyer and Champion, and 5 for Paper Process. It scores on evidence, and the bottom of the scale says so. Zero means no pain identified rather than pain the rep feels confident about. An unevidenced use case scores zero and the gap gets named.
Compelling events. Pain without a date is a roadmap item. A confirmed use case is what a compelling event attaches to, and the two together separate a deal that could close from a deal that has a reason to.
Generated content. Business cases, meeting decks, POC plans, proposals and follow-up emails are written against the deal's current state, so material produced after a hypothesis is revised argues the surviving use case rather than the discarded one. Documents already generated are versioned artifacts. They do not rewrite themselves, and Ruby surfaces the existing one rather than quietly producing a second.
Buying committee framing. The same use case has to be argued differently to a technical validator and to an economic buyer. Ruby frames it per stakeholder, using the roles already mapped on the deal.
Key takeaways
A hypothesis is not a pitch. It is a falsifiable claim you bring to discovery so the call can test it, which makes discovery sharper rather than redundant.
A qualified hypothesis has three parts: a named problem in the buyer's world, the specific capability that addresses it, and the proof point closest to their situation. Then one test on top, which is that it has to be specific enough to be wrong.
Organizational memory is the prerequisite. Researching an account is not the same as choosing which of your use cases fits it. The second needs your customer stories, sales plays, competitive material and ICP, typed so they can be retrieved rather than summarized.
Preparation is a stage, not a scramble. The account picture, the use case hypotheses and the first meeting kit are produced ahead of the call by the hypothesis stage.
The questions come from the hypothesis. The brief's discovery section is generated from the validation tests attached to each hypothesis, so what you walk in to ask is what would prove you wrong.
Pre-validation thinking gets labeled as such, down to the column header. A working hypothesis presented as a confirmed requirement does more damage than no hypothesis at all.
The state nobody tracks is "never tested." Validated and invalidated are easy. A hypothesis the call simply never touched is how an assumption survives three meetings without anyone checking it.
Two or three hypotheses, not one. The first use case you would have picked is often not the one that survives the call, so the alternatives get written down before the meeting instead of improvised after it.
Try this before your next first meeting
Twenty minutes, one account. Take the next new logo on your list and write one sentence before you do any research:
We think [company] is struggling with [specific problem], because [observable evidence], and [specific capability] addresses it. The closest proof we have is [customer].
If you cannot finish that sentence, you are not ready for the call, and no amount of additional tab-opening will change it. If you can finish it, take it into discovery and try to break it. Either the buyer confirms it and your business case is already half written, or they hand you the real use case in the first ten minutes.
Then look back at your last five closed-lost first meetings and ask what hypothesis you walked in with. If the honest answer is "the standard deck," the problem was never discovery. It was preparation, which is the one part of the sales process you can fix before the buyer is involved at all.