Imagine a medical test that tells a patient they have a serious illness, but is only right 5% of the time. Nineteen times out of 20, the result would be wrong: a false alarm that could leave a patient frightened and lead to unnecessary tests and treatment. No regulator would allow a test with that level of accuracy to be used in a hospital.
Yet for roughly two decades, B2B marketing has tolerated something remarkably similar in its lead scoring. Long-running Gartner benchmark data puts the rate at which sales accepts MQLs at around 5-6%. In other words, the typical scoring system is telling sales that a lead is ready when, in the vast majority of cases, it isn't.
That's the starting point for Episode 3 of B2B Effectiveness: Evidence-Based Marketing Ideas for B2B Practitioners, in which Dale W. Harrison and Liam Moroney examine why lead scoring continues to perform so poorly.
Their conclusion is that the problem runs deeper than the particular signals being used or the way those signals are weighted; the points-based model itself is fundamentally incapable of producing the kind of qualification accuracy marketers expect from it.
Why every proposed fix keeps failing the same way
Over the years, the industry has introduced one alternative after another: marketing qualified accounts, marketing qualified opportunities, AI-qualified leads and various other attempts to improve on the MQL. Yet most of these approaches leave the underlying scoring mechanism intact.
The mechanism is familiar. A collection of behavioral or demographic signals is assigned a numerical value and those values are added together. Visiting an About page might be worth one point; downloading an eBook might be worth ten. But where did those numbers come from? In most cases, they weren't derived from any meaningful relationship between the behavior and the eventual commercial outcome. They were chosen, tested and adjusted until the resulting scores seemed useful.
AI-driven scoring is often presented as a more sophisticated answer because it can process far more signals than a human-defined model. That gives it greater scale, but it doesn't necessarily address the fundamental problem. If the underlying logic still treats observed activity as evidence that can simply be accumulated into a score, adding more data and more sophisticated mathematics can make the model more elaborate without making its predictions more reliable.
Data is not information
Dale illustrates the distinction with a simple example. Ask someone a question and they give you an answer. Ask exactly the same question ten times and they give you the same answer each time. You now have ten data points, but you haven't acquired ten times as much information. You have received the same piece of information repeatedly.
That distinction matters enormously in lead scoring. If someone downloads ten articles from your site, the activity may tell you something useful about their interests and their fit with your target audience. But downloading 100 articles doesn't make you ten times more certain that they are a good prospect. At some point, additional activity stops telling you anything new.
A points-based model has difficulty recognizing that distinction. If every download continues to add points, the model keeps treating repeated behavior as fresh evidence even after the information contained in that behavior has largely been exhausted.
The real flaw is failing to distinguish wins from losses
This becomes even more important when models are built from historical sales data. A common approach is to take all of the accounts that became customers, identify the characteristics they share, and use those characteristics to predict future buyers.
The problem is that a characteristic shared by your winners only matters if it distinguishes them from your losers.
Suppose 31% of your closed-won accounts visited your LinkedIn company page. That might initially look like a useful signal. But if 31% of your closed-lost accounts visited the page too, the behavior tells you nothing about the likelihood of winning. It is common to both outcomes, so it has no discriminatory power.
The useful patterns are the ones that separate wins from losses. They need to occur at materially different rates between the two groups and, ideally, do so consistently enough to improve prediction beyond what you would achieve by chance. That sounds obvious, but many lead scoring systems are built around correlations that have never been properly tested against the outcome they are supposed to predict.
Measuring what you're actually missing
There is another problem with the way lead scoring is usually evaluated. Most teams focus on false positives: leads that the model scores as ready but that never become customers. The other half of the equation is just as important. What happens to the leads the model scores poorly that later become genuine opportunities?
Answering that question requires a closed-loop analysis. Dale describes tracking leads in real time and then returning to those records months later to see which ones actually entered an active, sales-qualified process, regardless of what their original score predicted.
That kind of analysis reveals some uncomfortable patterns. One is what Dale calls “click monkeys”: people who click on everything, download everything and appear highly engaged, yet never become customers. Their digital behavior can look remarkably similar to that of a serious buyer.
The reverse also happens. Some of the strongest buyers leave almost no meaningful digital footprint before they enter a sales process. They may have been referred by a trusted colleague, worked with the product at a previous company or inherited an existing supplier relationship when they changed jobs. By the time they appear in the database, they are already well on their way to a purchase.
This is part of what is often described as dark social. The buyer is engaging and gathering information, but much of that activity takes place outside the channels a marketer can observe and score.
The buyer is scoring you too
There is an important symmetry here that lead scoring tends to overlook. Companies spend considerable effort scoring buyers, but buyers are evaluating companies at the same time.
A product page filled with generic marketing language might carry very little weight with a buyer. A recommendation from a trusted colleague, or previous experience using the product at another company, could carry considerably more. The difficulty for marketers is that the signals they can observe digitally are often not the signals carrying the greatest influence.
This is also where product and company fit become important. Two businesses can sell similar products and achieve very different close rates because their experience, reputation and commercial capabilities are concentrated in different markets. One might have developed a strong understanding of healthcare and government procurement, for example, while another has built its expertise around industrial and manufacturing customers.
Those differences affect the likelihood of winning, yet they are difficult to capture through a conventional engagement score. The question is ultimately about the relationship between a particular buyer and a particular seller, rather than simply how active that buyer has been.
A coin flip is the baseline you have to beat
Strip away the framework names – MEDDIC, BANT and whatever comes next – and the underlying test for any qualification system is relatively simple. If you know nothing about a lead, your starting point is effectively random: heads, it becomes a win; tails, it doesn't.
A useful scoring system has to improve materially on that baseline. It needs to identify patterns that distinguish likely winners from likely losers and demonstrate that those patterns continue to work when tested against real commercial outcomes.
The fact that MQL acceptance rates have remained around 5-6% for so long suggests that much of the industry has not achieved that. We have built increasingly elaborate systems around a fundamentally weak premise, added more data to them, given them more sophisticated names and, increasingly, introduced AI into the process. The technology has changed considerably. The underlying logic has changed much less.
The next episode picks up a thread that initially sounds unrelated: how rising and falling markets disrupt the familiar “95-5 rule” – and why the effect is far more significant than most marketers realize.
.png)