Lead scoring software ranks your leads so your reps work the list from the top instead of from the middle. Every tool in the category does the ranking part well. What separates a scoring deployment that survives its first quarter from one that quietly gets ignored is not the engine, it is what you can feed it: a score is a function of fields, and a field you do not have does not score as zero, it scores as wrong.
So this guide evaluates lead scoring software the way you would evaluate a kitchen rather than a chef. We go criterion type by criterion type, firmographic, behavioural, intent and relational, and for each one we name the data it requires, what the score does when that data is missing, how fast that data goes stale, and what it costs to fill. If you are deciding which tool to buy, the shortlist you end up with will be shaped more by this than by any feature matrix.
What lead scoring software actually is
Lead scoring software assigns a number, a letter or a tier to each lead and account, so that routing, prioritisation and follow-up can be automated against it. It reads attributes about who the lead is and events about what the lead did, applies rules you wrote or a model trained on your closed-won history, and writes a score back to the record.
Lead scoring software comes in two families, and the distinction is real. Rule-based scoring adds and subtracts points you defined: plus fifteen for a director title, minus ten for a free email domain. Predictive scoring trains on your historical outcomes and infers the weights itself. Rule-based is transparent and reflects your bias. Predictive is more accurate and reflects your data coverage, which is a different bias and a less visible one.
Scoring, grading and routing
A score usually means engagement: what the lead did. A grade usually means fit: who the lead is. Many tools compute both and combine them, which is the right shape, because a perfect-fit account that has done nothing and a poor-fit account that downloaded everything need very different treatment and a single blended number hides both. Routing is the automation that consumes the result. Buying lead scoring software without deciding what the routing will do with the output is the most common way to end up with a number nobody acts on.

The four criterion types, and the data each one needs
Evaluate any lead scoring software against this table before you look at its interface. The third column is the actual system requirement. The fourth is what happens on the records where you do not meet it, which in most databases is a large minority of them.
| Criterion type | What it measures | The data it requires | What the score becomes when the field is empty |
|---|---|---|---|
| Firmographic | Whether the company fits your ICP | Headcount, industry, country, revenue band, legal entity | Silently neutral, so every unknown company is scored like an average one |
| Behavioural | What the lead did, and how recently | Tracked events with timestamps, identity resolution back to the record | Anonymous traffic scores nothing, so your most interested accounts can rank last |
| Intent | Whether the account is in market now | A dated signal the buyer emits: hiring, funding, technology change, research activity | The score becomes a fit score wearing an intent label, with no sense of timing |
| Relational | Seniority, role and buying authority | Normalised job title, department, seniority, tenure | Free-text titles bucket badly, so a Head of Growth and a growth intern can score alike |
The important column is the fourth one, and it is the one no vendor comparison shows you. Missing data does not raise an error in lead scoring software. It takes a default, and defaults are almost always neutral, which means every record you know nothing about is treated as mediocre rather than unknown. That is not a rounding problem. It systematically buries the accounts you have not enriched yet, which are frequently the new ones.
Firmographic criteria: the half that is easiest to fix
Firmographics are the fit backbone that lead scoring software leans on hardest: size, sector, geography, maturity. They are also the most commonly incomplete, because they usually arrive from a web form where the visitor typed a company name and nothing else, or from a list import that carried three columns.
The failure mode here is specific and worth naming. When industry is blank on a third of your records, a rule that awards points for target industries awards them only to the records that happen to be filled, so your score quietly becomes a measure of how completely a record was captured rather than how well the company fits. Enrichment coverage becomes the hidden variable, and it correlates with nothing you care about.
How stale it goes: slowly. Industry and country rarely change. Headcount moves enough over a year to matter for a banding rule, so an annual refresh is usually enough.
What it costs to fill: filling company attributes with Enrich Companies costs 1 credit per company, on the free plan included. Because companies are far fewer than contacts, this is the cheapest line item in the whole exercise and the one that moves the score the most.
Behavioural criteria: accurate, and narrower than it looks
Behavioural scoring is the part of lead scoring software that works best out of the box, because the tool generates its own input. Page views, email opens, demo requests, pricing page visits: the events are captured by the same system that scores them, so coverage is complete by construction.
The limit is not accuracy, it is population. Behavioural scoring can only rank people who already identified themselves. The account researching you anonymously for three weeks scores zero until someone fills in a form, which means a purely behavioural model ranks your known leads beautifully and is blind to your pipeline's actual top of funnel. Teams discover this when the model's top ten looks like a list of newsletter subscribers.
The second limit is decay. An email open from March is not evidence about today. Any behavioural criterion without a time decay rule accumulates points forever and eventually every long-lived record outranks every new one. If a tool you are evaluating cannot express decay, that is a genuine disqualifier, and it is worth more than three features on a comparison grid.
How stale it goes: fast. Behavioural relevance is measured in days and weeks. What it costs to fill: nothing extra, it comes with the tool. This is the one criterion type you do not buy data for.
Intent criteria: the only ones that answer "why now"
Fit tells you whether to talk to an account. Intent tells you whether to talk to it this week. Most lead scoring software sells intent as a premium add-on, and most teams buy it before they have fixed the firmographic half, which is the wrong order: an intent signal on an account you cannot qualify just makes you fast in a random direction.
Intent has one property that separates the useful signals from the decorative ones: a usable intent signal is dated. A company that posted three sales roles this month is a fact with a timestamp attached to it. A company that "shows interest in your category" is a score whose provenance and date you cannot inspect, and you cannot write a first sentence from it.
Two dated signals are cheap to resolve yourself. Hiring activity, through Company Hiring Signal at 1 credit per company, tells you an account is building the function you sell into. Technology changes tell you a stack moved. Both are verifiable, both carry their own date, and both survive the question a manager will eventually ask, which is "why is this account at the top".
How stale it goes: very fast, and this is the criterion most worth automating on a cadence. A hiring signal from four months ago is not an intent signal, it is history.
Relational criteria: the field that looks filled and is not
Seniority and department decide whether the person can buy. This is the criterion most likely to be present in your CRM and still useless, and it is where lead scoring software quietly loses precision, because job titles arrive as free text. "Head of Revenue Operations", "RevOps Lead", "Sr. Manager, Revenue Ops" and "revops" are four strings describing one role, and a points rule matching on a keyword list catches some and misses others without telling you which.
This is why title normalisation deserves an explicit question during an evaluation. Ask a vendor how their scoring handles title variants, and whether normalisation happens before scoring or not at all. It is a five-minute question that predicts more about your results than the demo does.
How stale it goes: moderately, and invisibly. People change roles inside the same company without any event reaching your CRM, so a title field ages without ever looking wrong.
What it costs to fill: refreshing a contact's role with Enrich Leads costs 1 credit per profile, free plan included. Note that this is per contact rather than per company, so it is the line that scales with your database size.

How to evaluate lead scoring software by its inputs
Feature grids compare the outputs of lead scoring software, which all look similar. These six questions compare inputs, which is where the tools actually differ. Take them to a demo.
Can it express time decay, per criterion? Not a global expiry, a per-criterion half-life. Behaviour needs weeks, firmographics need a year. A tool with one setting for both will be wrong on one of them permanently.
What does it do with a null, explicitly? Ask to see a record where a scored field is empty. If the answer is "it just does not add points", you now know unknown and unqualified are the same thing in this system, and you can plan around it.
Does it normalise before it scores? Titles, industries and country codes all arrive in several spellings. Normalisation is either in the product or it is in your ETL, and if nobody can say which, it is nowhere.
Can you write fields back in? The score is only as good as its inputs, so the ability to push enriched fields into the record on a schedule, by API or otherwise, is a hard requirement rather than an integration nicety.
Can it explain a single score? A per-record breakdown showing which criteria contributed is what lets a rep trust the ranking. Without it, adoption depends on faith, and faith runs out around week three.
What is the minimum history a predictive model needs? Predictive scoring on a few dozen closed-won deals is arithmetic dressed as machine learning. Ask for the number, and compare it honestly to what you have.

What lead scoring software costs to feed
The licence is the visible cost of lead scoring software. The invisible one is the data underneath it. Here is that arithmetic, using public prices and counting steps rather than features, because the most common budgeting error is to price a chain as though it were one action.
Start with the company layer, because it is where the fit criteria live and where the cost per unit is lowest. Enriching a company costs 1 credit. Adding a hiring signal costs 1 more credit per company. So a company made scorable on both fit and timing costs 2 credits, and a thousand accounts cost 2000 credits.
The contact layer scales differently. If you are importing profiles rather than starting from a list you already own, Import LinkedIn Leads costs 1 credit per profile, and enriching each one costs 1 more. That is two steps and therefore 2 credits per contact, not 1. On the free plan, 100 credits a month at zero euros makes roughly fifty contacts scorable, or fifty companies scorable on both fit and timing. Paid plans start at 9 euros a month on MINI and 20 euros on STANDARD, with the cost per credit falling as volume rises.
Two practical notes. First, prioritise companies over contacts: a company enrichment improves the score of every contact attached to it, so the same credit does more work. Second, enrich the segment you are actually going to work this quarter rather than the whole database. Scoring quality is a coverage problem within the population you touch, not a completeness trophy across a base where most records will never be called.
If you need to build the target population itself rather than enrich an existing one, Import Companies from a Prompt at 1 credit per company and Find Similar Companies at 1 credit per company build a list from a description or from your best existing customers, which is a more honest starting point for a fit model than a database export nobody chose.

The contract between your enrichment and your score
Lead scoring software recalculates continuously, so the layer that fills its fields has to satisfy three properties rather than simply exist. Whatever tool you use for the filling, check it against these, because a scoring model degrades in silence when any one of them is missing.
It writes back on a schedule tied to decay, not to the calendar. Weekly for intent, quarterly for contact roles, annually for firmographics. A monthly refresh applied uniformly is simultaneously too slow for the first and wasteful on the third. Setting that cadence field by field is what our guide on CRM data strategy works through, and CRM contact management has the one-hour audit that tells you where you stand before you set it.
It is idempotent and dated. Re-running it on an unchanged record must not churn the field, and every field it writes should carry the date it was written, otherwise you can never tell a fresh value from a two-year-old one and neither can the model.
It can run without a human remembering. This is the property that decides whether the score is still trustworthy in six months. A REST API call from your CRM automation when a record is created or a refresh falls due, available from the Standard plan, is what makes the difference between a scored database and a database that was scored once. The Google Sheets sidebar covers the first pass and the audit, where you discover which fields were actually missing, and those are usually not the ones you assumed.
Our sister guide on HubSpot lead scoring goes through this same plumbing inside one named CRM, property by property, if that is the implementation you are heading for.

Scoring accounts that are not in your CRM yet
Every guide on this subject assumes the leads are already in the database. That assumption quietly excludes the most valuable population you have, which is the set of companies that fit perfectly and have never heard of you. Lead scoring software cannot rank a record that does not exist, so outbound teams end up scoring inbound and calling it pipeline prioritisation.
The fix is to build the population before you score it, which inverts the usual order. Start from the accounts you already won, find their structural neighbours, enrich the resulting list on the same fit fields your model weights, and you have a scored target list rather than a scored inbox. Find Similar Companies does the neighbour step at 1 credit per company, working from your best customers rather than from a filter you guessed.
The practical benefit is that fit criteria behave much better on a list you constructed than on one that accumulated. Coverage is uniform because you enriched every row in the same pass, so the score measures fit rather than measuring how completely each record happened to arrive. That is the cleanest scoring population most teams can build, and it is usually the one nobody builds.
How to prove your model actually works
Configuration guides are everywhere; verification guides are not. Here is how to find out whether your lead scoring software is ranking or just sorting, in one measurement you can run next quarter.
Freeze the score at hand-off, not at close. Store the score as it stood when the lead was routed. Reading the current score against a closed deal proves nothing, because the score kept moving while the deal progressed and will look prescient by construction.
Compare deciles, not averages. Bucket the frozen scores into ten bands and compute the conversion rate of each. A working model shows a monotonic slope from the bottom band to the top. An average conversion rate for "high scores" against "low scores" can look healthy while the middle is pure noise.
Keep a holdout. Route a small random sample by arrival order rather than by score, and compare. Without an unscored control you cannot separate the model's contribution from the effect of reps simply working a shorter list more carefully, and that effect is real.
Wait a full sales cycle before judging. Reading outcomes at four weeks on a cycle that runs three months measures speed of response, not quality of ranking. Set the window from your own median cycle length and resist reading it early.
If the slope is flat after an honest measurement, the answer is almost never to re-tune the weights. It is to go back to the fill rate on the fields those weights read, because a model cannot extract a signal from a column that is empty on half the rows.

Five ways a lead scoring software project fails
Scoring before measuring fill rate. Lead scoring software will happily run on a database it has never audited, and our enrichment workflow guide covers how to fill the properties on a schedule instead. If you do not know what percentage of records carry the fields your model weights, you cannot distinguish a bad model from a thin database, and you will spend the next quarter tuning weights.
Buying on the vendor's dataset. Every demo runs on a database with complete fields, which is precisely the condition your own database does not meet. Ask to pilot on an export of your own records, nulls included, and watch what the tool does with the gaps, because that is the behaviour you are actually purchasing. Our CRM data quality report covers how to measure that starting state.
Scoring the whole base instead of the worked segment. Enriching and scoring two hundred thousand dormant records to rank the four hundred your team will actually call this quarter spends the budget where nobody looks. Scoring quality is a coverage problem inside the population you touch.
Thresholds set once. An MQL threshold picked in January against a pipeline that has since doubled sends a volume of leads nobody planned for. Thresholds are a quarterly decision.
No feedback loop from closed-won. If nothing compares the score at the time of hand-off to the eventual outcome, nobody ever learns whether the model works, and the honest answer after a year is that you still do not know.
Is your data ready to be scored?
Select every criterion type you can populate on most of your records today. The verdict tells you which half of the score is real and which is a default, not a mark out of ten.
Select the criteria you can populate on most records today.
Multi-select: click a criterion again to remove it. The verdict splits FIT (who they are) from TIMING (why now), because a score can be solid on one and entirely defaulted on the other.
Choosing lead scoring software: the short version
Pick your lead scoring software on three properties rather than on a feature count: it must express decay per criterion, it must let you see what a null did to a score, and it must accept fields written back on a schedule. Almost every serious tool in the category can rank; those three are where lead scoring software actually separates.
Then spend the first two weeks on the inputs rather than on the weights. Measure fill rate on the fields you intend to score, enrich the company layer first because it is the cheapest credit you will spend and it lifts every contact underneath it, add one dated signal so the ranking can answer why now, and normalise your titles before a single rule matches on them. A modest rule-based model on well-covered fields outperforms a sophisticated model on a sparse database, reliably, and it is the version your reps will still be using next quarter.
Continue exploring this cluster
Start enriching your sheet in 30 seconds
Free for 100 credits/month. No credit card.