Sales teams

How to build a B2B lead list that is worth working, instead of buying one

Most lead lists fail before the first email goes out. Wrong companies, stale contacts, addresses guessed from a pattern. A built system works the other way round: it turns your criteria into checks, runs them across a large set of companies, fills the gaps from public sources with the source attached to the row, and holds back the rows it could not confirm.

Sales teams: How to build a B2B lead list that is worth working, instead of buying one

Why bought lead lists and manual prospecting both run out of road

There are two normal ways to get a list, and both have the same ending. You buy one, and it arrives large, cheap and wrong: companies that no longer match, people who left, addresses assembled from a first name and a domain. Or you build it by hand, and it is accurate for as long as one person keeps working on it, which is never long enough to fill a pipeline.

The damage from the bought list is not the price. It is that no row can be argued with. Nothing on a line says which criterion it was meant to satisfy, where the company size was read, or when. So the first row that turns out wrong puts the whole file in doubt, and there is no way to correct one line instead of distrusting all of them. Sales stops trusting the file, works the twenty names they recognise, and the rest becomes an expensive spreadsheet.

The manual version has the opposite problem. The research is good and there is never enough of it. A rep who researches properly gets through a handful of companies an hour, and the moment they stop, the list starts ageing. Six weeks later the person has moved on and the trigger has passed.

Underneath both is a definition problem. Ask five people in the same company who the ideal customer is and you get five answers, all true, none checkable. Nobody can build a list from a description like mid sized, ambitious, professional, because nothing in that sentence can be looked up.

  • The file arrives as one block: you either trust all of it or none of it, one wrong row at a time.
  • Half the named contacts are no longer in the role, and nothing in the file says so.
  • Rows are filled with info@ addresses, so nobody in particular ever reads them.
  • Your ideal customer profile lives in a slide, in words nobody can check from outside.
  • Reps research one company at a time and stop when the day gets busy.
  • Nobody knows where a given row came from, so nobody can defend it when asked.

How a lead engine builds the list and keeps the evidence for every row

The system does the same work a careful researcher would do, at a volume a person cannot reach, and it writes down its evidence as it goes. You give it criteria in plain language. It turns them into checks, runs them across a large set of companies, and delivers what survived.

  1. 01

    Turn the request into criteria that can be checked

    You describe the target in a sentence: manufacturers in the DACH region, fifty to two hundred fifty employees, currently hiring production planners. That becomes separate criteria, each marked as hard, as a filter or as a timing signal, and each with a defined way of being confirmed from outside. Criteria that cannot be checked from outside are named before the run starts.

  2. 02

    Screen a large set of companies, not a shortlist

    The check runs across a broad set of company records rather than the first page of results. Most of what is checked never appears in your list, which is the point: the value is in what got excluded and why, not in the size of the file.

  3. 03

    Fill the gaps from public sources, with the source attached

    Company size, industry and site count are frequently missing or wrong in any base data. The system fills them from public sources and keeps the reference on the row, so every value can be traced back to where it was read.

  4. 04

    Find the person, not the mailbox

    The target is a named person in the relevant role, with their function on the row. Where no named decision maker can be found, the row is held back rather than filled with a generic company address. A held row is visible to you, it is simply not sold to you as a contact.

  5. 05

    Make every row disputable on its own

    Each delivered row carries what it was checked against, which check it passed, which value came from which source and on what date that source was read. That turns an argument about the list into an argument about a single line. You can strike one row, correct one threshold, or reject one source, without throwing away the file. A list that cannot be disputed line by line is a list nobody defends.

  6. 06

    Attach the timing signal with a date

    If the request includes a trigger, for example an open role or a new site, the system records where the signal was found and when it was published. A trigger without a date is not a trigger, it is a rumour, so undated signals are marked as such.

  7. 07

    Throw rows out and say why

    Wrong size, group subsidiary, no reachable contact: those rows are removed and the reason is kept. You get the discarded set as well as the delivered set, because the discards tell you whether the criteria were right.

  8. 08

    Stop where the evidence stops

    Anything unambiguous runs on its own. Anything ambiguous waits for a person: a company that may or may not be a subsidiary, a contact whose role changed recently, a signal that cannot be dated. The system flags. It does not fill the cell with its best guess, because a confident wrong row is more expensive than an empty one.

  9. 09

    Keep it current and keep the record

    The list is not a state you reach once. It is rebuilt on a schedule against the same criteria, so a company that stopped matching drops out and a company that started matching appears, without anyone maintaining a spreadsheet. What survives a rebuild is the record: origin and date of each row are stored, contacts who ask not to be approached are recorded and stay recorded across every future run, and deletion requests are honoured in the source rather than patched in one export. The criteria themselves keep moving with the company, because the grid is edited and the next run simply applies it.

Take this with you

A prompt that turns your ideal customer profile into a checkable qualification grid

This is the first step of the build, in a form you can run today. You describe your ideal customer in your own imprecise words. The prompt turns that into a grid: for each criterion, what you actually mean, how it is recognised from outside the company, and where that information comes from. It adds the hard disqualifiers and the signals of a current reason to talk. At the end it names the criteria that cannot be checked from outside at all, which is usually the most useful part.

prompt
You are a B2B research analyst. Turn my description of the customers I want
into a qualification grid another person could apply to a company without
talking to it. Work only from what I write. Invent no industries or sources.

WHO I WANT
[TWO TO SIX SENTENCES IN YOUR OWN WORDS. Keep the vague parts: "mid sized",
"ambitious", "professional enough to care about quality". Do not clean up.]

WHAT I SELL AND WHO SIGNS
[ONE OR TWO SENTENCES: the offer, typical deal size, the deciding role.]

THREE BEST CUSTOMERS
[EACH: industry, rough size, country, one line on why they bought.]

THREE WORST FITS
[EACH: one line on why it went badly or never closed.]

1. CRITERIA GRID
One row per criterion, vague ones made concrete. Columns:
  CRITERION    the plain name
  WHAT I MEAN  a testable definition with a threshold or range where one is
               possible. If I gave no threshold, propose one, mark PROPOSED
  TYPE         hard (must be true), filter (must not be true), signal (timing)
  VISIBLE AS   what is observable from outside: a stated employee count, a
               published job posting, a second address, a certification
  SOURCE       the kind of public source that carries it, as a category
  CONFIDENCE   high, medium or low, with a half line on why

2. HARD DISQUALIFIERS
Conditions that remove a company even if everything else fits, taken from my
worst fits. Each with how it is spotted from outside.

3. TIMING SIGNALS
Observable events suggesting a reason to talk now. For each: what it looks
like, where it appears, how long it stays meaningful.

4. NOT CHECKABLE FROM OUTSIDE
Criteria that cannot be confirmed without a conversation. Say so plainly, and
for each give the nearest observable proxy and what that proxy gets wrong.

5. WHAT I DID NOT TELL YOU
Up to five questions whose answers would change the grid, and which rows
above currently rest on a guess.

Output headed lists or a table. No sales advice, no summary.
  1. Write the who I want section badly on purpose. Vague words are the input, the grid is what makes them concrete.
  2. Use real customers for the best and worst fit sections. Invented examples produce an invented grid.
  3. Read section 4 first. Criteria that cannot be checked from outside are the ones that will silently break any list you build.
  4. Argue with the PROPOSED thresholds and rerun. Two rounds usually settle the definitions your team never agreed on.
  5. Take the finished grid into your own research, or hand it to whoever builds your lists, so at least everyone qualifies the same way.

The grid is a sheet of paper. A genuinely useful one, and most sales teams have never written it down, but it researches nothing. It does not look up a single company. It does not know who works there now, it cannot show you a source for anything it claims, and it has no idea which of your targets posted a job last week. It goes stale the day your market shifts, and nothing tells you that it has.

From a definition of the right company to a list you can defend line by line

The grid says what to look for. The built system does the looking, against a large set of companies, and comes back with rows you can actually work. That is the entire distance between the two, and it is where all the effort lives.

In the built version each criterion becomes a check that either passed or did not, on every company, in the same way. Missing attributes are filled from public sources with the reference kept on the row, so every value on the line can be traced back to where it was read and when. Contacts are named people with a function. Rows that fail are removed with the reason kept, so you can see whether the criteria were too narrow or too generous. The practical effect is that you argue with one row instead of the whole file.

It stops where the evidence stops. A company with no named decision maker is held rather than filled with a general mailbox. A signal that cannot be dated is marked as undated. Borderline cases go to a person. This is why the delivered list is smaller than what you would buy, and why it is worth working.

The legal side is stated plainly rather than skipped. Business contact data is personal data. Collecting, storing and using it for outreach is regulated, in Europe by the GDPR and by the rules on unsolicited commercial contact in each country. The system is built for that: it stores where each row came from and when, it records objections and carries them into every future run, it honours deletion at the source rather than in one export, and it works from sources whose terms permit the use. Nobody here will tell you that having an address means you are allowed to use it.

  • Your criteria as explicit checks, applied the same way to every company.
  • Named decision makers with their function, not generic company mailboxes.
  • Each row carries the result of every check, so one wrong line is corrected instead of the file being written off.
  • Missing company attributes filled from public sources with the source kept on the row.
  • Rejected rows delivered with their reason, so the criteria can be corrected.
  • Timing signals recorded with a date and an origin, or marked as undated.
  • Origin, consent status and objections stored per contact, and carried across every rebuild.
  • Rebuilt on a schedule instead of ageing in a spreadsheet.

Frequently asked questions

Can I not just build a lead list with ChatGPT?

For the definition work, yes, and the prompt on this page is exactly that. What a chat window cannot do is check anything. It has no reliable view of which companies match, it cannot show you where any claim came from, and asked for company names it will produce plausible ones, which is worse than none. The built system is mostly checking and recording evidence, and that is the part a chat window on its own cannot perform.

Is building a B2B lead list GDPR compliant?

Business contact data is personal data, so the answer depends on how the list is built and what you do with it, not on where it came from. A system can make the position defensible: record the source and date of every row, keep the lawful basis and any objection with the contact, honour deletion at the source, and use only sources whose terms allow it. What no system can do is give you permission to email whoever you like. The rules on unsolicited commercial contact differ by country and remain your decision to follow, with your own legal advice.

How is this different from buying a lead list from a data provider?

A bought list is a snapshot of someone else's database, sold to many buyers, correct on average rather than row by row. A built list is generated against your criteria, checked at the time of delivery, and handed over with its evidence and its rejects. It is usually much smaller than the file you would buy, and that is the point.

How do you keep a B2B lead list from going out of date?

The list is rebuilt on a schedule against the same criteria rather than maintained by hand. A company that no longer matches drops out, a company that started matching appears, and a timing signal that has expired stops counting as one. What carries across every rebuild is the record: the origin and date of each row, and any contact who asked not to be approached. Changing the list means changing the criteria, not editing cells.

What happens to companies where no decision maker can be found?

They are held, not filled with a generic company address. You still see them, marked with what was found and what was missing, so you can decide whether they are worth a manual look. Filling those rows with info@ addresses is how lists get large and useless at the same time.

How much does it cost to build a lead engine, and how long does it take?

The definition work takes days, because it starts from customers you already have. A first list against real criteria comes in weeks rather than months, and you work it while the rest is built out. We do not publish a price because it depends on how many sources have to be handled and how strict the checks need to be. What you get before the build is an honest view of which criteria are checkable, since criteria that cannot be checked are what make a build expensive.

Where does the data live, and can we keep the list if we stop working with you?

On your infrastructure: your server or your cloud account, your database, your exports. The list, the evidence behind each row and the objection records are yours and stay behind if the engagement ends. Nothing about your target market sits in a vendor system you cannot inspect.

Will the system contact these people by itself?

Not as part of this. Building the list and running outreach are separate decisions, and outreach is where the legal and reputational risk sits. The list can feed an outreach system, but sending is a step you switch on deliberately, with a person responsible for it.

Bring us your three best customers

If the prompt above showed you that half your ideal customer profile cannot be checked from outside, that gap is the project. Apply, bring three customers you would like more of and three you regret, and we will tell you honestly whether a list can be built against your criteria or whether the criteria need work first.

Apply

Related use cases