Home Guides How to Build a Target List for a Search Fund Thesis
Guide · Search funds & ETA

How to Build a Target List for a Search Fund Thesis

A 14-minute, method-first guide to constructing the one asset a search actually runs on. Every step works with or without us: census the real universe, translate your thesis into website-visible evidence, and rank on receipts instead of exports.

702
eligible machining cos. in one US census
~17,300
operating companies from a 25k-domain triage
1 in 10
keyword-perfect candidates already group-owned

The list is the search

A traditional search comes down to one asset: the ranked list of companies you will contact. Everything downstream — response rates, meeting quality, the odds the one company that says yes is worth buying — traces back to how that list was built.

The overlap problem

Funded searchers have grown severalfold in a decade, nearly all drawing from the same handful of databases. When 300 searchers filter the same index with similar criteria, exports overlap heavily — owners now receive multiple searcher letters monthly and reply rates keep drifting down.

A different construction method

Start from the whole population, not an index of it. Translate every thesis judgment into website-visible evidence. Screen out companies that structurally cannot transact. Rank on documented evidence, not intuition. The method works at manual or machine scale.

Step one: establish the denominator

Before ranking anything, answer the census question: how many companies actually fit inside my search box? Not how many rows an export returns — how many exist. The denominator changes strategy, not just morale.

Real numbers from a single industrial census
367,478
global domains in category
17,300
live US operating companies
93 – 702
eligible per subvertical
~150
Thesis too narrow
May not survive normal response rates — better to learn in week three than month fourteen.
~650
Sweet spot
Disciplined search contacts all of them in a year — careful sequencing and high-effort letters.
4,000+
Needs prioritization
You will never work the list uniformly — scoring and tranching become essential.
A denominator comes from sweeping a whole category — association lists, license registries, certification directories, and map searches at manual scale; classifying every web domain at machine scale. Either way, the census is a one-time cost that repays itself every week.

Translate the thesis into evidence a website can supply

A typical thesis reads “B2B services, $1–5M EBITDA, recurring revenue, owner-operated, defensible niche” — zero of those are database fields. Rewrite each phrase as a claim a company's website could prove or fail to prove, and you have your scoring rubric.

Recurring revenue
Maintenance agreements, PM programs, scheduled calibration, consumables reordering, “preventive maintenance plans” on services pages.
Owner-operated
Explicit founder or family self-description plus a visibly thin management bench. “Family-owned and operated” is identity language.
B2B focus
Named commercial customers, industrial case studies, and the absence of residential service pages.
Defensible niche
Certifications with real switching costs — ISO/IEC 17025 for calibration, AS9100 for aerospace, ASME stamps, UL 508A for controls.
Size band (EBITDA proxy)
Headcount claims, facility sq ft, fleet photos, branch count. Use proxies, write down error bars, let the first call do what screening cannot.
Independence
No “division of”, no group branding, no PE portfolio page. Footer and about page tell the story.
EBITDA is not on websites. Any vendor claiming to estimate financials from page text is guessing with confidence. Size proxies sort a universe into tranches reliably; the exact number is a first-call question.
Thesis phraseWebsite-visible evidenceWhere it lives
Recurring revenueService agreements, PM programs, consumables, contract languageServices pages, footer menus
Owner-operated“Founder-led”, “family-owned”, founding-year narrative, thin named benchAbout, history, team pages
B2B focusNamed commercial clients, industrial case studies, no residential offerCase studies, industries served
Defensible nicheAccreditations: ISO/IEC 17025, AS9100, ASME, UL 508A, NADCAPCertifications page, footer badges
Size band ($1–5M EBITDA proxy)Headcount claims, facility sq ft, fleet, branch countAbout, careers, facility pages
IndependenceNo “division of”, no group branding, no PE portfolio pageFooter, about, news pages

Screen for independence before you rank anything

Roughly one in ten keyword-perfect candidates turns out to be group-owned on inspection — branches of consolidated groups, franchise locations, and PE-owned operators still running legacy websites. Run the independence screen first because it is cheap relative to what it saves.

Where ownership discloses itself
  • Footers: “a division of” or “part of the X family”
  • About pages: “in 2021 we joined…”
  • News pages: acquisition announcements
  • Careers pages: redirect to parent's ATS
Reading rescues targets too
“The company was acquired by the Miller family in 2001” is a family purchase, not a consolidation. Two decades later it is exactly the independently held business your thesis wants. Naive keyword matching on “acquired” throws it away — reading keeps it.

Score on evidence, and weight the scores deliberately

With the universe censused and non-independent companies removed, rank the survivors. Every score should carry its receipts — the verbatim quotes that earned it, with the page they came from.

70%
Mandate fit
A company doing the wrong work is worthless however approachable it looks
20%
Outreach suitability
Identifiable decision-maker, reachable channel
10%
Transition context
Capped so a romantic story never outranks the thesis
Transition context is not succession prophecy. It means founder-associated, long-established, independently positioned businesses with an identifiable decision-maker — assembled entirely from published site content. Owner age, health, and intentions do not belong in a screening file.

Five signals that carry the most weight in a search screen

Our full framework reads fifteen signals per company; for a classical search thesis, five of them do most of the work. Each is worth understanding well enough to extract by hand.

Founder-led / family-led association

Explicit self-description — “founder-led”, “family-owned”, “second generation” — is the top tier, reliable because firms publish it deliberately as identity. In our industrial runs, roughly half of confirmed fits carried explicit founder or family evidence.

Visible leadership bench depth

One named owner-operator and nobody else means the seller and the operator are the same person — the classic search target and classic key-person risk. A site naming a CFO, ops director, and sales manager describes a business that can survive its founder’s exit — different deal, different transition plan, different letter.

Recurring-offering indicators

Look for the vocabulary of committed repeat revenue: maintenance agreements, PM programs, calibration recalls, consumables operations, rental fleets with service attached. This is the signal most correlated with what search investors underwrite, and it is almost entirely textual — invisible to firmographic filters.

Operating history & continued independence

Founding year plus independence language together define the seasoned-but-unconsolidated company a search wants. Record the stated year verbatim — long history with no independence language warrants a closer footer read; short history means a different conversation.

Website/news activity trajectory

Dead companies keep live websites for years. Before a company earns outreach, check the most recent dated artifact — a news post, a job opening, a copyright year — because a letter to a ghost costs the same postage as a letter to a prospect.

A worked example: one thesis, censused

Our published specimen includes a precision machining thesis run across the US — 702 eligible independent job shops after triage and independence screening. The specimen format shows a deliberate cross-section, not a highlight reel.

8
Top fits
Scored 95+ on mandate fit, 16 verified snippets each
5
Keyword-missed fits
Homepages lacked category keywords — caught by reading whole sites
5
Documented exclusions
Keyword-perfect candidates excluded, mostly group-owned
2
Insufficient evidence
Flagged honestly rather than guessed
That four-way split — fits, hidden fits, exclusions, unknowns — is what a census actually looks like. Any list lacking the last two columns has quietly hidden its own error rate.

From ranked list to outreach engine

Load the ranked universe into your CRM with scores and evidence quotes as fields — not attachments. The evidence quote is the personalization no mail-merge produces and no owner mistakes for spray.

Tranche the list

Top decile: researched, individually written letters. Middle: strong templates personalized from evidence. Tail: efficient sequences.

Log and learn

Record every call verdict back against the universe. If wrong-size rejections cluster, recalibrate proxies. If “not now” owners share a profile, move that profile up next quarter.

Refresh quarterly

New companies form, owners sell, sites go dormant. A 12-month-old list loses several percent of accuracy — and watching which fits get acquired is free competitive intelligence.

What this method will not do

A website census tells you what companies publish, and nothing else. The census answers “who exists and what do they say about themselves” — the questions before and after remain yours.

  • Will not tell you revenue, EBITDA, or what a business is worth — use size proxies and confirm by phone
  • Will not tell you whether an owner wants to sell — no visible signal supports that claim
  • Will not read minds about succession — only what a company states about its own structure and history
  • Degrades gracefully on thin websites: a few percent flagged insufficient-evidence, where the honest output is a phone call, not a guess
  • Will not replace judgment about thesis quality — a perfectly censused bad industry is a well-organized list of companies you should not buy

Build it yourself, or buy the machinery

Everything above runs at manual scale. A searcher with a narrow thesis can census it from registries, associations, and map sweeps, read every site personally, and hold a better list than any subscription export. The method is the value; the machinery is a multiplier.

Manual scale
  • Best for narrow theses (~172 companies)
  • Census from registries, associations, map sweeps
  • Read each site personally over a few weeks
  • Cost: your time only
Machine scale
  • Starts from 100M+ classified domains
  • LLM analysis of every candidate against your thesis
  • Full universe with verbatim evidence for every call
  • ICP re-runs included as thesis sharpens
  • Search fund engagement from €4,900

The sequence — census, translate, screen, score, work the evidence — is free, and it works.

Frequently asked questions

The honest answer is: all of them — the full eligible universe for your thesis, which in our specimen industrial verticals ran from 93 to 702 independent US companies per subvertical. The practical follow-up is prioritization, not truncation. A 600-company universe worked in evidence-ranked tranches outperforms a hand-picked 150 because the companies you would have cut are disproportionately the ones nobody else is contacting. If your census comes back under roughly 150 eligible companies, treat that as a thesis-viability warning rather than a convenient workload.

For a narrow, well-bounded thesis, yes. Combine trade association member directories, state licensing registries, certification-body directories, distributor and OEM dealer locators, and systematic map searches by metro area. Merge, deduplicate, then read every site against your written rubric. Expect two to five minutes per company for triage and fifteen to thirty for full evidence extraction on survivors. It scales to a few hundred companies of patience; beyond that, machinery earns its keep. Our method page documents the full pipeline if you want to replicate its structure.

Triangulate visible proxies and accept the error bars. Stated headcount, named staff counts on team pages, facility square footage, fleet size in photos, branch counts, and shift language (“three shifts”) each bound the plausible size range. A company citing 65 employees and a 44,000-square-foot facility sits in a knowable band. What we will not do — and recommend you distrust from anyone — is convert page text into revenue or EBITDA estimates. The proxies sort a universe into size tranches reliably; the exact number is a first-call question.

Convergence happens when everyone queries the same index with the same filters. A census diverges in two places: the companies that databases never indexed — small operators with no news trail — and the companies whose sites never use the category's obvious keywords, which in our runs is a fifth or more of confirmed fits. Those two pools are where reply rates are highest, precisely because the owners are not receiving three searcher letters a month. Coverage, not cleverness, is what remains proprietary in 2026.

Screen for what is visible and refuse what is not. Founder-associated, long-established, independently positioned businesses with an identifiable decision-maker and a limited visible leadership bench — every element of that description comes from published site content, and together they define the transition-relevant profile. Owner age, health, retirement timing, and willingness to sell are not website-visible, and any list claiming to score them is guessing. Our transition-context score is capped at 10% of the total for exactly this reason: it is context, not prophecy.

Keep reading

Search fund target lists (service)Screening out false positivesFounder-led website signalsMapping an industry's TAMPrecision machining universeOur method
What we refuse to sell: no “ready to sell” flags, no revenue or EBITDA guesses, no owner-age profiling, no distress detection — and no engagements in consumer-captive verticals. Read our standards; serious buyers tell us this page is why they trusted the rest.

See what the full universe looks like for your thesis

One email. We send the specimen report the same day — every company scored and ranked with full signal transcripts, plus the exclusions we documented and why.

Request the specimen report