Use case · Thesis-driven investors

Thesis universe mapping: know the whole denominator, not the visible slice

A search result is a sample. A universe map is a census — every company matching your written thesis, cut from 100M+ classified domains, with evidence attached to every inclusion and every exclusion.

100M+Classified domains at the start line
367,478Domains in one industrial category
~17,300Live operators after one 25,000-domain triage
3Scores per company — fit, outreach, transition
Data behind 300+ enterprise organisations
One of Europe’s largest telecoms runs on it
24.7M business & finance sites classified
700+ industry categories maintained
First, the term

What a thesis universe actually is

Most sourcing conversations skip this definition, which is how funds end up defending a “market map” that is really a list of companies a database happened to contain.

Thesis universe, defined

“The complete population of companies that satisfy a written investment thesis at a point in time — enumerated from the entire active web, with published evidence for every membership decision, including the negative ones.”

Two properties in that sentence do the work. Complete means the count is a denominator you can cite: not “we found 214 targets” but “214 of the 702 eligible independent US companies in this subvertical.”

Evidenced means each decision survives hostile review. When an IC member asks why a company is absent, the answer is a quoted sentence and a source URL — not a shrug.

The build

Four stages between the open web and your universe

Each stage narrows the population and raises the evidence bar. Nothing advances on a keyword match alone.

1
Classify

Classify the entire active web

Before any thesis arrives, the map exists: 100M+ domains classified into 700+ industry categories, covering 99.99%+ of active internet usage. Your universe is cut from this, not assembled by searching.

Scale example

One industrial category alone held 367,478 domains globally — before any thesis filter was applied.

Business layer

24.7M of the classified domains are business and finance sites — the layer acquisition theses actually draw from.

2
Triage

Triage: cheap breadth, honestly applied

A fast pass removes parked pages, directories, marketplaces, and non-operators. In the published specimen run, a 25,000-domain US-focused triage left roughly 17,300 live operating companies — the honest starting population.

What triage removes

Dead domains, resellers, pure directories, and sites with no operating business behind them.

What it never does

Triage never rejects on thesis grounds. Fit decisions belong to the deep pass, where evidence is captured.

3
Extract

Deep extraction against your written thesis

Every surviving company gets full-site LLM analysis against fifteen defined signals — founder association, leadership bench, recurring offerings, certifications, footprint, end markets, and more — each backed by verbatim snippets with source URLs.

Verbatim only

“Founded in 1981” is evidence. A hunch about company age is not, and never enters the file.

Exact claims

ISO, AS9100, ITAR, ISO/IEC 17025, ASME — captured as the exact claim text, page by page.

Source-linked

Every snippet carries its URL, so any reviewer can check any decision in seconds.

4
Score

Score and rank the survivors

Three weighted scores order the universe: Mandate Fit at 70%, Outreach Suitability at 20%, Transition Context at 10% — with an automatic zero on evidence of group ownership. The weights are published, so the ranking is arguable rather than mysterious.

The zero rule

Roughly 1 in 10 keyword-perfect candidates turned out to be group-owned — scored zero, documented, and filed as exclusions.

Ranked output

The specimen format ships 8 top fits, 5 keyword-missed fits, 5 documented exclusions, and 2 insufficient-evidence cases.

One real run

The specimen numbers, end to end

These figures come from the industrial-services run we publish openly — the same discipline applies to any thesis we map.

0Category domains, global
0Live operators after triage
0Eligible in the largest subvertical
0Eligible in the smallest subvertical

Between those extremes sat equipment repair (545), automation integration (534), material handling (513), compressed air (276), calibration and testing (254), water treatment (221), boiler and steam (172), and filtration (109) — ten defensible denominators from one map.

The undercount problem

Keyword search structurally misses your best names

Companies describe themselves for customers, not for your query — a precision machining shop leads with “aerospace components,” a water treatment operator with “plant reliability” — and in the specimen run a fifth or more of confirmed fits lacked the category’s obvious homepage keywords. Classification-first mapping finds them because it reads what companies do, not what they call themselves.

  • Keyword-missed fits are reported as their own named bucket — five of them in every specimen
  • Subvertical language handled explicitly: OEM authorizations, UL 508A, NADCAP, rental fleets, chemical service programs
  • Roughly half of confirmed industrial fits carried explicit founder or family evidence

Same universe, three verdicts

Target M-02 — founded 1942, second-generation ownerTop fit · 96.0
Hidden fit W-03 — no category keyword on homepageKeyword-missed
Excluded M-01 — group ownership quoted on siteExcluded · 0
Insufficient I-02 — site too thin to judgeHeld honestly
Evidence discipline

Exclusions documented as carefully as inclusions

Anyone can hand you a list of maybes. The harder deliverable is the documented “no” — every exclusion carries its disqualifying quote and URL, and every insufficient-evidence case is filed as exactly that instead of being silently dropped or optimistically included to pad the count.

  • Verbatim snippets and source URLs in client deliverables; anonymized on the public site
  • Sample verification rates published — e.g. “15/15 evidence snippets verified against site text”
  • No financial estimates, ever — websites do not publish financials, so neither do we

An exclusion, as delivered

Excluded R-01 — equipment repair, every keyword matchedZeroed
Disqualifying evidence, quoted

“The company, a subsidiary of a larger industrial group…” — captured from the target’s own site and filed with its source URL.

Why completeness pays

“We saw everything” is a sentence you can now defend

An IC debating a platform bet keeps circling one anxiety — what if the better asset is out there, unseen? A mapped universe retires that anxiety with a number and a method, and the same denominator settles market-structure questions from enumeration rather than analyst estimates: how fragmented the space really is, how many independents remain, and where they cluster.

In the IC memo, verbatim

Population: 702 eligible independent US precision machining companies, enumerated from 367,478 category domains
Method: two-pass screen, 15 signals, weights 70/20/10, group ownership zeroed
Audit trail: every inclusion and exclusion carries quoted evidence with source URLs
Quality bar

Five properties of a universe worth citing

Use these to interrogate any market map — ours included. A map that fails one of them is a list wearing a map’s clothes.

01

Enumerated, not sampled

The population is counted from a defined starting set — the classified web — not accumulated from searches until the analyst got tired.

02

Evidenced per decision

Each inclusion and exclusion carries quoted site text with a URL. If a decision cannot cite its evidence, it is an opinion.

03

Honest about ignorance

Thin or silent websites land in an insufficient-evidence bucket, reported as such — never guessed into either the fit list or the bin.

04

Re-runnable on demand

The thesis is a written instrument, so a pivot re-scores the same universe under a new ICP — included in every engagement, not sold as a change order.

05

Silent where evidence ends

No claims about anyone’s appetite for a deal, no financial guesses, no owner profiling. The universe states what companies publish — nothing further.

The uncomfortable arithmetic: if a fifth of true fits lack the obvious keywords and a tenth of keyword-perfect names are already group-owned, a keyword-built list is wrong in both directions at once — and nobody inside the deal team can see either error.

What arrives

The deliverable, itemized

A universe map is a working dataset, not a slide. Everything ships in formats your team can sort, filter, and import.

In every universe map

  • The full enumerated population with per-company signal extraction across all fifteen signals
  • Three scores per company — Mandate Fit, Outreach Suitability, Transition Context — with published weights
  • Verbatim evidence snippets and source URLs at cell level
  • Named buckets: top fits, keyword-missed fits, documented exclusions, insufficient evidence

And around it

  • Custom ICP re-runs on any subset, included — the map survives your next pivot
  • Universe statistics formatted for IC memos and LP updates
  • CRM-ready exports; see database gap analysis for reconciling against your existing pipeline
  • An optional standing refresh — the add-on radar — when the thesis becomes a program
Public pricing: proof project from €4,900 on a single subvertical; full universe with deep shortlist from €9,900; annual monitoring from €18,000 per thesis. Custom ICP re-runs are included at every tier. Full pricing →
Common questions

Universe mapping, asked and answered

How is a universe map different from a market map our analysts build?

Analyst maps are assembled forward — searches, directories, association lists — until the page looks full. A universe map is cut backward from a classified copy of the entire active web, so completeness is a property of the method, and the map reports its own edges including keyword-missed and insufficient-evidence buckets.

Our thesis is unusual. Can it be mapped at all?

Usually, yes — the screen runs on your written criteria (service mix thresholds, certification gates like ISO/IEC 17025 or ASME stamps, geography, independence), not a fixed taxonomy. The honest caveat: criteria that are not website-visible — margins, contract terms, culture — cannot be screened, and we will say so during scoping rather than pretend otherwise.

What does “evidence for every exclusion” mean in practice?

Every company that leaves the universe carries the reason, the disqualifying quote, and the source URL — and in the specimen run this discipline caught roughly one in ten keyword-perfect candidates as already group-owned, files that often earn their keep months later when a banker pitches you a “proprietary” name your exclusions already explain.

How current is the map on delivery, and how does it stay current?

Extraction runs against live sites during the engagement, so delivery-day currency is inherent. After that, annual monitoring re-screens the same universe on a quarterly cadence and ships changesets — new entrants, departures, score movements with their triggers — so funds running continuous programs typically pair the map with the standing radar.

Do you tell us which companies are open to a transaction?

No. Openness to a transaction is not a website-visible fact, and we do not simulate it. What the map can legitimately show is transition context from published evidence — founder-associated, long-established, independently positioned businesses with an identifiable decision-maker — describing what a company states about itself, never inferred personal circumstances.

What does a first engagement look like?

Most funds start with a proof project from €4,900 — one subvertical mapped end to end in the published specimen format — so you can judge evidence quality before committing. Full universe with deep shortlist from €9,900; annual monitoring from €18,000 per thesis. Send your thesis to start scoping.

Map the universe your thesis actually implies

Send the thesis as you have written it for your IC. We will come back with the population it defines, the method that cut it, and the evidence behind every line.

Our standards: no willingness-to-transact flags, no financial guesses, no owner profiling, and every exclusion documented. Read the full standards →