Most TAM figures are extrapolations wearing a decimal point.
A full-web census counts the companies themselves — every domain in the category triaged, classified, and segmented, so the market size you present is a list you can drill into, not a number you inherited.
Somewhere in most investment memos sits a market-size figure with a strange property: nobody in the room can say what it counts.
It arrived from an analyst report, which derived it from another report, which built it from a survey sample, a growth assumption, and a segmentation invented for a different question.
The number is defensible only in the sense that everyone else cites it too.
Ask the follow-up that actually matters to a deal team — how many companies is that, and which ones — and the report has no answer, because it never knew.
For fragmented service and industrial markets, the gap between report math and reality is at its widest. Survey-based sizing systematically under-counts the long tail of small independents, because small independents do not answer surveys.
Database-derived counts inherit their vendor's coverage: leading company databases index the companies they found, so a "market map" built on one is really a map of the index, with the thinnest coverage exactly where fragmentation — and opportunity — is greatest.
And standard industry codes mangle service businesses so reliably that even a perfect count of the wrong category would mislead: the water-treatment service firm coded as a plumbing supplier, the automation integrator filed under electrical contracting.
The strategic questions a TAM is supposed to answer are segment questions anyway.
An investor weighing a platform thesis does not need "the market is large"; they need the number of independent, service-led operators in scope, the consolidation rate consolidators have already achieved, and how thickly candidates cluster by region.
A corporate strategist weighing entry needs the split between product-led and service-led players and where the whitespace sits.
None of that survives aggregation into a single headline figure — which is why we stopped treating market sizing as arithmetic and started treating it as a census.
Every count decomposes into a company list with evidence per row. “221 independent water-treatment operators” is checkable — open the file, read the quotes. A number that survives diligence questions because it is made of answers to them.
Service-led versus product-led, independent versus group-owned, by region, by end market: each split is classified company by company from site evidence, so segment shares are measurements rather than allocation assumptions.
A purchased market report is frozen the day it ships. A census is infrastructure: re-run it next year, diff it against baseline, and market growth, consolidation, and churn become observed quantities instead of modeled ones.
Search any IAB industry category and see how many domains our classification database maps to it — the starting population before any thesis-level screening begins.
Our census work starts from the whole web rather than any index: 100M+ classified domains — including 24.7M business and finance sites — sorted into 700+ industry categories, covering 99.99%+ of active internet usage.
Your market definition maps onto those categories, and every domain in scope enters the funnel. From a recent industrial run, the actual numbers at each stage:
| Stage | Count | What happens |
|---|---|---|
| Category universe | 367,478 domains | One industrial engineering category, globally — the raw census frame before any judgment is applied |
| Triage pass | 25,000-domain US-focused sample | Pass one separates live operating companies from directories, marketplaces, dead sites, and out-of-scope pages |
| Live operators | ~17,300 companies | Confirmed operating businesses — the population deep extraction will read in full |
| Segmented & sized | 93–702 per subvertical | Ten subverticals counted: precision machining 702, equipment repair 545, automation integration 534, material handling 513, compressed air 276, calibration & testing 254, water treatment 221, boiler & steam 172, filtration 109, surface finishing 93 — eligible independent US companies each |
Two passes, deliberately. Triage is cheap and broad so that no candidate is excluded by a keyword's absence;. deep extraction is expensive and careful so that every retained company carries structured fields.
business model, end markets, footprint, ownership language, certifications. each backed by a verbatim quote and source URL.
The eligibility rules are stated, not implicit: when a company is excluded as group-owned, out of geography, or out of scope, the reason is logged with evidence.
Across these industrial runs roughly 1 in 10 keyword-perfect candidates fell out as already group-owned — which is not list shrinkage, it is the consolidation rate being measured.
And because the extraction is criteria-driven rather than keyword-driven, re-running a subset against a revised definition — narrower geography, a different independence rule, your own ICP — is included, not a new project.
The census applies the full 15-signal framework to every company; these five do the most work in converting a list into the segment tables a memo actually cites.
The census question is always someone's question — a fund's platform thesis, a strategist's category definition, a lender's collateral market — and fit is scored against that written definition, not against an industry code standing in for it.
This is what makes the boundary of the market defensible: when someone challenges why a company is inside or outside the count, the answer is a scoring rationale quoting the company's own description, not a taxonomy accident.
The market's edge cases stop being noise and become documented decisions.
The single most requested segmentation, and the one standard codes are structurally unable to deliver.
Whether revenue visibly comes from field service, manufacturing, distribution, or a hybrid changes margin structure, capital intensity, and buyer appetite — a “market” that mixes the four is several markets wearing one name.
We classify the model from what each site actually describes: crews and service areas, plants and product lines, brands carried, or some measured blend.
Segment shares in the deliverable are sums over those per-company classifications.
Consolidation state is the market-structure fact deal teams most want and reports least deliver.
Reading every company's site surfaces the acquirers' own announcement language — “now part of”, “a portfolio company of”, “completed twelve acquisitions” — so the census yields a measured roll-up penetration rate per subvertical and a named list of active consolidators.
For a platform thesis this is the temperature of the market: how much of it is already taken, and how fast the takers are moving.
Geo-segmented TAMs need stated footprints, not headquarters pins. We extract branch pages, declared service areas, and facility claims, distinguishing owned locations from partner networks, so regional counts reflect where companies actually operate.
Density maps built this way show clustering that HQ-based data hides — and for market-entry or add-on strategy, the sparse regions are often the finding: whitespace is a geographic statement, and it should rest on evidence.
Maturity structure separates a market of durable operators from one of recent entrants, and independence language — founding years, generational wording, “independent” and “family-owned” statements — feeds the independent-versus-consolidated split.
In our industrial censuses roughly half of confirmed independent fits carry explicit founder or family evidence on their sites; that share is itself a market-structure statistic, and one that matters to anyone whose thesis depends on a supply of long-established independents.
Three layers, one file set. The top layer is the market table: counts by subvertical and segment, the splits by business model, independence, and region, and the measured consolidation rate.
The middle layer is the company-level data behind every cell — one row per company, classified and scored, with quoted evidence per field.
The bottom layer is the audit trail: the eligibility rules as applied, the exclusion log with reasons, and the insufficient-evidence tier honestly separated.
Specimen reports use the same anatomy — 8 top fits, 5 keyword-missed fits, 5 documented exclusions, 2 insufficient-evidence flags per subvertical — so you can judge the discipline on a live sample before commissioning anything.
The keyword-missed tier deserves a word, because it is where census logic pays for itself in sizing terms.
Across our industrial runs, a fifth or more of confirmed fits lacked the category's obvious homepage keywords.
companies like the machining specimen whose homepage says only "family-built, American-owned, making metal work in America since 1965." Any keyword-bounded method undercounts the market by roughly that tier, and undercounts it non-randomly: the missing companies skew older, busier, and less marketed, which for most acquisition theses is the most interesting corner of the population.
To make the abstraction concrete, take the smallest-but-one segment from the industrial run: calibration, metrology, and industrial testing.
The census sized it at 254 eligible independent US companies — a number that would be hard to source anywhere else, because the segment sprawls across half a dozen industry codes and most of its members are small labs with modest websites.
The structure underneath the count is what a strategy or deal team actually uses. The confirmed fits read like this:
The keyword-missed tier held, among others, an accredited metallurgical testing lab whose homepage leads with failure analysis rather than any calibration vocabulary — a company a keyword-framed count simply never sees.
And the exclusion log measured the segment's consolidation directly: four of the five documented exclusions were group-owned, with the acquirers' language quoted.
"privately held by a national distribution group," "part of the family of companies." That ratio, read across the whole segment, is the consolidation-rate line in the market table;.
the names behind it are the active-consolidator list. One subvertical, one page of the deliverable — repeated for every segment in the frame.
Fund teams put census counts in IC memos and LP materials, where "we counted them;. here is the file" outperforms any citation to a market report.
and the same file then becomes the sourcing universe, so sizing and pipeline mapping are one purchase rather than two.
Corporate strategy teams use segment tables for entry and whitespace decisions, then hand the company layer to corporate development. Advisors cite census figures in mandate materials with a methodology note their clients can interrogate.
And diligence teams use a census to pressure-test a target's own market claims: when management says "we compete with a handful of regional players," the census says whether that is true.
Commercially it is project work: a Proof project from €4,900 establishes the method on one category; a full universe with deep shortlist runs from €9,900; annual monitoring from €18,000 per thesis keeps the census live with quarterly re-screens and diffs.
For the method in longer form, the guide on mapping the TAM of an industry for acquisitions walks through a full run.
Format follows the destination.
The market tables arrive as presentation-ready exhibits; the company layer ships as structured files built for whatever holds your pipeline — a CRM import, a data-room upload, or the model your analysts already maintain.
Because every row carries its evidence, the census survives handoffs: the strategy team that commissioned it, the deal team that inherits it, and the diligence provider who later stress-tests it are all reading the same auditable base rather than three generations of reformatted summary.
That persistence is worth more than it sounds — most market sizing dies in the appendix of the deck it was made for, while a census keeps working because the next question ("fine, so who do we call first?") is already answered by the same file.
A web census counts the web-visible market.
In B2B services and industrials that is very nearly the whole market — operating companies without any web presence are rare and getting rarer — but "very nearly" is a boundary we state rather than extrapolate past.
The census counts companies, not revenue: we do not estimate financials from page text, so a census yields market structure in units of firms, segments, and footprints, and if you need a currency-denominated TAM you will be combining our counts with financial assumptions you control.
explicitly, on top of an auditable base, instead of implicitly inside someone's report.
Classification is honest but not omniscient: companies that publish too little land in a flagged tier rather than being guessed into a segment, and websites lag reality by months, which is what scheduled re-screens are for.
If a market definition depends on facts websites never state, we will say so before you spend — the comparison guide on databases versus full-web mapping covers where each method's edges sit.
One email. We send a specimen census the same day — a subvertical sized company by company, with the evidence, the exclusions, and the boundary stated.
Request the specimen report