Method · Two passes, fifteen signals, three scores

Read every website. Extract the same signals. Keep the evidence.

Every candidate website is read against your written thesis — first in a fast triage pass, then in a deep extraction pass. Every claim in the deliverable traces back to a sentence a company published about itself.

2-passscreening architecture
15signals per company
3weighted composite scores
100%verbatim, sourced evidence
100M+ classified domains
24.7M business & finance sites
700+ industry categories
300+ organisations run on our data
Evidence for every call
The starting line

Screening begins with the entire category, not a curated list

We operate several AI platforms and maintain large-scale specialized datasets. More than 300 enterprise organisations run on our data — among them one of Europe's largest telecom operators, adtech and cybersecurity platforms, and a leading airline metasearch.

That classification of 100M+ active domains is where every engagement starts. Your thesis maps to categories; we pull every domain in scope. Nobody pre-filtered the list, so nothing was silently missed before you arrived.

Full-universe screening, defined

Starting from every classified domain in a category — not from a database of companies somebody previously found, tagged, and decided was worth keeping.

One real category, end to end

Domains in one industrial category, worldwide367,478
US-focused candidate set entering triage25,000
Live, operating companies after triage~17,300
Eligible independents across ten subverticals3,419

Every drop between rows is logged with a reason — nothing exits silently.

The pipeline

Two passes, five stages, no black box

A fast pass over everything, a deep pass over survivors. That split is what makes reading tens of thousands of websites economically sane — without ever sampling or skipping.

1

Universe pull

Your thesis maps to categories in the 100M+ domain classification. Every domain in scope is pulled — typically tens of thousands.

2

Triage pass

An LLM reads every homepage: operating company or directory? Independent or branch? Which subvertical, really? Dead domains fall away, with reasons.

3

Deep extraction

Survivors get a full crawl — about, team, history, services, certifications, careers — and structured extraction of all 15 signals.

4

Scoring

Three weighted composite scores rank the universe against your thesis. Group-owned companies zero out of outreach, whatever their fit.

5

Analyst QC

Evidence is re-verified against site text, exclusions are audited, and the package ships: ranked CSV, report, exclusion log.

Pass one — triage reads

Is this an operating company at all — or a directory, a marketplace, a parked domain, a blog?
Is it independent, or visibly a branch, subsidiary, or group member?
Which subvertical does the work actually belong to, keywords aside?
Which country does the page itself say the company operates from?

Pass two — extraction reads

The full site: about pages, team pages, history, services, case studies, certifications, careers.
All 15 signals, each recorded with a verbatim snippet and its source URL.
Subvertical-specific signals defined for your engagement — accreditations, stamps, authorizations.
A "not visible" record wherever the site is silent — never a guess.
Why not one deep pass over everything? Because most of a raw category is noise — directories, distributors' microsites, dead pages. Triage spends seconds on those so extraction can spend real attention on companies that might actually matter to you.
One specimen run, in numbers

The scale the method is built to survive

These are the real numbers from the published industrial-services specimen — the same run the sample report is drawn from.

0
Domains in the category, worldwide
0
US-focused domains entering triage
0
Live operating companies extracted
0
Signals per company, evidence attached
Ranked CSV, all signals Specimen-style report Verbatim evidence log Documented exclusion log Custom ICP re-runs
The framework

The 15-signal framework

Every signal is extracted only from what a company publishes about itself. Where the site is silent, we record "not visible" — never an estimate, never an inference from a name or a photo.

01

Founder-led / family-led association

Reads: explicit founder or family language — "second generation", "founder-led", "family-owned" — captured as the exact phrase on the page.

Why it matters: roughly half of confirmed industrial fits carry this evidence, and outreach to an identifiable owner is a different conversation entirely.

02

Visible leadership bench depth

Reads: how many principals are actually named on the site, and whether the visible bench extends beyond one or two people.

Why it matters: a shallow published bench shapes transition context; a deep one signals professionalized management worth meeting.

03

Operating history & continued independence

Reads: stated founding year, years in operation, and independence language, taken from the company's own telling of its history.

Why it matters: long-established independents are the population most theses actually target — and the hardest to find in profile databases.

04

Strategic fit to thesis

Reads: subsector, services, and customer types, matched line by line against the buyer's written thesis.

Why it matters: this is the principal signal — it drives the Mandate Fit score that carries 70% of the ranking weight.

05

Geographic & branch footprint

Reads: headquarters, branch locations, and the stated service area — with owned locations distinguished from partner mentions.

Why it matters: footprint decides whether a company is a platform, an add-on, or out of territory before anyone books a call.

06

Service-led vs product-led model

Reads: whether revenue visibly comes from field service, manufacturing, distribution, or a hybrid of them.

Why it matters: two companies with identical keywords can have opposite business models — this signal is how a thesis avoids buying the wrong one.

07

Recurring-offering indicators

Reads: service contracts, maintenance agreements, scheduled programs, consumables — the published language of repeat revenue.

Why it matters: recurring language on the site is the closest website-visible proxy for revenue quality a screen can honestly offer.

08

Vertical specialization & end-market exposure

Reads: customer industries evidenced by case studies and named end markets — not guessed from homepage keywords.

Why it matters: documented aerospace, food, or utility exposure is what separates a niche specialist from a generalist with a good copywriter.

09

Acquisition-program / roll-up readiness

Reads: signs the company already belongs to a group or runs its own acquisition program — press pages, "part of the X family" lines, investor tabs.

Why it matters: this is usually a disqualifier. About one in ten keyword-perfect candidates fails here and becomes a documented exclusion.

10

Management professionalization

Reads: visible non-founder functions — finance, operations, HR, marketing roles named on the site.

Why it matters: named second-layer management tells you how much of the company walks out the door with the founder.

11

Hiring posture & functional investment

Reads: open roles and their types, from careers pages and posted listings on the site itself.

Why it matters: what a company hires for is a candid window into growth, capacity constraints, and operational maturity.

12

Website / news activity trajectory

Reads: the most recent dated content anywhere on the site — news posts, project updates, certifications with issue years.

Why it matters: visibly active and visibly dormant companies deserve different outreach priority, and the difference is checkable.

13

Partner & channel ecosystem position

Reads: named OEM partnerships, distributorships, buying groups, and association memberships published on the site.

Why it matters: channel position often is the moat in industrial services — and it transfers, or doesn't, in an acquisition.

14

Compliance & regulated-market readiness

Reads: explicit certifications only — ISO, AS9100, ITAR, ISO/IEC 17025, ASME — captured as the exact claim text.

Why it matters: certifications gate entire end markets. A claim recorded verbatim can be checked; a checkbox in a database cannot.

15

Digital-commercial maturity

Reads: online quoting, customer portals, e-commerce, pricing transparency — how the company sells through its own site.

Why it matters: digital maturity is a cheap, honest indicator of how investable the operation is beyond its machines and vans.

Beyond the fifteen

Subvertical-specific signals, added per engagement

The fifteen common signals are the floor, not the ceiling. Each engagement adds signals that only make sense in your niche — the accreditations, stamps, and authorizations that decide who is real in that trade.

Automation & control panel builders

Industrial Automation & Controls Integrators
UL 508A panel shop listing Named PLC platform partnerships SCADA / HMI integration scope In-house panel fabrication

Calibration & metrology labs

Calibration, Metrology & Industrial Testing
ISO/IEC 17025 accreditation Published accreditation scope Accreditation body named Stated turnaround commitments

Boiler, burner & steam services

Boiler & Steam System Services
ASME code stamps R-stamp repair authorization Combustion tuning programs 24/7 emergency coverage

Coatings & surface finishing

Industrial Coatings & Surface Finishing
NADCAP accreditation AS9100 / aerospace approvals OEM process specifications In-house testing capability

Equipment dealers & material handling

Material Handling & Lifting Equipment Services
OEM dealer authorizations Factory-trained technicians Rental fleet on site Planned maintenance programs

Each added signal follows the same discipline as the core fifteen: exact claim text, source URL, and a "not visible" record when the site doesn't say. The signal list is agreed with you before extraction starts — it is your thesis, operationalized.

The ranking

Three scores, one deliberately lopsided weighting

Fifteen signals collapse into three composite scores. The weighting is intentionally unbalanced: fit does the work, suitability gates the list, and context is capped so it can never flatter a weak match.

70

Mandate Fit — 70%

How closely the company matches the written thesis: subsector, service model, recurring offering, footprint, compliance posture. The score that does the heavy lifting, and the one your ICP re-runs re-weight.

Principal score
20

Outreach Suitability — 20%

Is there an identifiable decision-maker and an independent owner to talk to? Founder or family association, visible principals, and no existing group ownership feed this score.

Gating score
10

Transition Context — 10%

Long operating history, limited visible bench, self-described generational ownership — website-visible context only, and deliberately capped at a tenth of the total.

Capped score

The zero-out rule: a company already owned by a group or consolidator scores zero on outreach, whatever its fit. Roughly one in ten otherwise-perfect candidates fails exactly here — and each becomes a documented exclusion, not a silent omission.

Evidence discipline

Every claim traces to a sentence a company wrote

Deliverables don't say "founder-led: yes". They quote the line, name the page, and link the URL — so your associate can check our work in one click. Below, real entries from the published specimen, anonymized for this page.

Target M-01Precision machining
"The company is a fourth-generation, family-owned precision CNC machining company."Industries page · signal: founder / family association
16/16 evidence snippets verified against site text
Target C-01Calibration & testing
"The company has been offering 17025-accredited calibration services you can trust since 1982."Homepage · signals: compliance + operating history
14/14 evidence snippets verified against site text
Target A-01Automation integration
"UL 508A — Industrial Control Panels Certification for USA."Certifications page · signal: subvertical-specific compliance
15/15 evidence snippets verified against site text
Excluded M-01Documented exclusion
"The company is a wholly owned subsidiary of a global industrial group."About page · rule: group ownership zeroes outreach
Excluded with reason — logged, not deleted
The "not visible" rule

If a signal is not on the website, the field says "not visible". It is never estimated, never borrowed from a third-party profile, and never inferred from a name, a photo, or a hunch.

Quality control

What happens between extraction and your inbox

LLM extraction at scale is powerful and fallible in equal measure. The QC pass exists because we assume errors, then go looking for them.

1

Snippet re-verification

Every quoted snippet is matched programmatically against the crawled page text. The verification count ships in the report — "16/16 verified", or honestly, "11/13". A snippet that can't be found is removed, and the signal reverts to "not visible".

2

Analyst review of top ranks and all exclusions

A human analyst reads the top of the ranked list and the complete exclusion log. Misclassified subverticals, mistaken independence calls, and over-generous fits get corrected before anything ships.

3

Zero-out audit

Every company flagged as group-owned is spot-checked against its own site, because a false ownership call costs you a real candidate. The one-in-ten exclusion rate is measured, not assumed.

4

Package assembly

Ranked CSV, specimen-style report — 8 top fits, 5 keyword-missed fits, 5 documented exclusions, 2 insufficient-evidence cases — plus the full evidence and exclusion logs, delivered CRM-ready.

Honest edges

What this method cannot do

A method you can trust has edges you can see. These are ours, stated plainly rather than discovered mid-engagement.

It only sees what companies publish

No financials, no revenue estimates, no valuation guesses. A website shows what an owner chose to say — the method refuses to pretend otherwise, which is why the standards page exists.

It never detects intent

We do not claim to know whether any owner would entertain a conversation. No website signal supports that claim, and vendors who sell it are selling a guess.

Thin sites limit extraction

A two-page website yields two pages of evidence. Those companies are flagged "insufficient evidence" — two per specimen report — rather than scored on imagination.

Websites drift

Sites change after we read them. A screen is a snapshot; the monitoring tier exists precisely because universes go stale. Re-screens catch what single passes cannot.

Related reading: the claims we refuse to make and the verticals we decline to serve are published in full on our standards page. The limits are the product.
Method FAQ

Questions deal teams ask about the pipeline

The specimen report answers most of these with worked examples — these are the short versions.

How does a written thesis become a screen?
You send the thesis as you'd write it for your IC — subsectors, service models, geographies, must-haves, dealbreakers. We translate it into category scope for the universe pull, prompts for both passes, and the weighting inside Mandate Fit. You review that translation before extraction starts, so the screen tests your thesis rather than our paraphrase of it.
What happens when a signal isn't visible on a site?
The field records "not visible" and the composite scores treat absence as absence — not as a soft negative, and never as an invitation to guess. This costs us tidy-looking completeness and buys you a dataset where every populated field is checkable. In practice, the pattern of what a company chooses not to publish is often informative in itself.
How do you keep an LLM from inventing evidence?
By never trusting it. Every snippet the model attributes to a page is re-matched against the crawled text of that page, and the match rate is printed in the deliverable — 16/16, 14/14, or honestly lower. Snippets that fail matching are dropped and their signals revert to "not visible". Hallucination isn't argued away; it's filtered out mechanically.
Why two passes instead of one deep read of everything?
Cost and attention. A raw category is mostly noise — directories, marketplaces, parked domains, branch pages. Deep-reading all 367,478 domains in the specimen category would spend most of the budget on sites that fail in the first sentence. Triage spends seconds each on those; extraction then gives real attention to the roughly 17,300 that survive.
Can the signal set be customized for our niche?
Yes — that's standard, not an upgrade. The fifteen common signals stay fixed so results are comparable across runs, and each engagement adds subvertical-specific signals: UL 508A for panel builders, ISO/IEC 17025 scope for calibration labs, ASME stamps for boiler services, NADCAP for finishing, OEM authorizations for dealers. The added signals follow the same evidence rules.
How current is the underlying domain data?
The 100M+ domain classification is updated continuously as part of the data business that 300+ organisations already rely on — it is not rebuilt per engagement. Site content is crawled fresh at extraction time, so signals reflect the website as it stood during your run, and the monitoring tier re-screens quarterly with monthly deltas.

See the method applied to a real category

One email. We send the specimen report the same day — every company scored and ranked with full signal transcripts, including the ones we refused to include and why.

Request the specimen report