Standard industry codes answer standard questions.
When a thesis lives or dies on a distinction no code set contains — field-service-led versus depot-led, OEM-authorized versus independent aftermarket — we take your scheme, sharpen it into screenable definitions, and classify the universe into it, with quoted evidence behind every assignment.
Standard codes flatten the splits that define a thesis — field-service versus depot, OEM-authorized versus independent, certified versus uncertified. Every one of those distinctions is visible on company websites; none exists in any standard schema.
Industry codes cannot separate a distributor with a service arm from a service firm that resells parts, or a UL 508A panel builder from an electrician with a website. Your thesis lives on those splits.
Route density, authorization status, certification tiers, revenue texture — companies publish these distinctions to win customers. We read them at census scale and classify every company into your scheme.
Force-fitting codes blurs the thesis. Hand-coding drifts between analysts. DIY with an LLM discovers the real work: definitions, edge cases, quality control, and re-applying refinements consistently.
700+ categories maintained across 100M+ domains — the same machinery, loaded with your scheme instead of ours. Your categories, your definitions, applied uniformly with evidence on every assignment.
The difference between a taxonomy that works and one that embarrasses everyone is entirely in the definitions. We spend the front of every engagement making yours precise enough to screen with — and honest enough to admit what a website cannot show.
Each category gets a written definition with inclusion rules, disqualifiers, and the visible markers that evidence membership — the language, pages, and claims a site would actually carry. Criteria that are not website-visible get flagged now, not discovered after the run.
Before anything runs at scale, we classify a few hundred real candidate sites against the draft scheme and review the results together.
Ambiguous cases become explicit edge rules; categories that will not separate get merged or redefined. The scheme that goes to production has already survived contact with reality.
The finished scheme applies across the whole target universe through the same two-pass pipeline as our standard categories — every assignment carrying its quote, page reference, and confidence grade.
Two passes turn a raw domain pool into evidenced category assignments. Your taxonomy's decision rules run against full site extractions, not homepage snippets.
Separates live operating companies from directories, parked domains, and dead sites. In one industrial category, 367K domains narrowed to ~17,300 genuine candidates.
Reads services pages, about pages, case studies, certifications, team pages. Your taxonomy's rules run against the full extraction — every assignment records verbatim evidence and source page.
Category membership and acquisition scores come from one reading: Mandate Fit 70%, Outreach Suitability 20%, Transition Context 10%. Group ownership evidence zeroes the total.
Sites without enough evidence are graded insufficient-evidence — reported honestly alongside fits and exclusions. An auditable "cannot tell" beats a coin-flip assignment.
Because the scheme is code plus definitions, it stays alive. Re-runs with refined ICPs are included; definition changes re-apply across everything already classified.
Most custom schemes are not exotic — they are ordinary strategic distinctions that standard data simply does not carry.
Field-service-led versus depot-led versus resident-technician programs. The economics differ completely — route density, response-time commitments, installed-base intimacy — and websites state the model plainly to their customers even though no database field captures it.
OEM-authorized versus independent aftermarket, quoted from authorization claims and partner pages. For consolidation theses in compressed air, material handling, or power transmission, this single split often defines the buy-box — and the pricing logic.
Categories bounded by compliance claims: ISO/IEC 17025-accredited labs versus uncertified calibration shops, ASME-stamped fabricators, NADCAP finishing houses, UL 508A panel builders. Exact claim text, captured per site, becomes a hard category edge.
Contract-and-program operators versus project-and-job operators, read from the language of maintenance agreements, chemical service programs, rental fleets, and scheduled offerings. The distinction most industrial buyers ask about first, applied as a classification rather than a hunch.
Who the customer base actually is, evidenced from case studies and named clients rather than inferred from keywords — municipal versus industrial water treatment being the classic example where the same equipment serves two entirely different theses.
Manufacturer, distributor, integrator, service provider, or hybrid — assigned from what the site describes doing, not from a self-selected directory label. The most common source of misfiled companies in standard data, and usually the first dimension a scheme needs.
Custom categories are not extracted from thin air — they compose from the same 15-signal extraction that runs on every deep-analyzed site. Five signals carry most custom schemes.
In a custom engagement this signal is literally your taxonomy — the classifier is built from your words, stress-tested with you, and every assignment traces back to a definition you signed off. Ambiguities surface as documented edge cases, not silent inconsistency.
The workhorse of industrial taxonomies: whether revenue arrives through field service, manufacturing, distribution, or a hybrid is stated on nearly every industrial site. The same "pump company" keyword covers a manufacturer, a distributor with a repair bench, and a field-service operation — no thesis treats those three alike.
Authorization-based categories classify cleanly because companies advertise their channel positions — "authorized distributor," "certified service center," named OEM partnerships. We quote claims verbatim; for aftermarket and channel theses, this signal usually is the taxonomy.
Certification claims make the sharpest category edges because they are binary, dated, and checkable: a lab either publishes ISO/IEC 17025 accreditation or it does not. Our capture of exact claim text keeps the edge auditable years later.
End-market categories fail when guessed from keywords, so we build them from documented exposure — case studies, project galleries, named customer industries. The confidence grade is part of the category assignment, not decoration.
The specimen work we publish is itself a custom taxonomy.
One buyer-shaped scheme split US industrial services into ten subverticals — precision machining, equipment repair, automation integration, material handling, compressed air, calibration and testing, water treatment, boiler and steam, filtration, surface finishing — none of which maps one-to-one onto any standard code.
Applied through the two-pass pipeline, the scheme produced eligible independent US counts per category: 702 in precision machining, 545 in equipment repair, 534 in automation integration, down to 93 in surface finishing.
Real denominators, per custom category, each company carrying its evidence.
The interesting rows are the ones a keyword approach gets wrong. A fifth or more of confirmed fits in these runs lacked their category's obvious homepage keywords — the specimen file's "hidden fit" class.
One equipment-repair fit's homepage never says the category's name; its About page evidence reads "Founded in 1991 by the founder and the founder", and its services live three clicks deep.
A calibration-testing hidden fit is a materials lab whose Careers page. not its homepage.
carries the line "we are a family-owned business." Keyword tagging files these companies under the wrong label or misses them entirely;. definition-driven classification, reading full sites, puts them where they belong.
The exclusions are equally instructive: about one in ten keyword-perfect candidates classified out as group-owned, evidenced in the companies' own words — "privately held by a national distribution group", as one excluded calibration business states.
A taxonomy that cannot document its exclusions is a list with opinions; the specimen format exists to show both directions of the discipline.
The fair comparison is not "custom taxonomy versus nothing" — it is against the three approaches teams actually use today.
| Approach | Distinction fidelity | Scale | Evidence | Maintainability |
|---|---|---|---|---|
| Standard industry codes | Coarse; thesis-critical splits usually absent | Universal | None — assignments are self-reported or inferred | Static by design |
| Keyword tagging | Brittle; misses the fifth of fits without obvious keywords, misfiles hybrids | High | The keyword itself, which proves little | Every refinement is a new regex debate |
| Manual analyst coding | High on a good day; drifts between analysts and across months | Hundreds of companies, then fatigue | Whatever the analyst noted | Re-coding after a definition change rarely happens |
| Full-web LLM taxonomy | Your definitions, applied uniformly; ambiguity surfaced as edge cases | Thousands to full-web | Quote + source page + confidence, per assignment | Definition changes re-apply across the classified universe |
Some category ideas cannot be built honestly, and the useful moment to hear that is before the run.
Categories defined by non-public facts — revenue bands, precise headcount, margin profile, backlog — are not classifiable from websites, and we will not pretend otherwise by proxy-guessing; that refusal is the same one our standards apply to every deliverable.
Categories that depend on intent or interior state ("companies open to partnership") are stories, not classifications.
And every scheme has an ambiguity floor: some real companies genuinely straddle two categories, and a hybrid distributor-integrator does not stop being hybrid because a schema would prefer it picked a side.
We handle those with explicit multi-membership or documented tie-break rules — but a scheme whose categories overlap heavily will produce arguments, not insight, and we say so in the stress-test phase.
Two more boundaries.
Coverage is web-visible coverage: a company with no meaningful web presence is outside the census, which in B2B is a thin sliver but not zero, and we state the boundary rather than extrapolating across it.
And confidence grades are real: an insufficient-evidence assignment means the site did not publish enough to classify, and buying a taxonomy from us includes accepting that class in the output.
typically a low single-digit percentage of deep-analyzed sites, honestly labeled instead of forced into a bucket.
The scheme and its classified output are engagement deliverables, confidential to you — including the definitions themselves, which by the end of the stress test usually encode real strategic thinking.
Output arrives as structured files keyed by domain: category assignments, confidence grades, evidence quotes, source pages, plus the standard signal extractions underneath.
Teams load it into CRMs as custom fields, into BI tools as the segmentation layer for market sizing, or into sourcing workflows as the buy-box filter that finally matches the thesis language in the IC memo.
A taxonomy also compounds.
The same scheme that sizes a market this quarter can drive an ICP discovery run for the commercial team next quarter and become the category backbone of ongoing monitoring after that — new companies classified on arrival, category migrations flagged as deltas.
Engagements start at proof-project scale (from €4,900, one scheme on one subvertical) so the definitions can prove themselves on a bounded universe before you commit them to a full category;.
details on pricing, and a same-day specimen report shows the evidence format before anything is signed.
One email starts the definition workshop. The specimen report shows the evidence format the same day — including how we classify the companies that refuse to fit neatly.
Request the specimen report