Use case — CRM & database gap analysis

Your CRM holds the companies everyone found. Audit it against the companies that exist.

Years of exports, banker books and badge scans have filled your pipeline with records — and quietly fixed its boundaries. We screen the entire active web against your thesis and show you exactly what your list is missing, with evidence.

1 in 5+confirmed fits lack the obvious keywords
1 in 10keyword-perfect records already group-owned
100%of calls shipped with quoted evidence
0
Classified domains screened
0
Business & finance sites
0
Domains in one industrial category
0
Live operators confirmed in one triage
The inherited pipeline

Every record in your CRM arrived through a door. Every door has a shape.

None of these sources is wrong. Each one is a filter that was applied before you ever saw the company — and the filters compound. A gap analysis measures what they filtered out.

Database exports

Profile databases index companies they previously found and tagged. Your export inherits their coverage, their taxonomy and their refresh lag — invisibly.

Inherits vendor blind spots

Banker books

Processes you were shown, teasers you passed on, lists assembled for someone else's mandate. Curated for a fee, not for your thesis.

Curated for other buyers

Conference scans

Exhibitors and attendees from the shows your team happened to attend. A marketing-budget sample of the market, three years stale.

Samples marketing budgets

Inbound & referrals

Companies that found you, or that a contact remembered. Valuable — and structurally biased toward the visible and the already-networked.

Biased to the visible
What a gap run is

An audit, not another list

Leading company databases start from the companies they found. We start from the entire active web — 100M+ classified domains, 24.7M of them business and finance sites — and run LLM analysis against your written thesis.

Gap analysis, defined

"Your target list, reconciled against a full-web screen of the same thesis: every fitting company you hold confirmed and enriched, every fitting company you lack surfaced, and every non-fitting record you hold flagged — each with quoted evidence and a source URL."

The output is three clean buckets, not a pile of lookalikes. Your team decides what to do with each bucket; nothing is overwritten, and nothing arrives without its reason attached.

Mechanics

Four passes from your CSV to a reconciled universe

The same two-pass architecture behind our full-web method — a broad triage, then deep extraction — bracketed by a normalization step in and a reconciliation step out.

1 · Normalize

We ingest your CRM export and database lists, resolve domains, and deduplicate. No credentials needed — a CSV of names and websites is enough.

2 · Triage

The relevant slice of the 100M+ universe is cut by category and geography, then triaged for live, operating, in-scope companies.

3 · Deep extraction

Survivors get full-site reading against your thesis: 15 signals, three scores weighted 70/20/10, and verbatim evidence per claim.

4 · Reconcile

The scored universe is matched back to your records. Every company lands in one of three buckets, every bucket ships with its proof.

The coverage question

What your stack answers today — and what it can't

This is the honest comparison. Your CRM plus profile databases are genuinely good at some of these rows. The rows they can't answer are where mandates quietly leak.

Coverage questionCRM + database exportsFull-web gap run
Starting universeThe companies someone previously found — thousands of records100M+ classified domains; 24.7M business & finance sites; 700+ categories
How a company gets inRegistered, exhibited, got profiled, or met your teamOperates a live website. That is the entire entry requirement.
Fits without the obvious keywordsInvisible — search runs on labels the company never usedRead, not searched: a fifth or more of confirmed fits lack their category's obvious homepage keywords
Group-ownership screeningRarely recorded; usually discovered mid-outreachExplicitly screened — roughly 1 in 10 keyword-perfect candidates is already group-owned
Evidence per recordA profile compiled somewhere, sometime, by someoneVerbatim snippet + source URL for every inclusion and every exclusion
RefreshWhenever the vendor or an analyst touches the recordRe-runs on any subset with a custom ICP; monthly deltas on the annual tier
The deliverable

Every record you hold — and every one you don't — lands in one of three buckets

Confirmed & enriched

Records in your CRM that survive the full screen — returned with 15 extracted signals, three scores, certifications as exact claim text, and quoted evidence, so your team can re-rank the pipeline before spending another outreach cycle on it.

Your best records, now defensible in an IC memo

Missing from your CRM

Companies that match your thesis and appear nowhere in your systems — consistently a larger bucket than teams expect, because many fits never used the keywords your existing sources search on.

A fifth or more of confirmed fits lack the obvious keywords

Held, but disqualified

Records you're carrying that fail the screen — group-owned subsidiaries, wrong service model, dormant sites — each with the documented reason, because deleting dead records is worth almost as much as finding live ones.

~1 in 10 keyword-perfect candidates: already group-owned
Scale, concretely

What one category actually contains

From a real run: a single industrial-services category held 367,478 domains globally. A 25,000-domain US-focused triage confirmed ~17,300 live operating companies — far more than any target list we have audited actually held.

367,478domains in the category, globally
25,000US-focused triage input
~17,300live operating companies confirmed
10subverticals with eligible independent counts
Precision machining702 Equipment repair545 Automation integration534 Material handling513 Compressed air276 Calibration & testing254 Water treatment221 Boiler & steam172 Filtration109 Surface finishing93
Signal deep-dive

The four signals that matter most in a CRM audit

All 15 signals ship with every record. These four do the heaviest lifting when the question is "what is wrong with the list we already have."

Acquisition-program / roll-up readinessThe silent disqualifier

The single most common defect in aged CRM data — a company absorbed into a group after its record was created, invisible until outreach. We read for subsidiary language, "part of the X family" phrasing, and acquisition announcements, zeroing the transition score on group ownership; roughly one in ten keyword-perfect candidates fails here.

Founder-led / family-led associationExplicit only

Database exports rarely carry ownership context; websites state it constantly — "second generation", "founder-led", "family-owned since 1981" — and we capture that language verbatim. In industrial categories, roughly half of confirmed fits carry explicit founder or family evidence on their own pages, making this the column your CRM is missing.

Operating history & continued independenceStated, not scraped

Founding year, decades-in-operation claims, and independence language, taken from what the company says about itself — separating a durable target from a two-year-old site wearing the same keywords, and correcting the quiet CRM failure where directory-sourced records carry no age context at all.

Website / news activity trajectoryDormancy detector

The most recent dated content on the site — posts, project pages, press items — tells you whether a company is visibly active or coasting on a 2019 rebuild, because aged pipelines accumulate dormant records and flagging them saves the outreach cycles that would have discovered the silence one email at a time.

From the specimen

What the three buckets look like on real companies

Our public specimen follows the same 8 / 5 / 5 / 2 discipline: eight top fits, five keyword-missed fits, five documented exclusions, two insufficient-evidence cases. Names are anonymized on the public site; clients receive source URLs.

Confirmed

Target M-01 — CNC, defense & medical devices

"a fourth-generation, family-owned CNC company"

Held in the client's CRM as a bare name. Returned with AS9100D and ITAR captured as exact claim text, founder context, and a 15-signal profile — the record went from a row to a case.

Missing

Hidden M-02 — stamping & assembly

"family-built, American-owned since 1965"

Absent from every list the client held. The site never uses the category's standard keywords — a keyword search cannot find it, and none had. Full-site reading scored it as a direct thesis fit.

Disqualified

Excluded M-01 — precision components

"a wholly owned subsidiary"

Sat in the pipeline for two years as an active target. One quoted line from its own About page ends the debate — already consolidated, transition score zeroed, record closed with the reason attached.

The uncomfortable arithmetic: a CRM audit almost never shrinks the pipeline. It swaps records you would have wasted cycles on for companies you did not know existed — and attaches proof to both moves.

What this does not do

Honest edges of a gap run

A gap analysis reads what companies publish. That boundary is what makes every claim checkable — and it excludes some things other vendors are happy to guess at.

No seller-intent flags. We do not claim to know who wants to transact. No website signal supports it, and pretending otherwise would poison the signals that are real.
No financial estimates. Revenue and earnings are not on company websites. Vendors who attach them to profiles are modeling, not observing. We decline to, and we say so in the deliverable.
No system-of-record surgery. We match on names and domains from your export — we do not log into your CRM, dedupe your contacts, or resolve entities against internal systems. You keep the keys.
Into your workflow

Built to land in the systems you already run

The deliverable is a flat, import-ready file plus a written read of the results. Most teams load the three buckets as list views and work them on different cadences.

  • CSV keyed on domain — the one identifier your CRM and our universe share
  • Bucket, three scores (70/20/10 weighting), and all 15 signals as columns
  • Evidence snippets and source URLs on every inclusion and exclusion
  • Custom ICP re-runs on any subset included — tighten the thesis, re-reconcile
  • Optional monthly deltas so the audit doesn't rot like the list it audited

A proof project from €4,900 audits one thesis-shaped slice of your CRM. The full universe with deep shortlist runs from €9,900; annual monitoring from €18,000 per thesis keeps it reconciled.

One row, reconciled

bucket: missing_from_crm
mandate_fit: 92.7  ·  outreach: 88  ·  transition: 71
founder_status: founder led — "the founder established the company in 1981"
certifications: UL508 — captured as exact claim text
evidence: 12/12 snippets verified against site text
Questions teams ask

CRM gap analysis, examined

What do you need from us to start — and do you need CRM access?

A CSV export with company names and websites, plus your thesis in plain language — no CRM access or credentials needed. Domains are the matching key; where your export lacks them, we resolve names to domains during normalization and flag ambiguous cases rather than guessing.

How is this different from just buying another database export?

An export adds rows from the same kind of source you already have — companies someone previously found and tagged — and cannot tell you what all such sources missed. A gap run starts from the entire active web, reads sites against your thesis, and reconciles to your records: an audit with evidence, not a bigger pile.

Our CRM has 4,000 target records. How big will the "missing" bucket be?

It depends on your niche's keyword density, and we won't pretend to know before running it — but the structural finding is consistent: a fifth or more of confirmed fits lack their category's obvious homepage keywords, so keyword-fed lists systematically under-cover. The proof project exists precisely to answer this on one slice before you commit to the full universe.

Will you tell us which of our targets are open to a transaction?

No, and we'd encourage suspicion of anyone who says yes — no website signal reveals intent, and we don't manufacture one. What we do document: founder-associated, long-established, independently positioned businesses with identifiable decision-makers, stated on their own pages and quoted verbatim, so your team draws conclusions from evidence.

What happens to the disqualified bucket — do you delete our records?

Nothing is deleted — every disqualified record ships with its documented reason and quoted source. Most teams archive them with the reason attached, which also stops the same companies re-entering the pipeline through next year's conference scan.

How current is the underlying data, and can we keep it current?

Every deep-extraction read happens at run time against the live site — no aged profile layer between you and the evidence. Annual monitoring (from €18,000 per thesis) re-screens the universe and delivers monthly deltas: new fits, ownership changes, and disqualifications reconciled to your list each cycle.

Adjacent use cases

If the audit reshapes the list, these keep it sharp

Held to our standards: no seller-intent flags, no financial guesses, no profiling of private individuals — in this deliverable or any other. Read the full standards.

Find out what your CRM doesn't know

Send us one thesis and one export. We'll return the three buckets — confirmed, missing, disqualified — with quoted evidence for every call, starting from €4,900.