To set up waterfall enrichment, you define the fields you need, filter inputs to your ICP, then route each field through a ranked chain of data providers, cheapest and broadest first. You add confidence thresholds and a stop condition so the waterfall stops paying once a verified value is found.
Then you merge results into one golden record, verify on a sample, and measure match rate and cost-per-match before you automate. This guide walks every step.
Waterfall enrichment setup at a glance
Here’s the whole setup in seven moves. Each line is one sentence, so you can scan it before you build anything.
| Step | What you do |
|---|---|
| 1 | Define the exact fields you need, and the canonical format for each. |
| 2 | Filter inputs to your ICP and required keys before you spend a credit. |
| 3 | Pick and order providers per field by coverage and cost, cheapest and broadest first. |
| 4 | Set confidence thresholds and a stop condition, so the chain stops on the first verified value. |
| 5 | Dedup and resolve conflicts into one golden record, using recency and source reliability. |
| 6 | Verify emails and phones, then test the whole flow on a sample. |
| 7 | Measure match rate and cost-per-match, then automate, sync to CRM, and refresh on a cadence. |
The body below splits these into eight H2 steps. That’s because verify-and-test and measure are each big enough to stand alone. The box keeps seven for scannability, which is intentional.
What “setting up” a waterfall enrichment workflow actually means
Setting up a waterfall enrichment workflow means building a ranked chain of data providers that fill each field in sequence, stopping the moment a verified value appears. So instead of trusting one data provider, you let several work in order.
The first cheap, broad source clears most records. Pricier, niche sources only run on what’s left.
Here’s why teams bother. A single source typically returns email match rates around 50-70%, and mobile phone coverage often sits lower. Single-source email coverage runs roughly 50-70%, while waterfall pushes toward 85-95%, and for phone numbers single-source covers 20-40% of records while multi-provider cascading reaches 50-70%.
Those gaps are exactly why teams stack providers. No single vendor owns every region or every data type, so you combine specialists.
If you want the full background before you build, here’s what waterfall enrichment is and how it compares to single-source. That concept piece covers the why. This guide covers the how.
The short version: single-source fill rates of 30-50% on contact data leave half your list unreachable. Reps then waste hours on manual research, or they skip those records entirely. Stacking providers in a waterfall closes that gap, and it does it without you buying two full subscriptions that each cover only part of the market.
This guide maps eight steps. First you define fields, then filter inputs, then pick and order providers. Next you set fallback logic, build the golden record, verify on a sample, and measure.
Finally you automate and refresh. Each step ends with a decision rule you can act on today.
You know the destination. So let’s start where every good setup starts, with the fields themselves.

Step 1: Define the exact fields you need
Start by listing only the fields that drive a decision: work email, mobile, title, seniority or function, employee count, industry, HQ country, tech stack, and LinkedIn URL. If a field never touches scoring, routing, or personalization, leave it out. Every extra field you chase costs credits and adds noise.
This is the foundation of how to set up waterfall enrichment well. Most teams skip it and pay for it later. A bloated schema means more conflicts to resolve, more verification to run, and a bigger bill for data nobody acts on.
Set a canonical format per field
Next, set a canonical format for each field. Store HQ country as an ISO code, not “USA” one row and “United States” the next. Bucket employee count into bands, like 51-200 or 201-500, instead of raw numbers that never match across providers.
Normalize titles into a seniority tier and a function. So “VP of Sales“, “Vice President, Sales”, and “Sales VP” all collapse into the same canonical value. This makes your data enrichment outputs comparable across vendors, and it makes your match rate honest rather than inflated by near-duplicates.
Field mapping starts here too. The canonical name you choose now is the name you’ll map to your CRM later. Pick it once, write it down, and reuse it everywhere downstream.
This pays off in CRM enrichment especially. When every provider’s “headcount”, “employees”, and “company_size” all map to one canonical field, your golden record assembles cleanly. Skip this, and you’ll spend Step 5 untangling three names for the same thing.
Keep the schema minimal for compliance and cost
There’s a compliance angle to a tight schema. Data minimization is a real principle, not a nicety. Enriching 50 attributes per contact when you only use 5 in your sequences runs against the data minimization principle.
So a lean schema keeps you cleaner under GDPR and CCPA. It cuts your bill at the same time.
💡 Pro Tip: Write your field list as a schema table first, with the field name, the canonical format, and the one downstream use. If a field has no use, delete the row. That table becomes the spec your whole waterfall enrichment setup follows.
When I built my first waterfall in 2022, I enriched 14 fields because the provider offered them. We actually used four. The other ten just inflated cost-per-match and made conflict resolution harder later.
Now I cut the schema before I touch a single provider.
Decision rule: enrich the minimum viable schema, not everything.
You’ve defined the fields. So which records actually deserve to enter the chain?
Step 2: Filter inputs before you enrich
Filter your inputs first, because enriching junk is the fastest way to burn budget. Drop any record missing the keys providers need to match: a domain, or a name plus company. Without a usable key, a provider returns nothing, and you still pay for the attempt on some pricing models.
So a pre-call field check pays for itself fast. Run it once at the top of the workflow, before any provider sees the record. It’s a tiny rule that prevents a steady drip of wasted credits across every batch.
Suppress out-of-ICP records
Then suppress out-of-ICP records. If a contact sits outside your ideal customer profile, enriching it changes no downstream decision. So why pay for it?
An ICP filter at the front of the workflow is one of the cheapest wins in the whole setup. It shrinks the data enrichment input set before a single credit moves. The smaller and cleaner your input, the more your match-rate numbers actually mean.
Audit data quality before the first credit
Quality matters as much as fit. Before you enrich, audit your data quality first so malformed domains and duplicate rows don’t poison your match-rate numbers. A clean input set makes every later metric trustworthy.
⚠️ Watch Out: A blank or malformed domain looks enrichable but isn't. It quietly drags your fill rate down and inflates your error count. So validate the key field before the record ever reaches provider one.
🔍 Did You Know? Sales reps spend just 28% of their week actually selling, with the rest consumed by tasks like data entry and research. Filtering out records nobody will ever work protects that scarce selling time.
Decision rule: enrich a record only when it changes a downstream decision.
Your input set is clean and in-ICP. Now comes the choice that decides your whole cost structure: who fills each field, and in what order?
Step 3: Pick and order providers per field by strength and cost
Order your data providers by coverage and cost, cheapest and broadest first, with premium or niche vendors last. The cheap, broad source clears most records before an expensive vendor ever bills. Brand reputation is the wrong sort key, and I learned that the hard way.
When I built that 2022 waterfall, I ordered providers by reputation. The premium vendor sat first, so it billed every record before the cheap one ran. Our cost-per-match tripled.
Reordering by coverage-and-cost fixed it in an afternoon.
Run a coverage audit by data type and geography
Run a coverage audit before you rank anything. Test each data provider by data type and by geography. A vendor strong on US contact data can be weak on EMEA firmographics, so one global ranking hides the truth.
Here’s the shape of a simple coverage audit. Score each provider per field, per region, then rank from there.
Treat firmographic data and contact data as separate columns, because a vendor that nails company size may be weak on direct mobiles. The audit is where a record-level instinct gives way to field-level discipline.
| Provider | Email (US) | Email (EMEA) | Mobile | Firmographic | Relative cost |
|---|---|---|---|---|---|
| Provider A | High | Medium | Low | High | Low |
| Provider B | Medium | High | Medium | Medium | Medium |
| Provider C | Low | Low | High | Low | High |
Use field-level cascades, not record-level
Now the move most guides skip: use field-level cascades, not one record-level chain. Build a separate ranked order for email, a separate one for phone, and a separate one for firmographic data.
A record-level chain re-queries a provider for fields it already filled, which wastes calls. Field-level routing sends each field down its own best path instead. This single design choice is where a lot of wasted spend disappears.
📌 Example: Your email order might run cheap-broad source, then a verification-focused finder, then a premium fallback. Your mobile order might invert that, because the vendor weakest on email is often strongest on phone. Same record, two different provider paths.
Where CUFinder fits, honestly
This is where CUFinder fits, and I’ll be honest about it. CUFinder works best as an affordable first or anchor node in a waterfall, not as a turnkey multi-vendor orchestrator. Its contact enrichment can clear a big share of records cheaply up front, so pricier vendors only run on the remainder.
Coverage and match rates vary by region and vertical, though. So test it on a sample before you commit it to a slot.
When you’re choosing vendors to slot in, the best waterfall enrichment tools is a useful shortlist to compare against your audit. Clay, for instance, documents conditional runs so a provider only fires when the prior field is blank. In Clay you configure run settings, including conditions for when the waterfall should run.
Decision rule: rank providers per field, cheapest and broadest first, premium last.
You’ve got an ordered chain per field. But what tells the chain when to advance, and when to stop paying?
Step 4: Set fallback logic, confidence thresholds, and stop conditions
Set fallback logic so the chain advances to the next provider only when a field is empty or below a confidence threshold. If provider A returns a verified work email, provider B never runs for that field. That single rule is where waterfall enrichment saves real money.
This is the step where most of the cost control in how to set up waterfall enrichment actually lives. Fields and providers matter, but the rules that govern when each provider fires are what keep your spend predictable.
Make the stop condition the heart of it
The stop condition is the heart of the whole design. The waterfall must stop the instant a verified value appears, on a per-field basis. Without it, you pay every vendor on every record, which defeats the entire point.
I shipped a waterfall once with no stop condition. It kept paying provider three and four after provider two had already returned a verified email. Adding stop-on-first-verified cut the bleed overnight, so now it’s the first rule I write, not the last.
Tune confidence thresholds by data type
Tune thresholds by data type. Set a strict confidence bar for email, because a wrong address bounces and hurts deliverability. You can run firmographics looser, since an approximate employee band rarely breaks a campaign.
A waterfall chains three to four providers in sequence so each fills gaps the last missed, and you verify everything before sending. The thresholds are what decide when “the last missed” actually triggers the next call.
💡 Pro Tip: Add a hard per-record provider cap, say three providers per field. It stops a stubborn record from cascading through ten vendors chasing a value that may not exist. Unify's setup guidance covers tuning these confidence thresholds and the cost benchmarks behind them, worth a read while you calibrate (Unify).
⚠️ Watch Out: Thresholds set too low never cascade, so you ship the first cheap guess. Set too high, and the chain runs every vendor needlessly. So calibrate against your sample test in Step 6, not by gut.
Decision rule: confidence threshold plus stop-on-first-verified plus a per-record cost cap.
Each field now has a verified value from somewhere in the chain. So how do you turn several provider outputs into one clean record?
Step 5: Dedup and resolve conflicts into one golden record
Merge every provider’s output into one golden record per contact, then resolve disagreements with explicit rules. First run deduplication, so the same person from three data providers collapses into a single row. Then handle the conflicts, because two providers will hand you two different titles.
Default to recency-wins
Use recency-wins as your default. The fresher value usually beats the stale one, since data decay is constant in B2B. B2B data decays at roughly 30-40% per year, as people change roles and companies restructure.
So a value crawled last week generally outranks one from last year. Recency is a sensible first tiebreaker because most conflicts are just one source lagging behind.
For persistent ties, fall back to a source-reliability score. Rank your data providers once on observed accuracy, then let the higher-ranked source win when recency can’t break the tie. This is overwrite governance, and it’s what turns merging from a guess into a rule.
Never let a weak source overwrite a strong one
The non-negotiable: never let a lower-quality source overwrite a higher-quality one. A cheap broad provider should not clobber a verified value from your strict email finder. Openprise’s guidance on assembling an enrichment waterfall covers this orchestration angle well (Openprise).
📌 Example: Provider A returns “VP Sales” crawled in March, while provider B returns an older “Sales Director” crawled in January. Recency-wins picks “VP Sales”, then you stamp the field with its source. So when a rep flags it later, you can audit exactly where it came from.
Stamp every field with its source
Stamp every field with its source provider. That source attribution lets you debug a wrong value months later, instead of guessing which vendor to blame. It also tells you which provider to swap when one slot underperforms.
CRM enrichment depends on this discipline. When the golden record flows into Salesforce or HubSpot, the source stamp travels with it, so your reps and your RevOps team can trust the lineage of any field.
Decision rule: explicit overwrite governance, recency first, then source reliability, with the source stamped per field.
You’ve got a clean golden record on paper. But is the data actually true? That’s the question Step 6 answers.
Step 6: Verify before use and test on a sample
Verify every email and phone before they reach outreach, because a high fill rate hides confidently wrong values. Validate emails on syntax, MX record, and catch-all status, and validate phones for format and line type.
A record that looks complete but bounces is worse than a blank, since it damages your sender reputation.
Run a truth-set sample test
Then test the whole data enrichment flow on a sample of 100 to 500 records against a known truth set. Run your real waterfall on data where you already know the right answers. Inspect fill rate per field and, more importantly, the false-positive rate.
Found emails that bounce at 20-30% degrade sender reputation and push campaigns into spam, so always include an email verification step and exclude catch-all domains when deliverability matters. The sample test is what surfaces that bounce risk before it touches a live domain.
We skipped the sample test once and pushed straight to a 40,000-record run. The fill rate looked great. But the emails were catch-all addresses that bounced, and the domain took a hit.
Now I never roll out without a truth-set sample first.
Check accuracy, not just match rate
🔍 Did You Know? Cognism's own guidance stresses that match rate and accuracy aren't the same thing, and that verification is what separates them (Cognism). A record can match and still be wrong.
Before you ship, run through a data enrichment checklist so verification, dedup, and source stamping are all in place. A checklist catches the step you skipped under deadline pressure.
Decision rule: no full run without verification and a truth-set sample.
Your sample passed. So before you scale, how do you know the workflow is actually paying off?
Step 7: Measure match rate and cost-per-match
Measure per field, not just overall, because one strong field can mask a weak one. Track overall fill rate, fill rate per field, cost per usable record, per-provider hit rate, and API error rate.
A 90% overall fill can hide a 40% mobile fill, and you’d never see it from the headline number. So the per-field view is not a nice-to-have. It’s the only view that tells you where to invest your next dollar.
Break cost-per-match down by field
Cost-per-match per field tells you which provider to swap. If your phone slot costs three times your email slot for half the coverage, that’s your next fix. So break the spend down by field and by vendor, every single run.
Set benchmark ranges so you know what good looks like. Running 1,000 contacts through a three-provider waterfall can lift verified email coverage from around 62% single-source to roughly 90%, at a comparable cost per usable record, near 16 to 17 cents.
So email fill in the 85-95% band is strong. Mobile fill in the 50-70% band is realistic, not low.
Review per-provider performance quarterly
💡 Pro Tip: Build a tiny dashboard with one row per provider: records attempted, hits, cost, hit rate. Review it quarterly. The provider that looked great at launch often slips as its coverage and your ICP drift apart.
🔍 Did You Know? One Datablist analysis found that moving from classic enrichment at 55% to waterfall at 80% can lift revenue by about 45%, simply because reps contact more prospects. The match-rate math compounds quietly through the funnel.
Watch your API error rate as closely as your fill rate. A provider returning errors isn’t free, since it still costs latency and sometimes credits.
It also quietly drags throughput down. So a rising error rate is often the first sign a vendor is about to slip in coverage.
Decision rule: measure per field and per provider, not just the overall number.
The numbers look right on the sample. Now you make it run without you.
Step 8: Automate, sync to CRM, and refresh on a cadence
Automate the data enrichment workflow through an iPaaS, a no-code tool, or direct API calls, then sync the golden record to your CRM. Map each data provider’s field names to your CRM fields once, carefully, so a “company_size” from one vendor lands in the right Salesforce or HubSpot property.
Bad field mapping is a silent source of dirty data. So get the field mapping right once, and CRM enrichment stays clean for every run after.
Schedule batch and real-time triggers
Schedule both batch and real-time triggers. Run nightly batches for bulk lists, and fire a real-time webhook when a new lead enters, so reps never wait on stale records. Use idempotency keys so a retried webhook doesn’t enrich the same record twice and double your bill.
A clean API key per provider, a webhook per trigger, and JSON payloads mapped to your schema: that’s the plumbing. Keep it boring and predictable, because clever automation is the kind that breaks at 2am.
For the CRM side, a reverse ETL tool like Hightouch can push the golden record from your warehouse into Salesforce or HubSpot on a schedule. So your real-time vs batch enrichment split stays clean: webhooks handle new leads, batches handle the backlog, and reverse ETL keeps the CRM in sync.
Refresh on decay, not on a whim
Then refresh on decay, not on a calendar whim. Re-enrich Tier-1 accounts around every 30 days, and lower tiers every 60 to 90 days. Roughly half of professionals change jobs every 18 months, and CRM records degrade steadily without refresh.
So the freshest tier gets the most frequent passes. That keeps your highest-value records accurate without re-enriching the long tail too often.
⚠️ Watch Out: Hard-coded integrations trap you. If swapping a provider means a rebuild, you'll tolerate a bad vendor far too long. So keep each provider modular, behind a clean interface, and you can drop one in or out without touching the rest.
I keep my waterfalls modular for exactly this reason. Last year a provider’s EMEA coverage slipped, and I swapped it in under an hour because nothing downstream was wired to its name. Modularity is cheap insurance.
Decision rule: refresh on decay, and keep every provider modular.
That’s the full setup. Now let’s watch one record actually travel through it.
Worked example: one record traveling the waterfall
Take one realistic record: a VP of Sales at a 200-person SaaS company, input as a name plus a company domain. Here’s the full trace, field by field, through the workflow you just built.
| Stage | What happens | Result |
|---|---|---|
| Fields needed | work email, mobile, title, seniority, employee count, industry, HQ country | schema set |
| ICP filter | domain present, SaaS, 200 employees, in-ICP | passes |
| Email chain, provider A | cheap broad source queried first | miss |
| Email chain, provider B | verification-focused finder runs because A was empty | verified hit |
| Stop condition | verified email found, email chain stops | provider C skipped |
| Phone chain, provider C | phone-strong vendor queried first | mobile hit |
| Firmographic chain | broad source fills employee count, industry, HQ | filled |
| Conflict | provider A title “Sales Director”, provider B “VP Sales” | recency-wins picks “VP Sales” |
| Golden record | merged, each field stamped with its source | assembled |
| Verify | email passes MX and catch-all check, phone format valid | clean |
| Cost | provider A attempt, provider B hit, phone, firmographic | logged per field |
What the field-level design bought you
Notice what the field-level design bought you here. The email chain stopped at provider B, so provider C never billed for email. The phone chain ran independently, because the email vendor was weak on mobile.
A record-level chain would have behaved differently. It would re-query providers for fields they’d already filled, wasting calls on data you already had. That’s the difference field-level routing makes on a single record, multiplied across every record in your run.
📌 Example: Because the title conflict resolved by recency, the rep saw "VP Sales" instead of the stale "Sales Director". The source stamp on that field means that if it turns out wrong, you know exactly which crawl to question. That's overwrite governance doing its quiet job.
This single trace is the clearest test of your setup. If you can follow one record from raw input to verified golden record and explain every provider call, your waterfall enrichment workflow is sound. If you can’t, find the step where the trace goes fuzzy and tighten it.
I run this trace on every waterfall I build, before the first real batch. It’s faster than reading dashboards, and it catches design flaws a fill-rate number never will. One record, fully understood, beats ten thousand records you can’t explain.
So that’s the happy path. Now the traps that knock it off course.
Common waterfall enrichment setup mistakes (and how to avoid them)
Most broken waterfalls fail in the same handful of ways. Here are the six I see most, each with its fix.

Record-level cascade instead of field-level
A record-level chain re-queries a provider for fields it already filled, so one missing mobile re-runs the vendor that already gave you the email. The fix is a separate ranked order per field. A record-level cascade once cost me about a third of my calls before I caught it.
No stop condition or confidence threshold
Without a stop rule, the waterfall keeps paying every provider after a usable value appears. With thresholds set too low, it never cascades at all. The fix is stop-on-first-verified, per-field thresholds, and a hard cap.
No overwrite governance
A cheaper or later source overwrites a higher-quality earlier value, and with no source attribution you can’t even tell what happened. The fix is recency-plus-reliability rules, with the source stamped on every field.
Skipping verification
A great fill rate can route reps to confidently wrong emails that bounce. The fix is to validate emails and phones, then sample-test against a truth set before any full run. Deliverability damage is slow to undo, so this one is worth the extra hour.
Enriching everything, including out-of-ICP records
This inflates cost-per-match for fields nobody uses. The fix is an ICP filter and a pre-call field check at the front of the flow. Filter first, enrich second.
Set and forget
No monitoring of per-provider hit rate, cost, or API errors, plus hard-coded integrations that can’t swap a vendor. The fix is to monitor, refresh on decay, and stay modular. A waterfall is a living system, not a one-time build, so it needs a quarterly review like any other piece of revenue infrastructure.
⚠️ Watch Out: These mistakes compound. A record-level cascade with no stop condition and no governance can quietly triple your bill while degrading your data. So fix them in order, starting with the cascade structure.
You know the build and the traps. Now the questions teams ask most.
FAQ
How do I set up waterfall enrichment in Apollo?
In Apollo, you turn on its waterfall enrichment feature, which routes records through Apollo’s data plus partnered third-party sources when Apollo’s own data misses. You map your input fields, enable the cascade, and Apollo fills gaps from additional providers automatically. Apollo’s knowledge base documents the exact toggle and setup (Apollo).
What is Clay waterfall enrichment, and how do I build one in Clay?
Clay waterfall enrichment chains multiple providers with conditional run logic, so each provider fires only when the prior field is empty. You add an enrichment column per provider, then set a condition like “run only if email is blank.” Clay’s docs walk through building the data waterfall step by step (Clay).
How do I turn off waterfall enrichment in Apollo?
You disable it in Apollo’s enrichment settings by toggling off the waterfall or third-party data option, so enrichment uses only Apollo’s native data. Teams turn it off to control cost, to avoid third-party sources for compliance reasons, or to keep one consistent data source. Check the same settings panel where you enabled it.
What does waterfall enrichment mean?
Waterfall enrichment means querying data providers in a ranked sequence until a verified value is found for each field, then stopping. Instead of one source, several run in order, cheapest and broadest first. Each data provider fills gaps the previous one missed, which lifts overall match rate well above any single vendor.
That cascading logic is what separates waterfall from flat data enrichment.
What are the best waterfall enrichment tools?
The best tools depend on your ICP and geography, not a universal ranking. Orchestrators like Clay and Apollo handle the cascade natively, while providers like ZoomInfo, Cognism, Clearbit, People Data Labs, BetterContact, LeadMagic, and Persana fill specific slots. Run a coverage audit by region and data type, then pick the providers that win on your actual list.
How much does waterfall enrichment improve match rates?
Waterfall enrichment typically lifts email coverage from a single-source 50-70% to roughly 85-95%, and mobile from 20-40% to 50-70%. Single-source providers generally achieve 40-60% match rates, while a waterfall can reach 80% or higher, often 30-40% more complete records. Gains vary by region and vertical, so measure on your own data.
How many providers do I need in a waterfall?
Two to four well-chosen providers cover most use cases. Adding too many sources creates operational complexity and cost without proportional coverage gains, so start small, measure, then add only if a real gap remains. Begin with one strong broad data provider, add a fallback, then layer a niche vendor only where your audit shows a hole.
Is waterfall enrichment GDPR-compliant?
It can be, provided every provider in your cascade collects data under a valid legal basis. Review each source’s compliance policy before integrating it, and apply data minimization by enriching only the fields you actually use. Consent-aware enrichment and a tight schema keep you defensible under GDPR and CCPA.
The bottom line
How to set up waterfall enrichment comes down to a repeatable sequence, not a tool purchase. You define the fields, filter to your ICP, order providers per field cheapest-first, set a stop condition, build a governed golden record, verify on a sample, measure per field, then automate and refresh.
The real advantage lives in three places: the stop condition that ends the bleed, the field-level order that kills wasted calls, and the sample test that catches confidently wrong data. So start from the fields and the use case, not the vendor. Get those eight steps right, and any stack you build on top will hold.




