Customer Data Platforms (CDP)
What a CDP actually is
A Customer Data Platform (CDP) is packaged software that ingests customer data from every source a business has — web events, app events, CRM records, support tickets, POS transactions, ad platform data — and stitches it into a single, persistent, unified customer profile. The defining trait isn't the ingestion pipeline; it's that the resulting profile is accessible to other systems. A CDP isn't a reporting tool that sits at the end of the pipeline — it's an identity and data layer that sits in the middle, feeding audiences and profiles out to ad platforms, email tools, personalization engines, and analytics warehouses.
The term gets abused in marketing copy constantly. The real test of 'is this actually a CDP': does it (1) ingest data from multiple first-party sources, (2) resolve those events to a single persistent customer identity, and (3) let you activate that unified profile in downstream tools without engineering work each time?
CDP vs CRM vs DMP — the confusion that never dies
These three acronyms get used interchangeably by people who haven't had to implement any of them, but they solve different problems. A CRM (Salesforce, HubSpot) is a system of record for known, named relationships — it stores deals, contacts, and sales activity, and its data is almost entirely first-party and identified. A DMP (Data Management Platform, e.g. legacy Adobe Audience Manager) was built for anonymous, cookie-based advertising audiences, historically blending first- and third-party data, and typically only retains data for a short window (30-90 days) because it existed to build ad segments, not customer relationships. A CDP sits between them: it's built for marketers (not just IT), unifies both known and anonymous behavioral data under one identity, and persists that data indefinitely as a durable asset — not just an ad-targeting cache.
CDP vs CRM vs DMP
| Dimension | CDP | CRM | DMP |
|---|---|---|---|
| Primary data type | First-party behavioral + identified | First-party, identified relationships | Mostly third-party/anonymous, cookie-based |
| Data persistence | Long-term, durable profile | Long-term, durable record | Short-lived (30–90 day cookie windows) |
| Owner/user | Marketing, analytics, growth teams | Sales, account management | Media/advertising teams |
| Primary use case | Unified profile + activation across tools | Managing deals and relationships | Building ad audience segments |
| Identity resolution | Core function | Manual/rule-based dedupe | Probabilistic, cookie-based, weak |
| Examples | Segment, mParticle, RudderStack, Tealium | Salesforce, HubSpot, Zoho | Adobe Audience Manager (legacy), Oracle BlueKai |
Identity resolution: the actual hard problem
The value of a CDP lives or dies on identity resolution — the process of deciding that an anonymous website visitor with cookie ID abc123, a mobile app user with device ID xyz789, and a CRM contact with email jane@company.com are all the same human. This happens through identity graphs built from deterministic and probabilistic matching.
Deterministic matching links identities using a hard key both records share — the same email address, the same hashed phone number, the same login/user ID. It's high-confidence but requires that key to actually appear in both records, which anonymous traffic doesn't have. Probabilistic matching infers a likely match from signals like shared IP address, device fingerprint, and behavioral pattern similarity when no deterministic key exists — useful for cross-device stitching, but inherently a confidence score, not a certainty, and increasingly limited by privacy-driven restrictions on fingerprinting.
anonymous_id: cookie_a1b2c3 (first website visit, no login)
|
| user logs in with email jane@company.com
v
known_id: user_9981 <-- email: jane@company.com
| phone: +1-555-0142 (added later, deterministic)
|
| same email seen on mobile app SDK
v
device_id: ios_44f2 --> merged into user_9981
Result: ONE unified profile with 3 merged source identities,
feeding a single "customer_id" out to every downstream tool.Bad identity resolution is worse than none
Real-time vs batch CDPs
CDPs split into two architectural families. Real-time (event-streaming) CDPs — Segment, mParticle, RudderStack — sit in the request path as data is generated: a website or app SDK sends events to the CDP first, which resolves identity and forwards the event to downstream destinations within seconds. This is the model needed for real-time personalization (e.g. showing a different homepage banner mid-session) or triggering an abandoned-cart email minutes after the event.
Batch (warehouse-native) CDPs — Hightouch, Census, and the broader 'Composable CDP' movement — instead run on a schedule (hourly, daily) against data already sitting in your data warehouse (BigQuery, Snowflake). They don't intercept live traffic; they read your existing warehouse tables, resolve identity there, and sync resulting audiences out to activation tools via reverse ETL. This model has grown fast because it avoids duplicating your data into a second system and keeps the warehouse as the single source of truth.
Real-time vs Composable (batch) CDPs
| Aspect | Real-time CDP (Segment, mParticle) | Composable/batch CDP (Hightouch, Census) |
|---|---|---|
| Where identity resolves | In the CDP's own event pipeline | In your data warehouse via SQL models |
| Latency | Seconds | Minutes to hours (sync schedule) |
| Good for | Real-time personalization, live triggers | Audience syncs, ad platform uploads, scheduled workflows |
| Data ownership | Data also lives in vendor's infrastructure | Data stays in your own warehouse |
| Setup complexity | SDK/event instrumentation across every surface | SQL model + existing warehouse pipeline |
The 'Composable CDP' pitch is really reverse ETL
Choosing between the major tools
Segment (Twilio) is the market incumbent — broad destination catalog, mature SDKs, but priced for scale and increasingly positioned as an enterprise product post-acquisition. mParticle leans mobile-app-heavy and enterprise, with strong data quality/governance tooling. RudderStack is the open-source-friendly alternative, self-hostable, and notably cheaper at volume since it can run on your own warehouse infrastructure rather than charging per tracked user. Tealium has deep roots in tag management and is common in large enterprises that grew out of a TMS deployment.
The decision usually comes down to: how much real-time activation do you actually need versus how much of this could be solved with warehouse SQL and reverse ETL, and how much budget exists for a dedicated identity layer versus building it yourself.
What's next
A CDP is only as good as the data flowing into it — and campaign attribution data is one of the most common (and most poorly governed) inputs. Understanding how UTM parameters populate traffic-source data is the natural next step before layering a CDP on top of it.
Next: UTM Parameters & Campaign Tracking →
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.