CRM Data Modeling
Every CRM is a relational database wearing a UI
Salesforce, HubSpot, and Pipedrive look like very different products, but underneath they're the same idea: a set of objects (tables), each with properties (columns), connected by associations (relationships). Contacts, Companies, Deals, and Tickets are the default objects almost every CRM ships with — the modeling work is deciding how your business's actual data maps onto them, and when to extend beyond the defaults.
Get this model wrong early and every downstream system pays for it: marketing automation sends duplicate emails, sales reps work stale records, and reporting can't answer basic questions like "how many deals came from this account" because the relationship was never modeled correctly in the first place.
The core object hierarchy
Almost every B2B CRM data model reduces to four object types and the relationships between them:
- Contact — an individual person (an email address, effectively)
- Company / Account — the organization a contact belongs to
- Deal / Opportunity — a sales process tied to a contact and/or company, carrying a stage, amount, and close date
- Ticket / Case — a support interaction, usually tied to a contact
The relationships matter more than the objects themselves. A Contact typically belongs to one primary Company (many-to-one), but a Company has many Contacts (one-to-many). A Deal can be associated with multiple Contacts (a buying committee) and one primary Company — a many-to-many relationship that trips up a lot of first-time CRM admins who model Deals as belonging to a single Contact instead.
Model for how deals actually close, not how they start
Standard properties vs. custom properties
Every CRM ships with default properties (email, phone, lifecycle stage) on its standard objects. The modeling discipline is resisting the urge to create a custom property for everything. Two rules keep a CRM from becoming an unmaintainable pile of one-off fields:
1. Reuse before you create. Before adding a new custom field, check whether an existing property (or a picklist value on an existing property) already covers the need.
2. Every custom property needs an owner and a source of truth. A property with no defined system-of-record (is it set manually by sales, or synced from a form, or computed by a workflow?) will drift out of sync within a quarter.
Property sprawl is the #1 cause of CRM rot
Lifecycle stage: the property everything else depends on
Lifecycle stage (Subscriber → Lead → MQL → SQL → Opportunity → Customer) is the single property most marketing automation, lead routing, and reporting logic branches on — yet it's the property most commonly left ambiguous. The failure mode: marketing and sales define "MQL" differently, so a contact can silently regress or duplicate stages depending which team last touched the record.
The fix isn't a better field — it's a written definition, agreed by both marketing and sales, of the exact criteria (behavioral or explicit) that moves a contact from one stage to the next, with the stage-advancement logic implemented as an automated workflow rather than manual updates. Manual lifecycle stage changes are where data integrity dies.
Deduplication: designing against the failure, not fixing it after
Duplicate contacts happen because there's more than one path data enters the CRM — form fills, manual entry, imports, integrations — and none of them reliably check for an existing match before creating a new record. The two-part fix:
1. A defined matching key. Almost always email address for Contacts, and a normalized domain (not free text company name) for Companies. Freeform "Company Name" fields are the single biggest source of duplicate Company records, because "Acme Inc", "Acme, Inc.", and "ACME" are three different strings to a database.
2. Dedup logic at the point of entry, not a cleanup job after. Every integration and form handler should upsert (update-if-exists, else-create) against the matching key rather than blind-inserting.
// Pseudocode: upsert against email, the standard Contact matching key
async function upsertContact(email, properties) {
const existing = await crm.contacts.searchByEmail(email);
if (existing) {
return crm.contacts.update(existing.id, properties);
}
return crm.contacts.create({ email, ...properties });
}
// Never do a blind create — it's the #1 source of duplicate records
// await crm.contacts.create({ email, ...properties }); // wrongAssociations and rollup fields
Once objects are correctly associated (Contact → Company, Deal → Company), most CRMs let you compute rollup properties — a Company-level field that aggregates data from its associated Deals, like total closed revenue or count of open deals. Rollups are what make account-based reporting possible without a separate BI tool.
The common mistake is building this logic in a downstream reporting tool (a spreadsheet, a BI dashboard) instead of natively in the CRM. When rollups live outside the CRM, sales reps working inside the CRM never see them, and the numbers silently drift out of sync with the live data.
Model custom objects only when the standard four run out
What's next
A clean data model is what makes the rest of the MarTech stack trustworthy — automation workflows, lead routing, and revenue reporting are all only as good as the objects and properties underneath them. The next module covers building the automation logic that acts on this model.
Next: Make.com Patterns →
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.