The Five Camps of Data Modeling (and Which One You’re Stuck In)

Follow into
Save into
Follow into
This article is the first in a series of free and paid adaptations from Mixed Model Arts. Here’s a look at the Five Camps of data modeling, adapted from Chapter 1 (The Dawn of Mixed Model Arts).
A few years back, an app team stored their entire product catalog in a massive JSON column in Postgres. Nested documents, pricing, reviews, arrays. For the app, it was great: fast reads, zero joins, fast deploys.
Practical Data Modeling is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Then analytics needed a product revenue dashboard. Categories were stored as arrays in some rows, nested objects in others, and plain strings elsewhere. Review counts in the JSON blob didn’t match the reviews table. Data engineering wasted three months hacking together a brittle pipeline that broke on every app release. The dashboard numbers were garbage, and leadership lost trust.
At the same time, the ML team built its own pipeline off the same column for recommendations, inventing its own taxonomy for categories. Recommendations contradicted the dashboard. The launch slipped four months, and the VP of Product had to explain to the CEO why their strategic AI initiative was sidelined by a database hack from two years earlier.
Every team was competent. Nobody looked across the whole architecture. Nobody asked how the app schema would break analytics, what ML needed, or what core terms actually meant across the company.
I see this exact scenario constantly. Teams model in isolation, talking past each other without shared vocabulary or coherent definitions.
Data Modeling Grew Up in Five Separate Camps
Different communities developed techniques for specific problems, and each came to believe its paradigm mattered most. You probably grew up in one. All five produced vital work, and all five have massive blind spots.
The Relational Camp dates back to Codd’s 1970 paper. It lives in normal forms, referential integrity, and eliminating redundancy. If you studied CS or database theory, you learned this first. This camp takes modeling more seriously than anyone else, but its blind spot hits when you pivot from transactional integrity to analytical queries. A schema requiring seven joins for a basic query is pure on paper and painful in production. For JSON events, raw text, or vector embeddings, relational modeling offers zero help.
The Analytics Camp took off in the 1990s to solve analytical queries against OLTP systems. Inmon’s enterprise data warehouse, Kimball’s dimensional modeling, and Linstedt’s Data Vault came from this wave. This camp proved that analytical workloads need different schemas than transactional apps. The blind spot? Tunnel vision. Ask traditional dimensional modelers to design real-time event streams, graph traversals, or feature tables, and they struggle. (And yes, people still waste energy arguing Inmon vs. Kimball.)
The Application Camp exploded during the web-scale 2000s and 2010s. Developers needed speed and flexibility, so they adopted NoSQL—MongoDB, Redis, Cassandra, Kafka. They model strictly around app access patterns. A document model might stuff an order, line items, and payment info into one JSON object. That makes normalization purists twitch, but it fetches everything in a single read. The blind spot is everyone downstream. Schemas tuned for one app’s reads create chaos when analytics needs millions of records or ML needs clean feature extraction.
The ML/AI Camp rarely calls what it does “data modeling,” but it lives in feature tables, embeddings, and vector spaces. Because the consumer is an algorithm, it organizes data for training and inference. The blind spot: many ML engineers have never touched relational design or dimensional grain. An ML engineer who doesn’t understand point-in-time correctness causes data leakage between training and test sets. Someone ignoring entity resolution trains models on duplicated customer records.
The Knowledge Camp comes from library science, philosophy, and enterprise architecture. Its toolkit includes ontologies, taxonomies, and knowledge graphs, and it asks core questions: What is a “Customer”? Is an “Order” an entity, a state, or an event? Historically, it operated on the periphery, detached from messy transaction logs and data lakes. But LLMs and AI agents have dragged this camp to center stage. An agent querying your stack must know what your terms actually mean.
Isolation Is the Core Problem
Practitioners rarely look beyond their niche. Data engineers who’ve never designed a normalized table. App devs who’ve never heard of slowly changing dimensions. ML engineers who’ve never modeled a core business domain. Ontologists who’ve never shipped a production pipeline.
Specialization worked when systems were isolated. Today, an e-commerce stack runs Postgres for orders, Kafka for events, star schemas in the warehouse, embeddings in vector stores, and a semantic layer defining “customer” and “revenue.” It exists simultaneously. It all gets modeled—intentionally or accidentally.
Three Waves Got Us Here
Wave 1: Operations meet analytics (1990s–2000s). Data modeling meant relational design for applications. Then business questions started hitting production OLTP databases, threatening uptime. Inmon published Building the Data Warehouse (1992), and Kimball introduced The Data Warehouse Toolkit (1996). Practitioners had to map normalized transactional sources into analytical targets.
Wave 2: Big data and lakes (2000s–2010s). Hadoop, NoSQL, JSON, and event streams brought schema-on-read: “Dump it in the lake; we’ll figure it out later.” Unmodeled lakes turned into data swamps. Successful engineers had to connect app documents, event streams, and warehouse models.
Wave 3: ML and AI (2010s–present). A customer is now an OLTP record, a warehouse dimension, a feature vector, and a graph node. These representations must agree. An AI agent facing ambiguous columns, missing business rules, or broken grain generates confident garbage at machine speed.
New waves don’t replace old ones. They stack.
Camps, Forms, and Layers
To navigate this landscape, keep three concepts clear: Camps are the practitioners and their traditions. Forms are the shapes of the data. Layers are where the data lives in your architecture.
Notice the asymmetry: structured data spans two layers, which is a reason the Relational and Analytics camps fought for decades. Unstructured data historically had no formal modeling discipline at all.
Take a single real-world transaction: I buy a red kettlebell to support the University of Utah (go Utes!). Order 808.
- Transactional: Normalized rows in Orders and OrderLines ensuring inventory integrity.
- Event: Clickstream events—searched “kettlebell,” added to cart, checked out—streamed through Kafka.
- Analytical: A fact table record joined to a surrogate CustomerKey to track dimension changes over time.
- Unstructured: A product review with a chipped handle image evaluated by sentiment analysis.
- ML/AI: Feature vectors for recommendations placing me near strength gear and far from yoga mats.
- Knowledge: An explicit business rule requiring a “Customer” to have at least one completed purchase, preventing AI agents from counting newsletter subscribers as active buyers.
One order, six representations. The CustomerID in OLTP must resolve to the CustomerKey in the warehouse, the user ID in Kafka, and the feature vector. When those keys break, your architecture falls apart.
Go Deep in One, Get Competent in the Rest
No. You don’t need master-level expertise in every discipline. You need cross-disciplinary literacy.
Keep your primary specialty, but build fluency in the others. A relational modeler who understands dimensional modeling builds source schemas that load cleanly into data warehouses. A data engineer who understands feature stores builds pipelines ML teams can actually use. A practitioner who understands semantics documents business logic so schemas explain themselves to both humans and LLMs.
Stop building models that solve problems for your team while creating technical debt for three others.
Which Camp Are You In?
- You immediately draw entities and foreign keys. Relational. You likely miss how data gets queried for analytics and how to handle non-tabular forms.
- You think in facts, dimensions, and grain, and hold strong opinions on Kimball vs. Inmon. Analytics. You likely miss real-time streams, graph relationships, and ML features.
- You model strictly for app reads, and “just add a field to the JSON blob” is standard practice. Application. You’re probably creating chaos for downstream teams.
- You live in feature stores, vectors, and embeddings, treating “data modeling” as someone else’s job. ML/AI. You likely miss grain, point-in-time correctness, and entity resolution—leading to data leakage and poor model performance.
- You start every meeting defining business terms and building ontologies. Knowledge. You likely miss the messiness of physical production data.
Take a recent design choice your team made—a schema, pipeline, or feature set—and identify which camp generated it. Then evaluate how the other four camps would approach it, and where cross-camp thinking would have saved you headache down the road.
Every team in that JSON-column story did its job, and the project still failed. Mastery of one camp isn’t enough when your data has multiple consumers. And today, your data always has more than one consumer.
Adapted from Chapter 1 of Mixed Model Arts (Book 1: Foundations). The book breaks down each camp, traces Order 808 across all six architectural layers, and establishes the core building blocks: entities, grain, time, semantics, and identity.
An annual subscription to this Substack gets you access to the Mixed Model Arts ebook. You can also find the book on Amazon.com (coming very soon).
Practical Data Modeling is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
