↵ select ↓ ↑ navigate esc close

This Year’s Better Mousetrap: New Tech, Same Old Problems, and The Semantic Swamp

Joe Reis ·

A very nice post from Joe Perez (follow him) featuring the slide where I discuss what’s in this article. Namely, we keep trying to solve the problem, but never actually solve it.

In my keynote at Big Data London a few days ago, I discussed a pattern that has bothered me for years (squint at slide in the above pic). Decades have passed with the data industry promising to finally make data useful or deliver business value, whatever that phrase is supposed to mean anymore. Every fresh hype cycle arrives carrying a rebranded label wrapped around the same underlying pitch. Data Warehousing, Hadoop, Big Data, the cloud, data science, the Modern Data Stack, Data Mesh, GenAI, and last year’s favorite, AI agents.

This year, it’s “context,” (a VERY overloaded term right now) and practically every vendor booth and keynote conversation revolves around context. Suddenly everyone is selling a “context layer” or a “context graph” (which is hilarious, because calling something a context graph is like calling water wet since graphs are contextual by definition). And yet when I speak with leaders and practitioners, many people still complain about the same stuff we’ve heard for decades - the business isn’t data literate, data initiatives are stalled, adoption is slow, etc. What gives?

Don’t get me wrong. Current tools are genuinely impressive. They operate faster, offer far more expressive capabilities, and strip away constraints that used to consume our daily focus. Nobody misses writing MapReduce jobs by hand or fighting clunky ETL pipelines. Even so, beneath this steady stream of buzzwords, organizations remain bogged down by fuzzy meanings, unclear ownership, lack of trust, and sheer organizational friction.

Whenever I travel, that old phrase comes to mind: “Wherever you go, there you are.” People often travel to escape themselves or hope a new destination brings self-discovery, but it rarely works out because they brought along the same old baggage. Enterprise technology behaves identically. Switch platforms tomorrow, acquire the latest AI suite, and the exact same organizational dysfunction will be right there waiting to disrupt your plans.

Data Swamp, Meet the Semantic Swamp

Consider the tired old debate: “What is a customer?” To Sales, it’s an account with an executed deal. Finance views it as a distinct entity generating recognized revenue. Customer support looks at it as anyone reaching out with a ticket. Each definition makes total sense within its own domain, but things fall apart the moment these groups try collaborating on a critical business decision.

You can dump all three definitions into a central warehouse, have an agent build a semantic layer to run the metrics, and let an AI agent pull the output. But someone still needs to determine which definition fits a given scenario and who holds the authority to decide. Unifying data on a single platform does not automatically resolve fundamental human disagreements.

During the Data Lake 1.0 era, dumping everything into HDFS or S3 bucket created the infamous Data Swamp. Earlier this year on a podcast with Juan Sequeda, we pointed out that the exact same “just throw it in” mindset is setting us up for a Semantic Swamp. Vendors now market “ontology in a box,” claiming LLMs can scrape your Slack threads, emails, and internal docs to auto-generate business domain models. But parsing corporate chatter doesn’t equal shared understanding; it mostly just automates and scales existing confusion, politics, and flawed assumptions.

Coordination hurdles, internal politics, and genuine consensus require deliberate, tedious work. Organizations repeatedly kick these responsibilities down the road because marketing promises a simpler shortcut, but the fundamental challenges don’t vanish.

Escaping this loop requires setting aside silver-bullet thinking to do the unglamorous slog of cross-functional alignment. Real progress happens when teams clearly establish data ownership, settle on definitions, and treat alignment as an ongoing organizational effort rather than a software feature. Technology accelerates your goals, but only if you have built the human groundwork first.


In this Freestyle Friday episode, I chat about the various craziness and trends of data and AI right now.

Freestyle Fridays and my other podcasts are available on Spotify, Apple Podcasts, and wherever you get your podcasts. Please support the show with a review. It means a lot.


Presented by Fivetran

Everyone wants to talk about autonomous AI agents, but if your agents make decisions on stale data from broken pipelines, they’re not delivering value. They’re actually an operational liability.

And if you’re still hand-rolling commodity SaaS connectors (or burning expensive warehouse credits just to stage, cast, and prep raw tables) you’re stuck doing undifferentiated heavy lifting that didn’t even make sense in 2016, let alone 2026.

That’s why Fivetran’s latest announcements at dbt Summit hit the mark. They’re extending managed pipelines straight into modern open lakehouses:

  • Automated Movement to Open Lakehouses: Land fresh, governed Parquet directly into Apache Iceberg and Delta Lake without babysitting pipeline drift or file compaction.
  • Lake Compute (Private Beta): Unveiled at dbt Summit, Lake Compute lets you run transformation models on your ingested lake data right inside your dbt DAG, freeing your primary warehouse from raw staging overhead.
  • Agent-Ready Freshness: Ensure the semantic layers, feature stores, and agents downstream are executing on trustworthy, ground-truth data.

Stop building plumbing and paying warehouse markups on basic transforms. Let Fivetran handle the data movement and lake staging so you can focus on models, semantics, and real business impact.

Try Fivetran free for 14 days.


Where I’m At

My fall calendar is shaping up, and here’s an idea of what I’ll be doing. In most cases, I’m giving a talk. In some cases, I’m just hanging out.

  • appliedAI. Munich. October 8.
  • Hex Prompt. San Francisco. October 27.
  • Stanford. November 2.
  • Data Outpost. November 4-5. San Francisco. Register here. Get $150 off - Use PracData-150off at checkout.
  • Connected Data London. November 11-12. Register here.
  • Forward Data Conference. November 16. Paris. Register here.
  • MLOps World. November 17-18. Austin. Register here.

More to be announced very soon.


Cool Videos and Reads

DuckDB - The Past, Present, and Future, DuckLabs, etc. w/ Hannes Mühleisen (co-creator of DuckDB)

My brother from another mother, Hannes Mühleisen (co-creator of DuckDB), joins me to chat about the past, present, and future of DuckDB, Ducklabs (recently acquired by AWS), and more. Whenever Hannes and I get together, it’s fun. Enjoy.


Here are some things I read this week that you might enjoy

No articles this week cuz didn’t read anything cuz conferences…

Note: These are articles I’ve read and enjoyed. I use AI to summarize my thoughts on the articles. I edit the summaries.

Find My Other Content Here

📺 YouTube - Interviews, tutorials, product reviews, rants, and more.

🎙️ Podcasts - Listen on Spotify or wherever you get your podcasts

📝 Practical Data Modeling - This is where I’m writing my upcoming book, Mixed Model Arts, mostly in public. Free and paid content.

If you’re interested in sponsoring my newsletter and podcast, Q3 and Q4 2026 are opening up. Space is very limited. Please fill out this form if you’re interested.

The Practical Data Community

The Practical Data Community is a place for candid, vendor-free conversations about all things tech, data, and AI. We host regular events such as book clubs, lunch-and-learns, Data Therapy, and more.

🤖 Join on Discord

Thanks for reading! Subscribe for free to receive new posts and support my work.