Are Data Teams Cooked?

Follow into
Save into
Follow into

“Data teams are cooked.” I’ve recently spoken with several vendors who each claim their product will do away with data teams. Their product is some variation of multi-agent data engineer/analyst that fetches data from various sources, compiles analysis, and delivers this to non-technical users. We’ve seen this playbook before, and it seems that every so often, vendors try to kill off data teams and empower non-technical users. The story is alluring, especially for execs - use our product and “binge watch your business” (that billboard is from 10+ years ago). Especially in an age where AI is allowing non-quant people to become their own quant fund, its certainly appealing to see data teams as “cooked.”
A decade ago, the promise was self-service BI: stick Tableau, Looker, Domo, or Power BI in front of non-technical stakeholders, and watch data teams magically dissolve into strategic obsolescence. Business users would ask their own questions, build their own charts, and find their own insights. It never happened. Seven years of tracking by BARC and Eckerson showed enterprise-wide BI adoption hovered stubbornly at around 25%. Instead of eliminating data teams, self-service spawned competing dashboards, fractured metric definitions, and left the data team buried beneath an avalanche of reconcile-the-numbers triage. Turns out that data is very messy. Same as it ever was.
Today, the playbook is back, but the pitch has escalated to a frenzy, cuz AI. Now we have The Autonomous AI Data Team. The sales deck is intoxicating for an executive looking at headcount. Why pay a team of data engineers and analytics engineers when an agent can sit in Slack, ingest your warehouse, write SQL, chain workflow steps across systems, and answer the CEO directly?
It makes for an incredible vendor demo. But if you look at the hard data - and what is actually breaking inside enterprise architectures right now - the reality looks radically different.
Data Teams Are…Growing
AI agents are getting very capable (try out Fable or Astra for an idea). On the surface, it would make sense to claim that they can replace pesky humans. If AI agents were replacing data teams, hiring would be declining. The actual figures show the exact opposite.
- Teams are expanding, not shrinking: In my survey of 1,101 data practitioners earlier this year, 42% expect their data teams to grow in 2026. Another 44% expect headcount to hold steady. Just 7% expect shrinkage.
- Budgets are moving upstream: The dbt Labs 2026 State of Analytics Engineering report (363 respondents) mirrors this: 36% report increased team budgets versus 14% reporting cuts, while 57% report increased warehouse and compute spend.
- Senior demand is accelerating: Robert Half’s Demand for Skilled Talent report found 78% of tech leaders plan to expand permanent headcount in H2 2026, with data engineering specifically flagged as a core bottleneck command ($127K–$180K national ranges). Projections put total US data engineering postings near 260,000 this year, growing roughly 23% year-over-year.
Clearly, the data profession isn’t dying. However, the ground is shifting beneath our feet, and it’s disproportionately impacting noobs. Data from Stanford’s Digital Economy Lab (revised through June 2026) tracks employment for 22–25-year-olds in highly AI-exposed fields. That cohort sits 19% below where it would be if it had tracked less-exposed peers, an ~11% raw decline since late 2022 against ~10% growth for unexposed roles. Across tech broadly, postings for 0 to 1 year of experience linger roughly 38% below 2019 levels, even while 10+ year postings sit 23% above pre-pandemic baselines.
Teams aren’t replacing their senior architects with an LLM (at least for now). They are using foundation models to handle the mundane tasks that used to train apprentices, like writing baseline regex, converting boilerplate SQL dialects, and drafting mock schemas. This is the grunt work we all did in our early jobs, and it helped shape who we are today. The structural risk is far worse than the data team disappearing tomorrow. It’s far, far worse. We’re screwing the next generation in favor of subsidized tokens. Experienced talent moves on at some point, and we’re creating a catastrophic talent vacuum for five years from now. Or maybe AI gets good enough to eliminate all jobs and companies (and therefore the economy), but that’s a scenario for another rant.
Why the “CEO as Data Team” Falls Apart
The idea that an executive can replace a data team works in precisely one scenario, taking the form of an early-stage startup that never had a data team in the first place, operating on a clean, isolated schema where basic questions are better served by an agent. It’s a neat and believable scenario. I do this in my business, and it works wonderfully (I’m also a skilled data professional, so there’s that).
Enterprise is different. Beyond that narrow band mentioned above, enterprise reality hits three immovable force fields.
First, there’s the C word - context. There’s a lot of murky context within the enterprise. Frontier models can chain tasks, run sub-queries, and use browsers. What they cannot invent is institutional memory. An agent doesn’t know that the revenue_actuals table from 2021 double-counts European refunds because of an unmerged Stripe migration, or why Product and Finance define “active account” with conflicting lookback windows. That context lives in the heads of the people who have been refereeing those political battles for years. Like the senior architect above, these people aren’t getting younger and will move on, taking their institutional knowledge with them. Of course, this is why there’s a bajillion startups trying to capture institutional memory and tacit knowledge, such as Google buying Spirit Airlines’ chat data for $10 million. We’ll see how they do, especially as workers are painfully aware of the motivations of these companies to replace human labor, and will try to sabotage these efforts. Hatred of AI is the one unifying force in the world right now.
Second, there’s the panic about governance. Yes, the G word is back. In the dbt Labs report, 83% of leaders cited data trust as a primary priority, while a staggering 71% reported active anxiety over hallucinated or incorrect data reaching executive stakeholders. In a world of agents, the bottleneck ceases to be writing SQL. The crux becomes verifying correctness. If an agent returns a convincing, beautifully formatted hallucination to the CFO, who takes the fall?
Finally, there’s the S word - semantics. Every vendor has a semantic layer (I joke my goldendoodle has one too). Despite some who claim that semantic layers are unnecessary, an agent operating without guardrails isn’t ideal. At that point, the agent depends upon clearly named database table columns, sampling, and descriptive metadata. If you feed ambiguous warehouse tables directly into a model context window, you’ll get trash results. Hence, there’s a push in the data industry to establish semantic layer standards. We’ll see how it goes.
Agents are genuinely more capable than the BI tools of 2016. Chaining execution across disparate databases, APIs, and documents is a major technical leap over static dashboards. The punchline is that capability does not equal context. The hard parts remain.
My Prediction - Vendors Are Squeezed First
The bitter irony of the current wave is that the vendors pitching the demise of the data team are in far greater peril than the teams themselves. Every week, frontier models push deeper into native long-context reasoning, tool orchestration, and autonomous execution. The industry has already watched OpenAI and frontier updates gut thin AI wrappers. I don’t think thin-layer agent startups (many of whom are claiming to kill off data teams) survive the year.
If a product’s primary value proposition is “we put a prompt-to-SQL wrapper in front of your database and wire it to Slack,” the release of models like GPT-6 Astra just eliminated their entire moat overnight. And given the rate of change of models, don’t expect this to improve the prospects of these vendors.
The defensible layer isn’t the interface, the “proprietary harness,” or the prompt router. Instead, it’s the “boring” infrastructure such as open, version-controlled modeling and data, agent-native databases, deterministic access controls, and strict semantic governance. Boring stuff that was written off for dead (like data modeling) is making a comeback. The enterprise doesn’t need to pay a middleman rent just to ask a frontier model a question. It needs its internal data organized so that when any model asks a question, the underlying truth is deterministic.
Then there’s the harder challenge of market traction and distribution. In a hyper-crowded space like AI tools, how do you stand out? If you and several other vendors are trying to kill off data teams in very similar ways, why do I choose your product over the others? I’m approached daily by vendors looking to get “exposure,” and overwhelmingly, these vendors are all the same. I almost always respectfully decline to help. There’s very little differentiation in either their approach or story. It’s a sea of sameness - the same Claude designed slide decks, the same pitch of “get rid of your data team,” etc. I don’t expect them to survive.
And finally, if I’m an enterprise, I need to ask a simple question. Since I have a team that understands the data and domains in my business better than you, why not just have the team build a solution with the latest frontier models, which are the same models you’re using to build your product?
Stop the Nonsense Talk of Killing Things Off
Stop trying to kill things off. Enough with the nonsensical combative language and focus on growth rather than having a fixed zero-sum mindset. It’s a bad look that never ages well. Instead, talk about how you can help serve the data teams who can help your product get more traction within the enterprise.
Instead of killing off data teams, self-service expands the scope of opportunities. More people and agents creating and using data is a good thing, since there’s more opportunity to improve the data and the systems supporting it. I’m hopeful self-service AI/BI finally allows us to tackle the data chaos we’ve avoided in the past.
In this Freestyle Friday episode, I chat about the various craziness and trends of AI right now.
Freestyle Fridays and my other podcasts are available on Spotify, Apple Podcasts, and wherever you get your podcasts. Please support the show with a review. It means a lot.
Presented by Fivetran
Your pipeline works great for 3 data sources. But what happens when you hit 30? Or 300?
Building custom connectors is fine until you have to scale. Fivetran takes the headache out of the equation by handling the schema changes, API updates, and retries across hundreds of sources, delivering analytics-ready data in hours.
They handle the ingestion mechanics. You handle the actual data strategy, modeling, and AI.
Try Fivetran free for 14 days.
Where I’m At
My fall calendar is shaping up, and here’s an idea of what I’ll be doing. In most cases, I’m giving a talk. In some cases, I’m just hanging out.
From AI Hype to AI ROI Roundtable (hosted by Revefi)
If you’re a data leader in the Bay Area, there’s a very special invitation-only roundtable for enterprise data and AI executives. I’ll be there too. Register here.
Other events:
-
dbt Summit. September 15-18. Vegas. Register here.
-
Big Data London. September 22-24. Register here.
-
Motherduck BDL after party. London. September 23.
-
appliedAI. Munich. October 8.
-
Stanford. This fall, I’m joining Bruno Aziza as a guest speaker in his Stanford class on building products in the era of AI (October 12 – November 9).
The course was just announced and it’s already filling up. Stanford has opened it up online, so you can join from anywhere. Sessions are recorded too, so a time zone isn’t a dealbreaker.
Reserve your spot here soon because the course is almost full.
- Motherduck’s Data Outpost. November 4-5. San Francisco. Register here.
- Connected Data London. November 11-12. Register here.
- Forward Data Conference. November 16. Paris. Register here.
- MLOps World. November 17-18. Austin. Register here.
More to be announced very soon.
Cool Videos and Reads
Bringing Distributed Compute & AI to the Edge w/ David Aronchick (CEO of Expanso)
In this episode of the Joe Reis Show, I sit down with David Aronchick, co-founder of Expanso and one of the early pioneers behind Kubernetes, Kubeflow, and the CNCF.
We dive deep into the evolving data landscape and discuss why the architectural pendulum is swinging back from pure cloud centralization toward distributed edge compute. David breaks down the real operational nightmares of edge data collection, why treating your bronze tier as a "toxic waste dump" breaks downstream pipelines, and how applying schema and context upstream makes your data truly AI-ready. We also get into the critical need for data provenance, data bills of materials, and what agents actually need to communicate reliably.
Here are some things I read this week that you might enjoy
If this is true, the hyperscalers are toast
Right after I said not to write things off for dead, but…Local vs public models is the build vs. buy argument for today. Joachim Klement highlights research comparing small local models against cloud-based large language models across chat and reasoning tasks. The data suggests local models can match or outperform LLMs in over 80% of typical workloads at significantly lower energy and compute costs. As a result, Klement argues hyperscalers risk massive over-investment in data center infrastructure that the market will not actually require. (Klement on Investing)
The Internet Is Kind of a Predatory Cesspit Now
Stephen Diehl argues that the internet has evolved from an amateur public square into an industrial system designed to exploit human vulnerability for financial gain, where platform algorithms and participatory grift models convert users into both victims and distributors of scams. He also warns that language models will reduce the marginal cost of producing synthetic content to zero, further degrading the web. I agree, and hope we find an alternative to today’s slopfest. (Stephen Diehl)
Agents Don't Query Like Humans Do
Analyzing MotherDuck query history, Alex Monahan finds that AI agents generate 29 times more queries than human users, executing rapid-fire, low-row-count queries to inspect metadata and run small experiments. He argues that modern data platforms must evolve to handle bursty, high-concurrency workloads using fast elasticity, context layers, and local caching. This also fits what I’ve been ranting about for ages - agents aren’t humans, and their architectures will differ from what we’ve built. So it goes. (Motherduck blog)
Live from ICM 2026: What Is Math For in the Age of AI?
You think tech is alone in its existential crisis cuz AI? Math gives our industry a good running. In a live podcast recorded at the 2026 International Congress of Mathematicians, researchers Akshay Venkatesh, Ravi Vakil, and Alex Kontorovich discuss how AI system advances impact mathematics. The panelists argue that while machines can automate proof generation and solve specific technical problems, the true value of mathematics lies in human understanding and storytelling rather than raw problem-solving. They also highlight immediate risks, including administrative budget cuts framed around AI efficiencies and degraded student learning from shortcut usage, which I’ve seen firsthand. Interesting times in academia… (Quanta Magazine)
Note: These are articles I’ve read and enjoyed. I use AI to summarize my thoughts on the articles. I edit the summaries.
Find My Other Content Here
📺 YouTube - Interviews, tutorials, product reviews, rants, and more.
🎙️ Podcasts - Listen on Spotify or wherever you get your podcasts
📝 Practical Data Modeling - This is where I’m writing my upcoming book, Mixed Model Arts, mostly in public. Free and paid content.
If you’re interested in sponsoring my newsletter and podcast, Q3 and Q4 2026 are opening up. Space is very limited. Please fill out this form if you’re interested.
The Practical Data Community
The Practical Data Community is a place for candid, vendor-free conversations about all things tech, data, and AI. We host regular events such as book clubs, lunch-and-learns, Data Therapy, and more.
Thanks for reading! Subscribe for free to receive new posts and support my work.

