↵ select ↓ ↑ navigate esc close

Seroter's Daily Reading — #871 (September 21, 2026)

Seroter's Daily Reading ·

Listen: https://blossom.buildtall.systems/612118ed05127d1a8154c3f1dad4fd8a0d6d8c8e6ca65e16650ccb573d7a2ba6.mp3

Source: Seroter's Original Post


Welcome to Seroter's Daily Reading, episode 871, for Monday, September 21st, 2026.

Let's start with an ambitious one this week. Thorsten Ball, best known for writing a pair of books on interpreters and compilers, published a piece he calls What I Believe About the Future of Software Development. It originally went out on X and blew up, so he put it on his blog to plant a flag. And it is a long list of genuinely radical predictions. He argues that code review is already dead, and unit tests might die too, because the models are getting good enough that the training wheels aren't necessary. He says the craft of writing code will disappear, the way making shoes by hand mostly has, but the craft of building software, understanding business problems and how to ship and get feedback, will matter more than ever. He thinks the terminal is dead, that tokens are the new computing paradigm, and that the classic triad of product manager, designer, and engineer is going to dissolve. There's a lot here I can't easily find fault with either, which is exactly what makes it uncomfortable. It's the kind of thing worth sitting with, even if you disagree.

Speaking of new paradigms, the model everyone was talking about this weekend is called Jev. It comes from a company called TypeSafe, founded by a former OpenAI engineer after two years in stealth, and the line everyone latched onto is that Jev is an LLM without the LL. It's the first model in what they're calling the System One series, and it's not designed to reason or write code. Instead it's built purely to make fast, structured decisions that software can consume directly. Think of it as a frontier-intelligence function call: unstructured information goes in, typed, probabilistic decisions come out, with the structure defined ahead of time by the developer. Because it's not carrying around the whole corpus of an LLM, it's cheap, something like four cents per million input tokens and free output tokens, which the company says are too cheap to meter. It won't replace your language model, but for the narrow job of an agent picking between predefined options, it's an interesting bet.

Now a cautionary thread. James Governor at RedMonk has a piece riffing on work from Per Buer and Konstantin Ryabitsev about The Agents Are Coming for the Web, and the web isn't ready. Ryabitsev runs the numbers for git.kernel.org, and they're wild. The Linux repo is about one and a half million commits with over nine hundred forks, and because every fork shares the same objects, it's efficient for humans, but a scraper sees several billion valid URLs to chase. Bots started out polite, then got sneaky, rotating through residential and mobile IPs making a few requests and vanishing. The frustrating part is that agents don't respect efficiency or robots.txt, and this isn't just crawlers. Every personal agent we all fire off asynchronously adds to the load. Cloudflare, Fastly, and GitHub are all trying to work the problem, and Governor's point is that agent traffic won't grow by a couple hundred percent but by orders of magnitude. Anything public facing is about to face serious scaling demands.

From there, let's talk business models. Ben at the Product-Led Geek wrote Popular Isn't a Business Model, and the core message is in Seroter's framing: open source isn't a business model, it's a distribution strategy. You still have to figure out what's worth paying for once someone is using the software. Ben lays out six proven paths: open core, where the individual adopts and the company pays to govern; managed cloud, where you're on call and charging for it; support and services, where you sell expertise and an SLA; commercial licensing, where you sell different legal terms; sponsorship, where someone funds work everyone benefits from; and companion products, where you sell a second thing to people who already love the first. He's lived this at CloudBees, watching upstream Jenkins and Kubernetes steadily erode the value they charged for, and his warning is that the line between free and paid is a one-way door. Move something that used to be free into the paid tier and you get something like the MinIO backlash, which ended in a fork. The takeaway is that open source gets you adopted, but not paid. That part you have to design.

Let's get concrete with an engineering tip. The team at CloudX published a post about scaling their Go CI by replacing GitHub's actions/setup-go. The default action builds its cache key from the operating system, architecture, Go version, and a hash of your go.mod files, which sounds reasonable until you realize that in an active codebase, almost none of your changes touch any of those things. So the cache goes stale, and every run redoes work from scratch. Worse, parallel jobs like lint and test race to write the same cache key, and whoever finishes first poisons it for the other. Their fix, which they've open sourced as cloudx-io/setup-go, includes the job identity and the run ID in the cache key so jobs write fresh caches every single time. The result was a sixty-nine percent cut in test job runtimes, and they backtested across four thousand commits to show that eighty-six percent of the default action's test work was unnecessary. It's a great reminder that a lot of slowdowns aren't your code, they're your tooling.

Switching to data, O'Reilly Radar published a useful explainer called Navigating the Modern Data Lexicon. It's aimed at the people who actually deploy these systems but keep running into jargon like semantic layer, ontology, and knowledge graph without a clear sense of how they differ. The article's framing is the most useful part: the single most important diagnostic is whether you're asking for a deterministic answer or a probabilistic one. A deterministic output is the same number every time, with a traceable calculation. A probabilistic output is a prediction that varies each run. The failure mode is asking a probabilistic system for a deterministic answer and not realizing it. From there, it walks through how these terms stack rather than compete: the ontology describes what exists, the knowledge graph holds the instances, the semantic layer defines the measures, context is how any of it reaches a model, and observability is how you find out when it stops working. If you've been nodding along while people throw these words around, this one's worth a read.

On the infrastructure side, Google announced a preview of cross-cloud caching for its borderless Lakehouse. The whole pitch of the Lakehouse is querying data in place across S3, Azure, GCS, and SaaS platforms without moving it into brittle ETL pipelines. But that cross-cloud data movement isn't free, so the new idea is to cache frequently accessed data locally. The clever part is the granularity: instead of pulling whole multi-gigabyte files, BigQuery caches only the specific column chunks and dictionary pages a query actually projects, at the sub-file block level. Combined with Iceberg columnar compression, they say you often only need to transfer under five percent of the data you process across clouds. They've also added freshness checks so you don't get stale reads, and everything is encrypted at rest by default. It's a sign that the multi-cloud analytics story is starting to get the cost mechanics figured out.

Here's a fun one about the agentic consumer. TechCrunch reported that Meta's AI agent has been blocked from using Amazon.com, while Shopify, in contrast, responded by building an integration into that same personal agent. Seroter's spin is the interesting part: the businesses that embrace the agentic consumer, that treat an AI shopper the way they'd treat a human one, are going to quickly surpass the ones that don't. When your customers increasingly arrive as agents rather than people, blocking them is a bet against your own checkout flow.

Back to infrastructure efficiency. Google also wrote about GKE Pod snapshots for scaling AI workloads, and the headline number is that in their examples inference startup time is down nearly ninety percent. The mechanism is a declarative policy using new Pod snapshot CRDs that captures a pod's compute state so you can resume it almost instantly. For agentic workflows in particular, this solves two problems at once: you can snapshot a sandbox once and spin up new ones quickly, and you can suspend idle sandboxes so you're not paying for GPUs sitting there unused. A customer called Retake replaced a complicated custom caching layer and cut their startup from a minute down to about eight seconds, which let them spin up H100s on demand and shut them down immediately.

And to round out the economics, Stripe published a piece called SaaS Platforms Are Surging Despite the SaaSpocalypse. A few months back the markets panicked that agentic AI would commoditize software, and software companies shed a trillion dollars in market cap in a month. But Stripe's payment data tells a different story on the ground. They added more new platforms in the last three months than they did in the final six months of 2025, and new platform businesses are up over one hundred eighty percent year over year. Their argument is that as software gets easier to build, the winners are the ones with deep industry knowledge and workflows that a customer can't turn off. The line they quote is a good one: a high-signal indicator of defensibility is what breaks the day the customer turns it off. If claims don't pay or cars don't sell, the product is hard to dislodge. So SaaS isn't dead, it's just differentiating harder.

Finally, a nice counterpoint to the Thorsten Ball piece. Tech World with Milan interviewed Sebastian Raschka, the author of Build a Large Language Model from Scratch, on a simple question: do engineers still need to understand how LLMs work? His answer is yes, to a point. He's not saying everyone should train their own model, but understanding what's happening under the hood helps you write better prompts, be more critical of outputs, and keep up as the field evolves. He uses Claude's watermarking as an example; if you know how token sampling works at inference time, the whole watermarking mechanism becomes obvious. His practical advice for a backend engineer is to rebuild something you've already built, this time with an agent, and then let the agent extend it. And his one wrong prediction was that agents matured much faster than he expected. There's a nice tension between these two ends of the episode: Ball says the craft of writing code is disappearing, while Raschka argues that understanding the machine is still a worthwhile investment in your future self. Both are probably right.

That's it for episode 871. Thanks for reading along, and I'll see you on the next one.

  1. What I Believe About the Future of Software Development
  2. Runtime: Jev is an LLM without the LL
  3. The agents are coming for the web and the web isn't ready
  4. Popular isn't a business model
  5. Scaling Golang CI by Replacing actions/setup-go
  6. Navigating the Modern Data Lexicon: A Working Vocabulary for the Semantic Era
  7. Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching
  8. Meta's AI agent has been blocked from using Amazon.com
  9. Scale your AI workloads faster and more efficiently with GKE Pod snapshots
  10. SaaS platforms are surging despite the SaaSpocalypse
  11. Do engineers still need to understand how LLMs work?