↵ select ↓ ↑ navigate esc close

Seroter's Daily Reading — #862 (September 8, 2026)

Seroter's Daily Reading ·

Listen: https://blossom.buildtall.systems/797602fb23eafe6cf8995cd84b2b6c7d715b6c1a93727fe2bceecac20ca49a9f.mp3

Source: Seroter's Original Post


Episode 862, for September 8th, 2026. Welcome back.

Today's list has a clear through-line: what happens to software engineering when agents do more and more of the writing. There's a lot of strong thinking here about where the human still matters, plus a few interesting detours into Go's standard library, genomics, and the business of actually deploying AI. Let's dive in.

First up, from Telerik, a piece called Coding Agents vs. Workflows vs. Orchestration vs. Platforms: How to Architect the AI Dev Stack. The argument is that these four things are genuinely different jobs, and swapping one for another shows up in production as an incident no test suite predicted. The author walks through a scenario: two teams point two coding agents at the same authentication module, both ship green pull requests, and then production starts handing out 401 errors. Each agent only tested its own branch, so neither saw that merging the two would break a token-refresh path. The takeaway is that each layer answers a question the one below it can't answer: agents execute, workflows define what counts as done, orchestration owns state across parallel agents, and the platform governs who can run what. And the author flags something worth remembering: orchestration holds state you can't regenerate, which makes it the hardest layer to swap out later. Worth planning for before you're locked in.

Next, from O'Reilly Radar, Inside a Software Factory. Seroter noted he's not sure the software factory as currently described will survive, but this post does a fine job outlining where humans fit. The author built his own factory called Squid and maps it to eight stages: triage, brainstorming, planning, implementing, review, review-CI, release, and monitoring. The key insight is that humans are indispensable at brainstorming and planning, and they return for the final check, while agents own everything in between. He also makes a great point about planning quality: a strong plan lets cheap models execute cheaply, while a weak plan forces them to retry until the extra tokens erase any price advantage. And his warning about overbuilding is refreshingly honest: his first attempt was a monolith he couldn't debug, halt, or redirect mid-run.

That connects directly to the third piece, from Milan's newsletter, What is the future of Software Engineering when nobody needs to write code? Seroter says if you read just one article that sums up the current discussion, make it this one. The core idea is that code was never the product, and now that code is cheap to produce, the hard part moves to judgment: deciding what to build, writing the specification, and verifying that the generated code actually does the right thing. The author lays out nine points, and a few stand out. Specification becomes the real programming, because agents can't translate vague requests. Verification becomes harder than generation, which means you need automated gates that can scale. Architecture becomes a control mechanism that defines boundaries for the agent. And there's a junior developer problem: if agents do all the implementation, how does anyone build the judgment that comes from failure and debugging? It's a thoughtful, wide-ranging read.

Seroter also flagged a genuinely fun engineering detour: a dev.to post called Replaced 15 Go Packages With Nothing But the Standard Library. The author built a 2FA authenticator and encrypted vault named stdotp using only Go 1.27's standard library: no go get, no vendor folder, no network calls. He reimplemented TOTP and HOTP straight from the RFCs, and the punchline is that the algorithm is about fifteen lines; the hard part is trusting your own implementation, which means writing test vectors by hand. There are nice details too, like binding the AES-GCM header as additional authenticated data so tampering with the iteration count causes decryption to fail. As Seroter points out, external packages absolutely have their place, but when you don't need them, there's real advantage in not depending on them.

Then a question that would have sounded silly six months ago: What happens when AI burns down the backlog? from Dave Porter. His answer to whether AI truly burns down the backlog is: kind of. AI ships PR-ready code orders of magnitude faster than humans, and that puts pressure on testing and CI systems that can't keep up. And it compounds: AI is a fast car that needs good roads, and if your pipeline is a monolith with long build times, it'll compound a wreck. But the more interesting question is whether the backlog itself runs out. He used to say no, and now he's not sure, because AI is ultimately a business transformation problem, not a technical one. His advice: invest at the front of the pipeline, in figuring out what's actually worth building.

Then a quick but weighty note from Google DeepMind: AlphaGenome Atlas: A Predictive Map of Every Possible DNA Letter Change in the Human Genome. Seroter's comment is exactly right: there aren't many frontier labs doing work like this consistently. The acknowledgments alone read like a who's who of medical research institutions, and the ambition of mapping the effect of essentially every possible single-letter change is the kind of thing that only happens with frontier-scale compute.

From the getdx newsletter, a fascinating data point: The quality paradox of AI-generated code. Across more than five hundred companies, code maintainability improved almost four percent in Q2, while change confidence fell six percent. Those two have historically moved together, because code you understand is code you feel safe changing. The author uses an older framework called TRUCE to explain why they've split, and the core idea is that quality was never a scalar. Locally good code doesn't guarantee a healthy system; an AI can write code that's perfectly readable in isolation while duplicating something that already exists or adding an unnecessary dependency. The essay ends on a striking note: we can now measure the machinery of software production with extraordinary precision while remaining almost blind to whether the software actually serves the people using it.

Speaking of deployment, TechCrunch covered Google Cloud Races to Catch Up in the AI Deployment Wars with Accenture Deal, a joint unit with Accenture called the Accenture Gemini Enterprise Business Group that will train up to a thousand forward-deployed engineers to build AI applications on the Gemini Enterprise platform. It's part of a broader bet: OpenAI, Anthropic, Microsoft, and Amazon have all launched similar programs, betting that implementing AI is its own trillion-dollar business. The stakes are high. Google accounts for only about six percent of enterprise AI spending according to Ramp, versus Anthropic's forty-plus percent, even as hyperscalers pile up hundreds of billions of dollars in commitments.

On a related note, The New Stack covered research from a company called Armature on how coding agents pick their tools. They watched seventeen thousand tool-choice sessions and found, as one founder put it, that twenty years of brand building simply froze in time. Reputation follows a tool into the model weights, but it operates under different gravity now. A few findings: repository context matters enormously, so the winning email provider differed by language. Resend won on TypeScript, SendGrid on Python. And the agents themselves disagree: Cursor leans on web search, Codex searches in almost every session, and Claude Code relies mostly on its priors. Maybe the wildest stat is that PayPal was mentioned 139 times and never picked. As Seroter notes, a model's bias toward web search over what's frozen into its weights has a big impact.

Then a useful benchmarking piece from Google Cloud: Not All LLM Workloads Are Equal: Benchmarking TPU Performance on Classification vs. Generation. Comparing Gemma 3 twelve-B and twenty-seven-B on TPU v6e, they found that for generation-heavy tasks the twenty-seven-B model hits a hard wall past sixty-four concurrent users, while the twelve-B keeps scaling. But for classification, where there's lots of input and a tiny output, parameter size barely matters and both models scale similarly. The practical advice: you can deploy bigger models for classification and summarization without a throughput penalty, but if your workload is decode-heavy and high concurrency, downsize or cap concurrent requests.

Jeff Gothelf tackles How to Write a Definition of Done for an AI Feature, and his warning is sharp: don't mistake evals for a requirements spec. Evals tell you the system behaves the way you specified, but not whether anyone wants it. He points to Albertsons, whose AI shopping assistant gets shoppers to spend more, partly because they stop forgetting items. An eval wouldn't catch that, because forgetting an item is a quality of the shopper, not the system. His fix is to attach a specific user outcome to every eval, so "done" means both passing the tests and moving the behavior you actually care about.

And finally, from InfoWorld, a piece called What It Took to Triple Our Software Engineering Output in 18 Months. The author's advice: set a public, aggressive goal and design for the velocity you're about to create. His underestimated lesson was that the bottleneck moves downstream. Once engineering sped up, the constraint moved to go to market, and now they treat go to market as its own automated phase. His bigger point is that the tools matter less than the organization. Everyone can buy the same tools; what actually changed was the system around them.

That's the list for episode 862. The thread that ties it all together: we've made producing code dramatically faster, and that success keeps pushing the hard problems elsewhere, into judgment, specification, verification, and the work of going to market. The engineers and companies that figure out where the human still belongs are the ones positioned to win.

Thanks for reading, and I'll catch you next time.


Articles covered in this episode:

  1. Coding Agents vs. Workflows vs. Orchestration vs. Platforms: How to Architect the AI Dev Stack
  2. Inside a Software Factory
  3. What is the future of Software Engineering when nobody needs to write code?
  4. Replaced 15 Go Packages With Nothing But the Standard Library
  5. What happens when AI burns down the backlog?
  6. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
  7. The quality paradox of AI-generated code
  8. Google Cloud races to catch up in the AI deployment wars with Accenture deal
  9. “Twenty years of brand building simply froze in time”: How coding agents select their tools of choice
  10. Not All LLM Workloads Are Equal: Benchmarking TPU Performance on Classification vs. Generation
  11. How to write a definition of done for an AI feature
  12. What it took to triple our software engineering output in 18 months