↵ select ↓ ↑ navigate esc close

Seroter's Daily Reading — #852 (August 24, 2026)

Seroter's Daily Reading ·

Listen: https://blossom.buildtall.systems/4cae14e31a71823f67444f37494a3d6eefd1333afd3935148dcf5857c3f70c4f.mp3

Source: Seroter's Original Post


Episode 852 for Monday, August 24th, 2026. Let's get into today's reading.

First up, a piece from Milan over at Tech World called 20 Lessons From 20+ Years in Tech. This one is a list of hard-won lessons from a couple decades in the industry, and the point is that we need to be reminded of these things and be more intentional about how we go about our professional lives. It's the kind of post that's easy to skim and nod along to, but the value is in actually sitting with each lesson and asking whether you're living it. Worth a slow read.

Next, Derek Comartin at Code Opinion has a post on event-driven architecture and coupling, with the blunt title You're Not as Decoupled as You Think. His argument is that a lot of teams break apart a monolith, drop in a message broker, and assume they've achieved decoupling, when all they've really done is start replicating data everywhere. The tell is the events themselves. If you're publishing things like ShipmentStatusChanged or ShipmentAddressUpdated, you've basically turned your internal data model into a public API. Every consumer now has to understand your internals, and if you change what a status means, you break them. His fix is to model events around business concepts and behaviors instead, things like ShipmentDispatched or ShipmentDelivered, and to treat integration events as a contract you version, just like an HTTP API. The key distinction he draws is between domain events, which are private and inside your boundary, and integration events, which are public. It's a good reminder that removing the synchronous call doesn't remove the coupling, it just moves it into the semantics of your messages.

Third, the Zed team published a piece called The Case for Software Craftsmanship in the Era of Vibes. The opening line is great: they keep hearing that developers will soon be replaced by autonomous agents, yet every single day they encounter bad software. The argument is that when the constraints on code production get lifted, the bar for quality should go up, not down. We should measure our contribution not in lines of code generated, but in reliable, well-designed systems that are easy to change and pleasant to use. There's a nice point that a gnarly codebase now hinders not just your own ability to work in it, but also the ability of AI tools to be effective in it. And they announce a new series called Agentic Engineering, combining human craftsmanship with AI tools. The throughline is that software isn't solved just because AI exists; system design, product thinking, and quality still matter enormously.

Fourth, The New Stack covers a new benchmark called SWE-Bench ProMax, focused specifically on large-scale refactoring. Most coding agent benchmarks skip this kind of work because it's hard to make deterministic and fast to run, but that leaves a big blind spot. The headline number is sobering: the best model only achieved a 41.2 percent resolve rate. The piece digs into why refactoring is so hard for these models. One expert makes the point that token proximity doesn't guarantee structural understanding, and that feeding a whole codebase into a big context window and asking for multi-file diffs is treating refactoring as a text-generation task, which it isn't. Another calls out time as a fundamental problem, race conditions, idempotency, atomicity, retry handling, the stuff that breaks when things happen in a different order. The benchmark spans 170 instances across seven languages, and the researchers went through multi-stage curation to fix the quality problems that plague other benchmarks. It's a solid, realistic test, and right now the models aren't doing great at it. Yet.

Fifth, a quick one from Google Cloud on where Antigravity looks for configurations. The gist is that if you use a handful of AI tools, you're probably copying config files around between locations, because the industry hasn't standardized on where we stash them. It's a small but real friction point, and it's the kind of thing that quietly eats time when you're juggling multiple agents and IDEs.

Sixth, a Research-Driven Engineering Leadership piece asking how engineering leaders should measure the ROI of AI. This is an all-important question, and the framing is sharp: most leaders can name their usage stats, seats activated, suggestions accepted, tokens consumed, but far fewer can quantify the value they got back. And the distance between those two numbers is creating friction as AI costs skyrocket. The piece walks through a methodology from Quotient that converts a productivity gain into dollars, then sizes it against spend. The key insight is that you have to separate usage metrics from value metrics. How much AI you're using is an input; what it returned is the actual question. There are also guardrails worth reporting alongside any ROI number, like cycle time, change failure rate, and developer experience, because weighted throughput can rise without a team actually improving. It's a genuinely useful framework for anyone who has to defend an AI budget.

Seventh, another Google Cloud post introducing the Antigravity Extension, described as Gemini Code Assist on steroids, in whatever IDE you love. The observation here is that for many people, the IDE is now mostly a read-only code viewer, and the real work happens in an agent. So it makes sense to put your coding agent in the IDE of your choice rather than forcing you into one specific environment. It's a sign of where the tooling is heading.

Eighth, a post called Three Clouds, Three Native Agents, where William looks at coordinating agents across the three major hyperscalers. It's a telling effort, because it gets at the reality that most enterprises aren't single-cloud, and if you're going to use native agents, you're going to have to make them work together across providers. Worth watching how that coordination actually plays out.

And ninth, a Google Developers Blog post on how to evaluate live and voice agents in ADK. This is a great one, because getting a live agent into production is about more than a good demo. It has to take the right actions across a spoken conversation, turn after turn, where timing and recovery matter as much as content. The post shows how to drive a live, voice-based agent with a simulated user that speaks its turns as audio, score the spoken replies, and do it all inside the same eval loop you already run for text agents. You can author conversation scenarios where the simulator improvises, or fixed conversations where you script every turn verbatim. And you can inspect the results in ADK Web, with transcripts and playable audio clips for each turn. The punchline is that your voice agent doesn't have to ship on vibes, it can ship measured.

That's the list for today. A few threads worth pulling. There's a clear tension between the abundance AI promises and the quality it actually delivers, whether that's the 41 percent refactoring score or the argument that software isn't solved just because AI exists. And there's a quieter theme about measurement, whether it's measuring the ROI of AI or measuring whether your voice agent actually works. The tools are getting more capable, but the discipline of knowing what good looks like is still on us. Thanks for listening, and I'll see you tomorrow.


  1. 20 Lessons From 20+ Years in Tech
  2. Event Driven Architecture and Coupling: You're Not as Decoupled as You Think
  3. The Case for Software Craftsmanship in the Era of Vibes
  4. Most coding agent benchmarks skip large-scale refactoring. Not this one
  5. Summary: Where does Antigravity look for configurations?
  6. How should engineering leaders measure the ROI of AI?
  7. Meet the Antigravity Extension: Gemini Code Assist on Steroids, in Whatever IDE You Love
  8. Three Clouds, Three Native Agents
  9. How to Evaluate Live & Voice Agents in ADK