Seroter's Daily Reading — #863 (September 9, 2026)

Follow into
Save into
Follow into

Source: Seroter's Original Post
Welcome to the Daily Reading list for Wednesday, September ninth, twenty twenty-six. This is post number eight hundred sixty-three. And Seroter's in a good mood today — he writes that it was a good day, and that tomorrow he's jetting up to Sunnyvale for some fun meetings at Cloud HQ. His note is a nice reminder: it's easy to get caught up in everything changing all around us, but he reminds himself to enjoy the present moment, learn new things, and not worry about tomorrow. Good advice.
Now, today's list has a strong thread running through it — agents doing real, production-grade work. And refreshingly, people being honest about both the wins and the places where agents still struggle. Let's dig in.
First up is a genuinely impressive story from the monitoring company Checkly, in a post called We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go: Zero Incidents. They took a background worker called the Results Daemon — a Node.js service processing roughly ninety-two million messages a day — and rewrote it in Go, entirely using AI agents, with zero incidents. The results were striking: a seventy percent reduction in running pods, lighter database load, and Go's stronger type system proved a better fit for agents than JavaScript. But the thing that really made this work wasn't the model — it was the test harness they built before the agent wrote a single line. They treated the component as a black box, generated inputs from their production data lake, and recorded every expected output in golden files, asserting byte-for-byte parity. They used real infrastructure like actual PostgreSQL containers rather than emulators. And they learned some hard lessons along the way: on first deploy, retries started failing because their local queue topology didn't match production, where routing depended on priority and hosting type across eighteen possible queues. The fix was a rule worth stealing: align your harness with production as closely as possible, define every boundary down to its lowest unit, and write down every assumption rather than inferring it. It's a really thoughtful blueprint for anyone considering an agentic rewrite of a critical service.
Next, a short but provocative piece from Jacob Gold, whose post is titled I Trust My Coding Agents with Production Secrets Now. He hands his agents access to his SSH keys, API tokens, even a Gmail app password — all stored in a vault so the agents generally use them without ever seeing the values. His framing is useful: security is always a productivity-risk tradeoff. He gives agents production access the same way he'd give access to an inexperienced teammate — knowing a new teammate could make mistakes, but they still need the keys to do their job. He also points out that frontier models like Astra and Fable are simply much harder to trick with prompt injection than older models were. But notably, he still runs those agents in isolated Docker containers, and he still does least privilege where reasonable. It's a balanced, realistic take rather than a cavalier one.
There were also a couple of entries Seroter flagged that I couldn't pull full text on this morning. One is Karl's piece, Five Things You Shouldn't Vibe Code, and What to Use Instead — with the headline advice being, don't build your own authentication system, it's a solved problem.
And the other is a hands-on guide to safely running untrusted code in Cloud Run sandboxes. Both are worth a look, even off the one-liners.
Sticking with agents, Meta this week introduced Muse, which they're calling the world's first personal AI agent built for everyone. What's interesting here isn't just the assistant — it's the infrastructure around it. Muse runs on something called Muse Secure VM, a dedicated virtual machine that houses both the agent and a person's data. A separate "Sentinel" agent sits on that same machine, kept apart at the system level, and nothing Muse does reaches the internet unless Sentinel approves it. Muse never sees passwords or payment methods — credentials go into secure storage and the agent uses them without reading them. It checks with the person before sensitive actions like sending an email or making a purchase. And later this year, Meta plans a Confidential VM where the whole thing is encrypted with a key only the user holds. It's a signal that the next wave of agents is going to be defined as much by trust and safety as by raw capability.
Then there's a short piece from the Fowler-adjacent camp asking a pointed question: Do You Even Need a Presentation? It argues something we've probably all felt but rarely said out loud: we conflate the presentation with the slide deck. The author's definition is worth repeating — a presentation is an act of storytelling in which the presenter orchestrates narrative, timing, and emotion. The slides are not the presentation. His thesis is that most corporate communication is just conveying information, and for that, a well-structured document is almost always better than slides — documents have linear structure, are easier to scan, and reduce the chance that even an LLM summarizing them will hallucinate. Live presentations, he argues, should be reserved for moments that genuinely merit synchronous interaction. It's a refreshing push back against deck-for-everything culture.
Now a quick one that Seroter says he thinks he loves. It's a post from Package Main titled Gitignore Everything by Default. We've all been there — committing a .DS_Store file, a node_modules directory, or an environment variable, then doing the embarrassing cleanup. The idea here is to flip the approach: instead of allow-anything and selectively ignore, start with an asterisk that ignores everything, then use negations to whitelist exactly the files you want — your Go files, your go.mod, your README. Only what you explicitly allow gets tracked. It's a clean, disciplined alternative worth considering for repos drowning in local junk.
On the agent tooling side, Google Cloud published their Data Agent Kit, for what they call agentic analytics. The scenario they lead with is relatable: your director asks why average order value dropped seven percent in January while revenue stayed flat, and suddenly you're bouncing between a warehouse, a production Postgres instance, and raw JSON files in object storage. The Data Agent Kit bundles a set of MCP servers and agent skills — markdown files that teach an agent how your specific stack works — so the agent can run the queries and read the results on your behalf, rather than you copy-pasting SQL between consoles all afternoon.
And I want to close on the most contrarian piece of the day, from Armin Ronacher, who asks bluntly: Astra for Coding: Why Are We Doing This Again? He ran what he calls a "slop factory" over the weekend — letting the model manage its own context, spin off subagents, and decide its own workflow on a single prompt. Thirty-five hours and roughly four billion tokens later, it delivered almost nothing of value, produced seventy-five thousand lines of code, and never stopped. His observation is sharp: the model is heavily rewarded for token efficiency and long-horizon task completion, but rarely punished for writing bad code. So it code-golfs — reaching for obscure one-liners, unreadable string-splicing, and bizarre nested tool calls — optimizations that look efficient but are a nightmare for a human to maintain. His broader point is sobering: we easily measure token efficiency and task completion, but we don't have simple metrics for what makes code understandable, and the less humans look at the output, the less any of that matters.
Also worth a mention, Seroter flagged one more from Abi at Google: a real-time pickleball agent built on Spanner Omni, using on-device GraphRAG. I couldn't pull the full text either, but the gist is an educational look at an edge-style deployment — running a downloadable version of Cloud Spanner on-device. It's a nice concrete example of bringing a serious database down to the edge.
So there's a nice tension in today's list: on one hand, Checkly's harness-driven rewrite shows agents can ship real production code safely when you define what "correct" means first. On the other, Armin's slop factory shows exactly what happens when you don't. The difference isn't the model — it's the guardrails.
That's it for post eight sixty-three. Thanks for reading along, and we'll catch the next one tomorrow.
- We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents
- I trust my coding agents with production secrets now
- 5 things you shouldn't vibe code (and what to use instead)
- Introducing Muse: The World's First Personal AI Agent Built for Everyone
- Do you even need a presentation?
- Safely Running Untrusted Code: A Hands-On Guide to Google Cloud Run Sandboxes
- .gitignore everything by default
- Agentic analytics with the Data Agent Kit
- Astra for Coding: Why Are We Doing This Again?
- Building a Real-Time Pickleball Agent with Spanner Omni: On-Device GraphRAG