Seroter's Daily Reading — #865 (September 11, 2026)

Follow into
Save into
Follow into

Source: Seroter's Original Post
Post number 865, September 11th, 2026.
This is a shorter list today, and I want to acknowledge the date up front. It's September 11th, and for a lot of us that date just lands differently. It probably always will. So on that note, let's get into a compact but really good set of reads.
We're going to start with a new leaderboard that Wiz has put out called the Cyber Model Arena. The idea is simple and kind of fun: fight. Take a bunch of different models, pair them with different agent harnesses, throw them at cybersecurity tasks, and see who wins. The results are genuinely interesting because they make a point that keeps coming up in this space: the harness matters as much as the model. The top of the board right now is Gemini 3.8 Flash Cyber running on ADK React, which scores just under 75 percent and finishes in about five minutes for under two dollars. But drop that exact same model into Claude Code as the harness, and the score falls to around 68 percent. Claude Opus 5 with Claude Code is right behind the leader at just over 71 percent, but it takes more than twice as long. What I like about this is that it finally gives us a side-by-side view of the thing that usually stays hidden: how the scaffolding around the model changes the outcome. You can't just say "Gemini is best" or "Claude is best" anymore. It's the pair that matters, and it also bakes in cost and time, which is exactly the conversation a security team should be having.
Next up is a Harvard Business Review piece with a question we don't ask ourselves often enough: Is Your Strategic Plan Too Ambitious? Or Not Ambitious Enough? Seroter flags some really useful advice here. Before you set a target, have you actually defined the problem? Are you getting the investment you need, or are you setting a big goal with a small budget and pretending that's fine? And are you looking out over the right time horizon? The trap on one side is aiming so high that you never stand a chance, and the trap on the other side is a plan so safe that it isn't worth doing at all. The piece is essentially a checklist for that moment before you commit, when you have benchmarks in mind and a rough sense of cost but no idea whether the whole thing is worth it. Worth a read if you're weighing a new market or a new offering.
The third item is a blog post from East River Source Control called What Comes After Git, and it's a thoughtful one. These folks have a controversial premise wrapped in a very pragmatic strategy. They don't think Git is the future of source control. It was designed around the constraints of 2005, built for the Linux kernel, and it's missing a lot of what organizations need at scale. But they also know you can't just walk away from Git, because Git is baked into everything. It's GitOps, not SvnOps. So their answer is clever: keep the Git protocol, but swap out the storage engine underneath. Your normal git client still connects, still speaks Git, but on their servers they don't actually store Git repositories. They store your code in a custom engine that scales horizontally in a way a Git repo as the source of truth can't. And because they're already speaking a protocol rather than shipping one storage format, the door is open to speaking multiple protocols down the road, including a native protocol for Jujutsu, the jj tool, which they're big fans of. The honest note at the end is good too: none of this is available yet, but they'll be opening things up soon. It's a bridge from the present to the future, and I think it's a smart way to frame the version control problem.
Fourth, Google shipped Genkit Go 1.13, and there's a lot here if you're building agentic apps in Go. The headline feature is resumable control flow. If your agent is in the middle of a multi-step tool loop and it fails on step five, you no longer lose steps one through four. Generate returns what it finished alongside the error, and you can send that history right back in to resume instead of redoing work. The same applies to full agent turns: a failed or aborted turn saves its completed tool rounds as a snapshot, and you can resume it later, even from a different process. There's also support for background subagents, where a delegation tool returns a task ID immediately and the orchestrator can check, wait for, or abort it while it keeps working. And there's a new A2UI plugin, which lets an agent stream interactive user interfaces straight to a frontend. Seroter also notes the Python version got a rev at the same time. It's a solid release for anyone doing multi-step agent work.
Fifth, Anthropic published a genuinely interesting writeup on what 1,000 small business owners taught them about AI. They ran a tour of workshops across ten cities, pairing AI fluency training with hands-on Claude Cowork sessions. A few lessons stood out. First, AI levels the playing field: eighty percent of attendees run companies with five to fifty employees, and people with no software background were building real tooling for their specific pain points in fifteen minutes. Second, owners are using AI for important work but worry about accuracy, and the advice that landed best was to treat the model like a new employee and build trust over time. Third, data and governance were the top concern everywhere they went. Fourth, building agents felt intimidating until people kicked off their first workflow, after which they felt empowered to keep going. And fifth, owners overwhelmingly wanted to learn from other owners. The through-line Seroter calls out is the timing: this is the moment to be out listening to a whole new set of users, plus existing users who are adapting to new ways of working.
Sixth is a piece on using the Open Knowledge Format, or OKF, to give your agent long-term context. This tackles a problem anyone who's done serious agent work knows too well: the agent finally gets up to speed on your codebase, feels like a real colleague, and then the context window fills up or you start a fresh session and it's back to square one. The author's point is that the tools already write plans and memories, but they stash them under your home folder, where colleagues can't see them and no other tool can read them. OKF is a fix: it's an open format from Google Cloud that's basically a directory of markdown files with YAML front matter. One file is one concept, the file path is the identity, and each directory has an index file and a log file. Version 0.2 adds provenance fields like sources, generated, verified, and stale_after, so you can decide how much to trust a document before you spend tokens reading it. The framing I liked best is the difference between retrieval and a wiki: retrieval re-derives the answer every time and throws it away, while a wiki is a compounding artifact where the synthesis sticks around. It sits naturally next to spec-driven development, where the plan is disposable and the durable decisions migrate into the knowledge base.
And last, a quick nod to a Medium post from Google on saving a billion tokens with Firebase AI Logic and on-device AI for Chrome. The article itself is paywalled or blocked, so I couldn't pull the details, but the title tells you the theme, and Seroter's note is telling: he built something with on-device AI this week and came away seeing both the power and the cost savings. On-device inference is quietly becoming the way to dodge the token bill on things that don't need a frontier model.
And that's the list. If there's a theme tying today together, it's that the layer around the model keeps getting more interesting than the model itself. The harness changes the cyber score, the storage engine changes what Git can do, the format changes what your agent remembers, and on-device changes what you have to pay for. Thanks for reading, and see you next time.
- Cyber Model Arena
- Is Your Strategic Plan Too Ambitious? Or Not Ambitious Enough?
- What Comes After Git
- Genkit Go 1.13
- What 1,000 small business owners taught us about AI
- Using OKF to provide long term context to your agent
- Three ways to save a billion tokens with Firebase AI Logic and on-device AI for Chrome