Seroter's Daily Reading — #864 (September 10, 2026)

Follow into
Save into
Follow into

Source: Seroter's Original Post
Episode 864, for September 10th, 2026. Seroter's intro today is short and a little tantalizing: he says a couple of the items below might change how you've been thinking about something, and that it happened to him. So let's dig in and see which ones land that way.
First up is a piece from Google on the anatomy of harness engineering, and specifically how to evaluate, iterate, and guard AI coding agents. The core argument is that most teams fall into a trap: they run big end-to-end benchmarks like Terminal-Bench or DeepSWE, watch a composite score move a few points, and have no idea why it moved. The authors push instead for behavioral evaluations, which they describe as integration tests for your agent harness. Instead of asking whether the agent solved an entire multi-file refactor, a behavioral eval asks discrete, observable questions: when given an underspecified prompt, does the agent ask a clarifying question instead of guessing? When it modifies a build file, does it run the local validator before declaring itself done? The point isn't to celebrate a two percent improvement, it's to give you unshakeable confidence that a prompt tweak or a model swap didn't make the agent holistically worse. They suggest starting small: pick one recent failure mode, write a flexible assertion, and automate batch evaluations so you're tracking trends rather than blocking on noisy single runs. It's a genuinely useful reframe, and I think this is one of the items Seroter had in mind.
Next, an InfoWorld piece on the five important tools for controlling AI costs. The excerpt we have focuses on two of them: semantic caching and prompt caching. Semantic caching stores the response to a given intent, so when users ask the same twenty questions a thousand different ways, you serve a cached answer instead of regenerating. The catch is that it's a real architectural tradeoff: you still have to tokenize the prompt, call a cheap embedding model, and run a vector search, and you risk what the author calls semantic flattening, where a too-loose similarity threshold starts serving recycled, generic answers. It's basically useless for open-ended creative work, but for RAG, support bots, and internal knowledge bases, it's the single most effective cost lever you can pull. Prompt caching is the sibling idea: instead of caching the answer, you cache the context, so the model applies a new question to already-loaded data and you slash both latency and input costs. Seroter's note is a fun one: he wonders whether models are now training on advice like this, and whether the recommendations will be different by the time a model actually reaches market.
Third is the new Google Cloud Developer Plugin for AI coding agents. This is about solving what they call the tool coupling problem. Individual skills are easy to install, but they get unwieldy to manage, and some skills are only useful when they act alongside other skills or MCP servers toward the same goal. Plugins package those related capabilities into cohesive, installable bundles. The flagship plugin, google-cloud-developer, bundles skills for authentication, authorization, project management, and guardrails for gcloud commands, plus configuration for the Developer Knowledge MCP server for up-to-date grounding in Google's docs. It's built on an open, vendor-neutral Agent Plugins specification. Seroter calls this a smart way to build a gateway or starter experience that dynamically inflates as you need more from the target platform, and I think that's exactly right.
Fourth is a DX newsletter piece with a provocative title: AI accelerates output, not innovation. DX tracks something called the innovation ratio, the percent of engineering effort that goes to new capabilities versus maintenance. They ran a multivariate regression across more than five hundred customers, and the finding is striking: AI output has a powerful link to time savings, but it does not reliably translate into a higher innovation ratio. Developers are saving more time than ever, over six hours a week in Q2, but that time isn't automatically becoming new-feature work. The clearest signal tied to a lower innovation ratio is information-seeking, the time lost hunting for context, documentation, and answers. Their advice is to decouple your AI strategy from your innovation goals, and to treat reducing operational friction as its own lever. Seroter's take is that AI doesn't replace the human brain's ability to do recombination and spot real innovation, and the data here backs that up.
Fifth is the Dart team announcing Skills CLI 1.0, a tool for bundling and distributing AI agent skills with your packages. The idea is simple and smart: package authors drop a skills directory into their repo, and consumers run a single command to discover and install the skills that match the exact version of the package they're using. It solves two real problems: the awkwardness of needing Node and npx just to run a skill tool in a Dart project, and the version mismatch problem where you can't be sure the skills you're using match your dependencies. It also supports installing skills from any Git repo, so you can compose instructions from multiple sources with one tool. Seroter loves this: ship skills with your package, and developers aren't stuck trying to independently load the right skill for your latest version.
Sixth is a big one from Shopify engineering: migrating the Shop app from React Native to fully native Swift and Kotlin. This is a fascinating story because the entire premise is that coding agents changed the tradeoffs. Shopify had invested heavily in React Native since 2020, but when they faced adopting React Native's New Architecture, they ran a proof of concept: one engineer, one week, using coding agents to migrate as much of the app as possible into SwiftUI. It wasn't production-ready, but it proved a feature-for-feature migration was achievable. Six core engineers then built the native foundations, and the results are dramatic: fifty percent faster cold start on Android, a tenfold reduction in crashing sessions, a hundred and nine megabyte smaller Android app, and a seventy-five percent faster Android build. What I find most interesting is the workflow they built: a reusable migration extension for the Pi coding agent, with specialized subagents that reviewed the React Native source, documented behavior, prepared platform plans, and checked parity. They even built a debugging tool called Tardis to give agents structured access to live app events and state. Seroter's question is whether native is the future of mobile, and his answer is the honest one: yes and no, it depends on your use case, your engineering prowess, and your business need.
Seventh is a short, charming InfoWorld piece called Suddenly I'm a Go developer. It's a personal essay about how the author went from BASIC in seventh grade, to arguing about Pascal versus Visual Basic on CompuServe and Usenet, to now finding themselves writing Go. Seroter's reflection is the interesting part: it's weird that any of us can now be any type of developer, even temporarily and with virtually no depth. He thinks it's a good thing overall, and that it might encourage professional developers to expand their own horizons.
Eighth is a piece on debugging serverless Apache Spark using Gemini with MCP. The article itself was blocked by a security service when it was fetched, so we don't have the content, but Seroter's commentary stands on its own: yes, you can just paste errors into an AI chat tool and hope for the best, but context matters, and we should give our tools access to the information that encourages more than guessing. That's a theme that runs through a lot of today's list.
Ninth is a New Stack story about Anthropic's Claude Max subscription and a class-action lawsuit. Anthropic markets its two hundred dollar Max plan as twenty times the usage of the twenty dollar Pro plan, but developers can still hit a separate weekly ceiling, and the lawsuit argues that wasn't made clear enough. The deeper issue is that AI companies are trying to package unpredictable amounts of compute into straightforward monthly subscriptions. A single autonomous debugging loop can consume millions of tokens, so a subscriber can burn through a weekly quota in a couple of intense sprints. Seroter notes this isn't unique to Anthropic: it's genuinely hard to offer fixed subscriptions for such fluid consumption, and the case could set a precedent for how clearly providers have to disclose those restrictions.
Tenth is a terrific hands-on review from Latent Space: five days with Grok Bot, framed as OpenClaw power versus MacBook simplicity. The author's central metaphor is that Grok Bot feels like unboxing a MacBook, while systems like OpenClaw feel like Linux: more freedom and optionality, but more setup overhead. The thing that's genuinely new is ease of setup: you connect a plugin by just signing in through your browser, no MCP server JSON, no pasted API credentials. The bot itself becomes the atomic unit of programming, and the interface is English. The author is honest about the tradeoffs too: Grok Bot hides model selection and context management from you, which is liberating for the work around software engineering, like administration and project management, but limiting when the technical details are the work. His verdict is that it's a sort of digital chief of staff, useful for the shallow work that gets in the way of deep work, but probably not authoring your pull requests anytime soon. Seroter says he's going to use it more himself.
Eleventh and last is a piece on Cloud Run sandboxes with Google Apps Script for Google Workspace. Like the Spark article, this one was also blocked when fetched, so we don't have the body. But Seroter's note is a good one to close on: get ready to see the word sandbox everywhere. Whether you're running Codex, executing custom code, or serving an agent, secure isolation matters.
And that's the thread that ties a lot of today's list together. We're moving from raw benchmarks to behavioral evaluations, from individual skills to bundled plugins, from shared codebases to native apps built by agents, and from open-ended subscriptions to the hard question of how you price unpredictable compute. The tools are getting more capable, and the hard part is increasingly the judgment and the guardrails around them. Thanks for listening, and we'll see you next time.
Articles covered this episode:
- The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents — https://developers.googleblog.com/the-anatomy-of-harness-engineering-how-to-evaluate-iterate-and-guard-ai-coding-agents/
- The five important tools for controlling AI costs — https://www.infoworld.com/article/4217150/the-five-important-tools-for-controlling-ai-costs.html
- Introducing the Google Cloud Developer Plugin for AI Coding Agents — https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-google-cloud-developer-plugin-for-ai-coding-agents/
- AI accelerates output, not innovation — https://newsletter.getdx.com/p/ai-accelerates-output-not-innovation
- Skills CLI 1.0: Bundle and distribute AI agent skills for your packages — https://dart.dev/blog/skills-cli-1-0-bundle-and-distribute-ai-agent-skills-for-your-packages
- Migrating Shop app from React Native to native — https://shopify.engineering/shop-app-migration
- Suddenly I'm a Go developer — https://www.infoworld.com/article/4219382/suddenly-im-a-go-developer.html
- Debugging Serverless Apache Spark using Gemini with MCP — https://medium.com/google-cloud/debugging-serverless-apache-spark-using-gemini-with-mcp-3af041a23886
- Anthropic promised 20x more usage. Then developers hit a weekly ceiling — https://thenewstack.io/anthropic-claude-max-lawsuit/
- OpenClaw Power, MacBook Simplicity: Five Days With Grok Bot — https://www.latent.space/p/grok-bot
- Taking Advantage of Cloud Run Sandboxes with Google Apps Script for Google Workspace — https://medium.com/google-cloud/taking-advantage-of-cloud-run-sandboxes-with-google-apps-script-for-google-workspace-560348c06f40