↵ select ↓ ↑ navigate esc close

Seroter's Daily Reading — #850 (August 20, 2026)

Seroter's Daily Reading ·

Listen: https://blossom.buildtall.systems/7214be72ca737bd5363537dfa6d9d54dbe9caa0883191efa98d196c8f9e12832.mp3

Source: Seroter's Original Post


Seroter's Daily Reading, episode 850. August 20, 2026.

I'm adding things to my "you should try this technology" list much faster than I'm getting through it. At some point I'm going to declare bankruptcy and start over. Tomorrow's looking like a clean "no meeting day" so I stand a chance to get through two of my top ones. Stay tuned.

Let's start with Expanding Google Antigravity for enterprise customers landing in enterprise. If you've been following along, Antigravity is Google's AI coding agent. It started as something individuals could use, and now it's graduating into Gemini Enterprise. The big additions are IDE integrations across popular tools, pooled usage across teams, and the backend controls that enterprise IT teams demand before they'll approve anything. The blog post is full of quotes from Accenture, Cognizant, Deloitte, Wipro, AirAsia, and Datamatics, all talking about how the combination of agentic AI and enterprise governance lets their teams move faster without giving up control. What stands out to me is the framing: these organizations aren't talking about AI experimentation anymore. They're talking about measurable outcomes, real software delivery improvements, and teams evolving from writing code to orchestrating outcomes. That feels like a meaningful shift in how the enterprise AI story is being told.

Next up, Addy Osmani has a piece on Practical Loop Engineering, and this one is really useful if you're trying to figure out how to actually run AI agents at work. He opens with the recognition that most of us now have five to ten agents running concurrently, and not all of them deserve the same level of attention. Some tasks, like writing documentation or checking test coverage, are safe to fully delegate with clear stopping conditions. Others, especially anything touching auth, finance, or system access, deserve closer watching. He walks through the core primitives in Claude Code and Codex: goal-based loops that keep iterating until a measurable condition is met, and time-based loops that run on a schedule like a cron job. The goal primitive is best for bounded, specific work with deterministic completion criteria, like running Lighthouse scores until a page hits a performance threshold. The loop primitive is for recurring checks, like polling a GitHub repo for new issues every hour. He also describes what he calls proactive loops, the highest rung, where the system runs on events or schedules without any human in real time, managing things like bug triage or dependency upgrades. One piece of advice that landed with me is to never let the agent that did the work decide the work is good. One sub-agent drafts, a separate one verifies. The goal evaluator behind the scenes doesn't judge quality, it only checks whether hard rules were met. Addy learned this the hard way when he almost pushed changes an agent had generated without reviewing the actual implementations closely enough. Delegating the task is fine. Delegating the taste and judgment is not.

On the observability side, Jason Davenport has a post on Add telemetry and observability to your AI agents. He makes a useful distinction right at the top. In traditional application development, traces were expensive telemetry you deployed mostly to debug failures. In agent development, traces are central to measuring system performance and keeping things well tuned, not just finding what broke. In agent telemetry, a trace is called a trajectory, and it captures every step, every prompt, every response, across the entire reasoning loop. He also covers session state, which is often an afterthought in agent examples. In-memory session stores are fine for demos, but production agents need durability: handling connection failures, letting users resume sessions later, storing information securely. He walks through a reference implementation using Google Agent Development Kit, Cloud Run, and Cloud Spanner for session management, with BigQuery for trajectory logging. He notes something interesting: as models get more capable, the math changes between using a framework with all the tools out of the box versus having agents code the telemetry framework themselves. Gemini 3.6 Flash essentially single-shot the telemetry implementation he needed.

Here's a piece that made me laugh and then think harder than I expected. The New Stack ran an article titled Your coding agent got the onboarding your developers never did. It opens with Steve Yegge's argument that coding agents are basically sentient, deserve recognition, and should be able to refuse tasks. You don't have to agree with that framing to notice the pattern underneath. Whatever you think about model sentience, look at what teams are doing to make their agents productive, then compare it to what those same teams were willing to do for the actual developers sitting next to them. Agent onboarding documents like AGENTS.md are now written carefully, kept under a few hundred lines, optimized for every token because every line gets re-read in every session. That same discipline almost never gets applied to human onboarding docs. Google's DORA research found that 90% of organizations have adopted AI-assisted development, a 14-point jump in a single year, but the benefits aren't being realized evenly. AI amplifies existing software delivery quality. Strong systems get stronger, dysfunctional ones get more chaotic. Small-batch changes, fast test suites, clean deployment pipelines. All things DORA has been saying for a decade, all things that got waved off as inconvenient when humans were making the changes, and now they're suddenly urgent priorities because agents need them to work reliably. Telemetry from Faros AI found that developers working with AI assistance are juggling 67% more pull request contexts and 18% more task contexts per day, up sharply from 47% and 9% the prior year. Work restarts are up 14%. More than a quarter of in-progress tasks sit untouched for a week or longer. Both humans and agents are drowning in context switches that better batching would have prevented. And teams are learning that lesson for the sake of their agents' throughput. It took DORA a decade of survey data to get the same lesson taken seriously for people.

On the human side, CNBC and SurveyMonkey ran a survey of workers and students about AI and job security in their piece on Only daily AI users feel positively about job security. Only daily AI users feel positively about the technology's effect on their job security. People who use it a few times a week or less say AI makes them feel less secure. Daily users were also the only group optimistic about the overall job market. The rest are worried, even when replacement isn't the near-term reality. They worry that AI will reduce the value of their expertise or change the work they're known for. More than half of employers don't have an official AI policy, which adds to the uncertainty. Employees feel they could be held accountable for a tool's mistakes, and that AI is something being done to them, not with them. The pattern connecting this to the previous article is pretty stark. When organizations invest in making agents productive, the humans around them often feel the opposite.

Let's talk about Git. Datadog published a deep dive on 20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog, the system they built to serve code to CI at massive scale. Their context is extreme. Millions of CI fetches per week across thousands of repositories, including monorepos with years of history and hundreds of thousands of files. But the lessons are broadly applicable. The problem started with AI coding agents hitting Git far harder and more often than even their most active human contributors ever could. The traditional fixes didn't work. Adding nodes made replication overhead worse. A CDN in front doesn't help because the expensive part of a fetch isn't a static byte range, it's computation specific to each client request. Cloning on demand from GitHub just moves the thundering herd upstream and into rate limits. Their insight was that the CPU concentrating on a few nodes was the real bottleneck. Gitretriever uses many independent mirror pods, each keeping a fresh local copy, with a separate relay fleet fanning out reads to CI jobs. Mirrors sync from GitHub in a tight loop, pulling each changed reference in parallel so no single request concentrates expensive delta compression work. Relays reuse that packfile work instead of fetching again. They also added a pack cache, and about half of all pack-building fetches are served directly from it. Traffic grew about twenty times in four months while median latency stayed around 40 milliseconds. The old backend's fetch-serving CPU dropped by three to four times. One observation that stuck with me: most non-CI workloads don't actually need a full repository clone. They wanted a single file at a commit, the SHA a branch pointed to, the list of changed files. Cloning an entire repository for that was enormous overkill, yet internal services, developer tools, and AI agents were doing it constantly. They added a read-only HTTP API for those queries. Resolving a ref or reading a file takes single-digit to tens of milliseconds. A shallow clone of a large monorepo takes around 75 seconds. The cheapest answer is often an API call rather than handing out a repository clone.

Harvard Business Review has an article on Is Your Organizational Culture Too Nice?. The argument is that cultures optimized for comfort, where everyone gets great performance reviews and every opinion gets a seat at the table, often underperform cultures that aim for something harder: caring deeply about people while delivering candid, results-focused feedback. This isn't a new idea, but it connects nicely to the AI adoption theme. The organizations investing heavily in AI tools while leaving their people underinvested in are essentially choosing the comfortable path for the machines and the hard path for the humans.

Google published on how AlloyDB's ScaNN index now scales to 10 billion vectors in their post on How AlloyDB ScaNN scales vector search to 10 billion vectors. This is a relational database, not a purpose-built vector store, and they're handling that scale through a four-level tree architecture, which is a significant jump from the previous two and three-level configurations. The technical details include Top-K branching, SOAR algorithms, centroid adjustment, and balanced tree shapes to maintain accuracy while keeping build efficiency reasonable. This matters for enterprise agentic AI applications where vector search underpins memory, retrieval, and knowledge grounding at scale.

Finally, O'Reilly has a piece on When Your Buyer Is an AI Agent, and it's a fascinating look at how autonomous procurement is already happening. Maersk deployed Pactum's AI negotiation agents to handle freight lane contracts with carriers. The agents achieved a 96% agreement rate, requiring no human intervention, and secured rates 22% lower than human negotiators for identical lanes. G2's research found that 51% of B2B buyers now start their research with AI chatbots, up from 29% the prior year. By the time buyers reach your website, much of the funnel has already been run by the answer engine. Sales cycles used to start with relationship building. That initial conversation has shifted from buyer-focused exploration to a process of self-discovery. The piece covers three pressure points on traditional B2B commercial models: per-seat licensing breaks when agents operate continuously across time zones at machine speed, relationship-driven sales breaks when agents evaluate purely on measurable metrics before passing results to humans, and customer success renewals break when an agent computes ROI from API logs and cross-references competitor pricing without any susceptibility to social influence. The strategic recommendations are outcome-based pricing, machine-readable product surfaces, and agent-compatible commercial authorization. This is a space worth watching closely because the buyers are already changing.

That's episode 850. A theme that kept showing up across these articles is the gap between the investment and care being poured into AI systems versus what human practitioners receive. Whether it's loop engineering discipline applied to agents while developers are left without it, or the sudden urgency around documentation and test quality for the sake of agent throughput, or the daily AI users feeling secure while everyone else feels left behind. The agents are getting excellent onboarding. The humans are still waiting.


  1. Expanding Google Antigravity for enterprise customers
  2. Practical Loop Engineering
  3. Add telemetry and observability to your AI agents
  4. Your coding agent got the onboarding your developers never did
  5. How Google Cloud MCPs Give Claude Code Inside Access
  6. Only daily AI users feel positively about job security
  7. 20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog
  8. Is Your Organizational Culture Too Nice?
  9. How AlloyDB ScaNN scales vector search to 10 billion vectors
  10. When Your Buyer Is an AI Agent