select navigate esc close

Seroter's Daily Reading — #829 (July 21, 2026)

Seroter's Daily Reading ·

Listen: https://blossom.buildtall.systems/a7f5b982fbd674568b094ed261dc908fac4034eda1924e9d55c590bb1356699a.mpga

Source: Seroter's Original Post


Seroter's Daily Reading, episode 829, July 21, 2026.

This week my team spent time in San Diego with a customer taking a use case from idea to MVP. It's a really cool exercise to watch a team come together and iterate quickly, learn, and ship. Good reminder that the craft of building is still alive and well. And on that note, let's get into the reading list.

Google had a big model release this week, introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The headline on 3.6 Flash is that it's a solid step up from 3.5 Flash in coding and knowledge work, but the interesting part is the efficiency gains. According to the Artificial Analysis Index, it consumes 17% fewer output tokens, and in some coding benchmarks like DeepSWE from Datacurve, they see up to 65% fewer tokens. That's not just saving money — it's meaningfully faster responses and less context bloat for multi-step agentic workflows. 3.5 Flash-Lite is the budget option, prioritizing speed and cost above all else. And 3.5 Flash Cyber is a new specialized model for cybersecurity tasks, paired with Google's CodeMender agent for code security work. These are well-priced, efficient models, and the efficiency story is the most interesting part to me. If you're building agents at scale, cutting token usage by double-digit percentages compounds fast.

Next up, there's a piece from the Aha engineering blog on How do you stay familiar with the code when it's written by an LLM? If you used an LLM to write code you shipped a month ago, how well do you actually remember it? Would you know if something you committed today broke it? The author is essentially asking whether code you don't understand is code you're responsible for. And the answer is yes, it is. So the piece offers concrete practices for staying engaged with LLM-generated code. Make mistakes — meaning create opportunities to be surprised, since smooth success doesn't build memory the way failure does. Type things yourself, even if the LLM could do it faster, because the act of typing reinforces understanding. Ask questions constantly: how would I have done this? Why did it do it this way? Explore the codebase actively, not just through GitHub diffs. Generate an HTML explainer for complex features. Review the LLM's work like it's a pull request from a colleague, leaving comments and asking for revisions. And build your awareness — when something feels wrong or overcomplicated, trust that instinct. The overall argument is that LLMs amplify what you already have. If you have nothing, they produce nothing useful, very fluently. And if you're not careful, they keep you from building the intuition that makes you valuable.

Google Cloud published a guide to AI tokenomics with Eleven Principles for Token Efficient Software Engineering. This one has a lot of practical advice for teams trying to work efficiently with AI coding assistants. Start with a balanced model like Gemini 3.5 Flash and scale up only when needed. Use reusable skills from the beginning so you're not re-explaining your context in every prompt. Delegate output-heavy tasks like deep research to sub-agents and only reconcile the final results. Divide and conquer by using high-reasoning sessions to generate a detailed plan, then execute that plan in a clean, low-token session. Shift verification left by running local builds and tests early. If the agent drifts, use Undo rather than piling corrective prompts on top of a broken state. Be specific with context rather than micromanaging — pointing to an exact file or adding a comment like "SHOULD BE X, NOT Y, FIX THIS" goes further than a vague request. Iterate on rules rather than repeatedly correcting the agent, so the fix persists. Avoid uncontrolled supervisor loops that scan for pending work since they can burn your entire token budget. And start new sessions for each new topic so the agent doesn't carry irrelevant context. The underlying message is that tokens aren't infinite, and token optimization is really about directing the AI's attention where it matters most.

Then there's a piece from Theocharis.dev that hits a lot of the same themes but from a more personal and philosophical angle. The title is "The LLM Critics Are Right. I Use LLMs Anyway." The author attended Local-First Conf in Berlin The author attended Local-First Conf in Berlin and noticed something striking: Armin Ronacher, who built Flask and founded the company behind Pi, was on stage auto-closing LLM-generated PRs and issues because the flood was overwhelming. And yet the audience was largely running Claude Code. That dissonance — using LLMs while applauding talks that critique them — is the subject of the piece. The author walks through the fair criticisms: LLMs produce a lot of low-quality content that clutters open source, making it hard to trust contributions. Junior engineers can't signal effort anymore because anyone can generate a plausible PR in minutes. Seniors have no incentive to mentor juniors if the mundane tasks are outsourced to LLMs. Geopolitically, there's real risk — the US already cut off non-citizens from Anthropic's frontier models under an export control order. And LLMs tend to slowly absorb the dominant viewpoint of their training data, which has its own risks. But the author's conclusion is that LLMs amplify what you already have. Opinions, structure, frameworks — those come out sharper and faster. Without them, you just get fluent nonsense. The test the author applies: would you stand in front of an audience and read it out loud? If not, it's slop. That framing — human credibility and judgment as the irreducible core — keeps showing up this week, and I think it's right.

Moving to product management, there's a piece from David Pereira on Substack asking whether PMs are becoming irrelevant? His take is that the outdated PM role is over — the inside-out version that spends most of their time writing backlog tickets, coordinating tasks, and managing predictability. Those things can be AI-augmented now. But the strategic PM, someone who identifies problems worth solving and creates genuine customer and business value, is more relevant than ever. The interesting warning in the piece is against chasing the "AI PM" label. Yes, organizations are paying for it, and yes, there's demand for prompt engineering and agent skills. But the author's point is that positioning yourself as the AI PM makes your job more about doing AI and less about doing product management. The money might be good short-term, but the bottleneck has shifted from delivery speed to direction quality. Companies don't need more feature factories. They need people who can figure out what to build and why.

From CIO, there's a piece on the career path from How do you go from junior to staff engineer when AI writes the code? The question is essentially: if the traditional path to seniority was writing code, reviewing it, making mistakes, and building intuition, what happens when LLMs handle most of the implementation? The piece references Kent Beck, who recently spoke at an event titled "Juniors FTW." His argument is that the industry is moving so fast that being new is actually an advantage. Junior engineers are too fresh to have absorbed what everyone "knows" is impossible, which makes them less biased and more creative. The real risk isn't that juniors can't write code — it's that the signal gets lost. When a senior reviews a junior's PR, they used to be able to gauge effort and understanding. Now they can't tell if the junior spent hours or minutes. And if seniors don't need juniors for the mundane work, what incentive remains to teach them? The piece suggests that mentoring may actually get easier if you bring agents into the mix to get juniors productive faster, while keeping the human mentorship focused on judgment and direction rather than syntax.

From Harvard Business Review, there's a piece on Frontline Workers Know How to Solve Your Organization's Biggest Problems. A pharmaceutical company had a Patient Service Group that handled daily calls with patients about a critical medication. They introduced an AI-powered dashboard to automate call planning, but the representatives felt "behind the glass" — disconnected and overlooked. The company ran a structured process: confidential interviews with about 20% of the group, custom surveys, focus groups to review findings, and then co-creation of solutions with employees and leaders. What emerged was a set of changes that employees themselves generated and ranked. Two were implemented immediately: adjusting the algorithm to give representatives more leeway on refill reminder calls, and introducing a clearer competency framework for coaching conversations. The pilot was treated as an experiment with defined metrics. Refill rates held steady. Pulse survey scores improved. The lesson is simple but easy to skip: senior leaders are often too far from day-to-day reality to know what needs to change. The people doing the work usually do. This shows up in the AI space too — the best AI deployments come from people who understand the domain deeply enough to know when the AI is wrong.

Sentry published a guide on How to structure a log, and it's a great reminder that good observability hygiene matters regardless of whether AI is in the picture. The core advice: use stable event names in a consistent pattern like "domain.action", use snake_case across all attributes, flatten nested objects into dot notation, include a low-cardinality "result" field like "succeeded" or "failed" for grouping and visualization, and only use primitive values in attributes, not objects. The author even offers an ESLint plugin to enforce these patterns. If you're building AI systems that generate logs, or using AI to debug systems, the quality of those logs becomes even more critical. Garbage in, garbage out — and this is especially true when the garbage is in your observability data.

Google also published a list of 13 hands-on demos to build on Gemini Enterprise Agent Platform. These cover the full lifecycle: building basic agents with ADK, connecting them to data via MCP, building event-driven workflows with human-in-the-loop approval gates, deploying to production with Agent Runtime, securing agentic coding with threat modeling and pre-commit hooks, controlling access with Agent Gateway and IAM, and running evaluation flywheels to measure whether your prompts are actually improving things. The Agents CLI integrates directly into your coding agent and scaffolds, deploys, and monitors agents for you. This is a practical resource if you're trying to move beyond toy demos into real agent deployments.

From the PostHog newsletter, there's a piece on building what they call "2030-shaped software." The argument is that by 2030 The argument is that by 2030, AI agents will be the primary users of software, not just the assistants. This reshapes how products should be designed. Human work splits into two types: judging and deciding on direction, and understanding and verifying what happened. The actual "doing" UI — buttons, forms, task completion — goes away, replaced by surfaces that support judgment and verification. For product builders, the practical implication is that if a human can do something with your product, an agent should be able to do it too. That means complete APIs, MCP servers, and proper tooling as first-class surfaces, not afterthoughts. PostHog is walking through this themselves, starting with their Desktop app as a blank slate for building agent-first. The question to ask of any UI surface is: is a human looking at this to verify trust and judgment, or to act? If it's an action, give the agent the capability. If it's to verify or judge, build for that specifically.

And finally, from the Netflix Technology Blog, there's a detailed look at In-House LLM Serving at Netflix. They run the full stack themselves They run the full stack themselves — model deployment through inference inside their production environment rather than using hosted APIs. The architecture runs on vLLM under NVIDIA Triton, with a Java control plane handling deployment and autoscaling. A few production lessons stand out. First, the vLLM backend is architecturally better than the Python backend for packaging because it generates I/O specs dynamically rather than freezing them at packaging time — which means frontend and model upgrades can happen independently. Second, the OpenAI-compatible API frontend is essential for ecosystem compatibility, but they had to patch it because response_format was silently dropped before reaching vLLM, meaning JSON output requests weren't getting guided decoding constraints. They git-subtreed the fix. Third, their Prometheus metrics endpoint had a gap: Triton surfaced only 9 of 40-plus vLLM metrics, missing critical ones like KV cache utilization. They built an HTTP proxy to merge both endpoints. And fourth, their constrained decoding implementation at scale exposed a GIL bottleneck in vLLM V0 — custom logits processors ran per-request on CPU sequentially. The fix was migrating to vLLM V1, which moved logits processing to batch level and let them rewrite the hot path in C++ with multithreading. They also had to handle edge cases like partial prefills and request preemption that stateful constraint logic in the decode loop introduced. This is the kind of operational detail you only learn under production load, and it's exactly the kind of thing worth sharing.

Across this week's articles, a few threads stand out. First, staying involved in what AI produces is a recurring theme — whether that's staying familiar with LLM-generated code, keeping humans in the loop on PM decisions, or mentoring juniors through the shift. The instinct to step back and let AI handle it is understandable, but the best outcomes come from human engagement guiding quality. Second, efficiency and token discipline come up repeatedly — not as a cost-cutting measure, but as a sustainability and focus discipline. And third, there's a consistent message that the bottleneck has shifted from building to deciding what to build and whether it worked. That's the work that humans need to own.


  1. Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — Google
  2. How do you stay familiar with the code when it's written by an LLM? — Aha Engineering
  3. Guide to AI Tokenomics: Eleven Principles for Token Efficient Software Engineering — Google Cloud
  4. The LLM Critics Are Right. I Use LLMs Anyway — Theocharis.dev
  5. Are PMs becoming irrelevant? — Substack (dpereira)
  6. How do you go from junior to staff engineer when AI writes the code? — CIO
  7. Frontline Workers Know How to Solve Your Organization's Biggest Problems — Harvard Business Review
  8. How to structure a log — Sentry
  9. 13 hands-on demos to build on Gemini Enterprise Agent Platform — Google Cloud
  10. 2030-shaped software — PostHog
  11. In-House LLM Serving at Netflix — Netflix Technology Blog