[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Follow into
Save into
Follow into

Today was a tough news cycle to launch anything; we ordinarily promise to cover any new decacorn fundraises so Cognition’s $48B round and Mistral’s $24B round would normally have made it; we love imagegen so GPT Image 2.5 would have been its own headline; we covered the Dreamer story closely so their relaunch as Meta’s Muse agent should have made it; but.. yknow… the bar is higher these days.
The summaries below capture the substantive facts; we recommend not looking too deep into the authorship drama as OpenAI and the authors have pretty much laid out enough detail to conclude that OpenAI’s achievement is real though the process is in some despute.
AI News for 9/7/2026-9/8/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
OpenAI’s Navier–Stokes Result, Credit Dispute, and the Emergence of Massive Test-Time Compute
- OpenAI’s proposed Navier–Stokes solution dominated the day. OpenAI said an internal model “significantly more capable than GPT-6 Astra” produced a proposed proof in 88 hours using roughly 10,000 agents, followed by another 17 hours of Lean formalization/verification with Astra, according to summaries and reactions from @TheTuringPost, @polynoamial, and @sama. OpenAI stressed that its proof differs from the independent researchers’ work and addresses a different Euler setting; it also said no specific user data was accessed for this effort, while conceding it cannot rule out de-identified derivative data from product usage having helped model improvement more generally in the past @OpenAI.
- The technical meta-point is test-time compute scaling. Several observers highlighted the implied economics and trajectory: what cost millions today could become consumer-accessible quickly, just as ARC-AGI costs collapsed from hundreds of thousands to tens of dollars @polynoamial. Others estimated the proof run at 130B output tokens and perhaps $10M–$40M API-equivalent cost depending on input-token scale assumptions @scaling01. The strongest consensus signal was that unstructured parallel test-time compute plus orchestration is now a first-class scaling axis, not just pretraining or post-training @eknight, @eliebakouch.
- The controversy centered on priority, data contamination, and norms. Sam Altman and Sébastien Bubeck argued OpenAI heard rumors that Anthropic-associated researchers had solved a Millennium problem, then tested whether OpenAI’s models could do the same; when OpenAI learned the other team had Euler but not Navier–Stokes, it says it offered coordination, priority on Euler, and possible lead authorship for Tristan Buckmaster on a rewrite of OpenAI’s proof @sama, @SebastienBubeck. Critics focused less on direct spying—which many deemed unlikely—and more on whether derived user data or public rumors should have triggered stricter checks, and on whether this behavior will chill open scientific exchange @aidangomez, @johnschulman2, @simonw.
- Mathematicians’ reaction is becoming a substantive governance issue. Terence Tao’s cautionary comments, amplified by @fchollet and @GaryMarcus, framed the key risk: if even rumors of progress can trigger industrial-scale AI efforts that “flatten” a research direction, fields may move toward secrecy and away from long-standing open-science norms. Separately, @stevenstrogatz emphasized that prior public work by Córdoba and Martínez-Zoroa supplied the key strategy that others built on.
Meta’s Muse Launch and the Personal-Agent Security Architecture
- Meta launched Muse, a consumer-facing “personal AI agent” positioned as always-on, app-connected, browser-capable, and goal-oriented, with strong distribution through Meta properties and integrations @finkd, @alexandr_wang, @MetaNewsroom. Product details repeatedly surfaced: persistent isolated Linux VMs, browser use, WhatsApp/app interfaces, and connectors to services like Gmail, Calendar, Outlook, Plaid, OpenTable, Docs, Spotify, Peloton, plus unique Meta-native connectors for Instagram, Messenger, Facebook, and Marketplace @alexandr_wang.
- Security architecture is the differentiator being pushed hardest. Meta’s team said each Muse runs in its own secure VM, actions are mediated by a separate Sentinel, secrets are never directly exposed to the agent, sensitive actions require approval, and there is a public bug bounty up to $300k @shengjia_zhao, @alexandr_wang. There’s also explicit commerce infrastructure: Stripe Link for payments with an agentic payment protection / refund guarantee, plus incoming Shop Pay integration @alexandr_wang.
- Early reception from practitioners was notably positive, especially on permissioning, secrets management, and consumer utility. Commentary from @matthuang, @signulll, and @lilyjclifford suggests Muse may be one of the first broadly legible personal-agent products where context and access, not raw model IQ, are the bottleneck. Meta also said usage exceeded internal projections by 10x on day one @alexandr_wang.
- Model and ecosystem placement: Meta’s Muse Spark 1.3 was quickly exposed in third-party tooling like Cursor @cursor_ai, while arena-style benchmarking positioned Muse Spark 1.3 Max as price/perf competitive in web-dev coding workloads @arena.
OpenAI’s Image 2.5 Release and Astra Rollout
- OpenAI also shipped ChatGPT Images 2.5, though it was partially overshadowed. The release emphasizes up to 50% lower latency vs Images 2.0, better realism, stronger edit consistency across repeated edits, comment-based localized changes, transparent backgrounds, and a new Sketch tool for guided generation @OpenAI, @ChatGPT, @sama.
- Two API variants were introduced: GPT-Image-2.5 Flare for speed/quality and Sunburst for higher-precision detailed work @reach_vb. Arena results claimed #1 and #2 positions across text-to-image, image-edit, and multi-image-edit leaderboards, with especially large gains in multi-image editing @arena. Integrations landed quickly on fal, Higgsfield, Manus, and Hermes Agent @fal, @higgsfield, @ManusAI, @Teknium.
- Astra availability widened materially. OpenAI said GPT-6 Astra is now fully rolled out to Plus, Pro, Business, and Enterprise users in Codex and ChatGPT Work @OpenAI. Community demos showed strong practical computer-use performance: @theo reported Astra compiling and running Super Smash Bros. Melee on macOS at 120 FPS after a roughly 6-hour loop, while Vals reported Astra nearly saturating an unreleased computer-use eval by building a Minecraft Nether portal in under 3 hours with no specialized harness @ValsAI.
Agent Harnesses, Post-Training, and Serving Infrastructure
- Harvey + Baseten’s M&A diligence work is one of the clearest model-harness co-optimization case studies. Their recursive language model (RLM) harness uses a root agent to search a data room, delegate to sub-agents for document review, and aggregate findings over corpora up to 80M tokens. On the synthetic LAB Diligence benchmark, moving from a standard tool loop to the RLM harness raised mean rubric pass rate from 23% to 62% across models @harvey, @nikogrupen.
- Post-training inside the harness mattered at least as much as the harness itself. Harvey reports self-distilled SFT on GLM-5.2 improved pass rate 46% → 60%, while GRPO on Qwen3.5-122B-A10B lifted pass rate 30% → 63% on held-out rooms and improved document coverage 62% → 96% @harvey. The broader implication, echoed by others, is that agent benchmarks increasingly need to treat orchestration and post-training as part of the model system, not external glue.
- LangChain/deepagents shipped quality-of-life primitives for harness design, including subagent forking that passes supervisor context down to subagents, plus managed connections to abstract OAuth/token/consent flows for either agent-owned or user-owned identities @colifran_, @hwchase17, @caspar_br. This is a useful sign of the stack maturing around long-horizon agent workloads.
Inference and Systems: Sparse Attention, Agentic Serving, and Decode Megakernels
- vLLM’s long-context serving work is notable. The project described Hybrid HiSparse for sparse-MLA models: KV stays on GPU while possible, then cold KV pages are offloaded to host memory, while a hot buffer serves the indexer. On GLM 5.3 with 1M context on an 8×H200 node, configured concurrency 32, plain offloading sustained 5–6 requests while Hybrid HiSparse sustained 19–25 @vllm_project. This matters directly for RL rollouts and long-context concurrency, where VRAM-bound decode otherwise kills throughput.
- vLLM also published a full-stack optimization pass for real-world agent traffic, benchmarked on AgentX. Key takeaways: pipeline parallelism helps cold long prompts but loses on warm short turns; decode context parallelism depends strongly on the model’s attention stack; and session-sticky routing can beat naive load balancing because warm KV caches matter more than even queue distribution in fast-turn agent settings @vllm_project.
- Cohere introduced an open-source serving stack built around a “decode megakernel,” claiming up to 1.58× faster performance than vLLM on North Mini Code and 1.25×–1.41× end-to-end gains at higher batch sizes @cohere. Combined with Baseten’s note that frontier RL rollouts now get new policy weights live in under 40 seconds globally with only a 6-second pause @baseten, the clear trend is toward infra specialized for continuous post-training and rollout refresh, not static model serving.
Top Tweets (by engagement)
- Anthropic resignation / safety warning: Jacob Hilton resigned from Anthropic, arguing both Anthropic and OpenAI are racing toward self-improving superintelligence irresponsibly and that insiders privately treat extinction risk as real @hilbertspaess, with follow-up claims that current systems could soon hack infrastructure and transform fields rapidly @hilbertspaess.
- OpenAI’s user-data clarification: OpenAI’s formal statement that no specific user data was accessed for Navier–Stokes, alongside the caveat about possible de-identified derivative improvement, became a major flashpoint @OpenAI.
- Cognition financing: Cognition announced a raise of $2B+ at a $48B valuation, saying run-rate revenue grew from $492M to nearly $900M since May @cognition.
- Meta Muse launch: Mark Zuckerberg’s launch post for Muse was among the highest-engagement product tweets of the day @finkd.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Chinese Multimodal AI Releases: Driving and Flash APIs
-
Qwen/Qwen-Drive-1.0-4B · Hugging Face (Activity: 549): Qwen released
Qwen/Qwen-Drive-1.0-4B, an open-weight autonomous-driving VLM derived from an unchanged Qwen3.5 4B VLM, with a full BF16 checkpoint around9Band extraplanner-sft,planner-rl, andperceptionmodules. Per the linked technical report, Qwen-Drive-1.0 adds an external BEV perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, plus a Planning Expert for future ego-trajectory generation, trained via staged mixtures of driving supervision and general VLM data. The release reports competitive performance across WOD-E2E, NAVSIM, driving VQA, and open-/pseudo-closed-/closed-loop planning evaluations while largely preserving general multimodal capability. -
DeepSeek Flash 4.1 is already being tested via API and rolling out. (Activity: 528): DeepSeek V4.1 Flash is reportedly in internal beta via API: keep the existing
base_urland call modeldeepseek-v4.1-flash-expires-on-0910, with pricing unchanged fromdeepseek-v4-flashand a20concurrent request/account limit (source). The translated announcement claims a “new model architecture” with native multimodal support, stronger capability, faster throughput, and lower cost; commenters report roughly2.24×speedup and up to~30%better token efficiency in benchmarks, though one edit speculates the observed speed gain may be partly due to lower beta concurrency rather than architecture alone. Comment sentiment is strongly positive toward DeepSeek/open-weight progress, but the only substantive debate is whether the claimed performance improvement reflects a genuinely new architecture or simply lighter API load during beta testing.- Users report that DeepSeek Flash 4.1 appears to be around
2.24xfaster via API testing, with some speculation that the observed speedup may come from lower concurrent load rather than a fundamentally new architecture. Other comments suggest it may be multimodal, though this is not yet confirmed in the thread. - One technically relevant claim is that some users are seeing up to
30%better token efficiency in benchmarks, which could explain DeepSeek’s reported “lower costs” messaging if fewer tokens are needed for comparable outputs. The comment frames this as benchmark-dependent and not yet independently validated. - There is some discussion of release cadence and migration complexity: users mention not having fully moved from the 0731 model to the newer vision variant before another release appears imminent. This highlights a practical API-integration issue where fast model iteration can outpace downstream evaluation, regression testing, and deployment workflows.
- Users report that DeepSeek Flash 4.1 appears to be around