Astra flies a drone, Nvidia weighs 10B for Anthropic, oversight gains tailwind
Agents get more practical, capital and governance follow, and new models hit search and music.
The key stories
Astra flies a surveillance drone and runs a small business on its own. It earns nearly three times Claude Fable 5.1. The figures come from the vendor, there are no independent audits. You need clear approvals and telemetry before you ship this live.1
Early benchmarks point to a jump in spatial reasoning for GPT-6 Astra. That means thinking about position and movement in space. The model improves sharply on a new robotics benchmark. Independent confirmations are missing here too. You can test planning tasks with less supervision once you get access.2
Nvidia is negotiating up to 10 billion dollars for Anthropic’s planned IPO. These are talks, not a deal. That would cement the compute alliance between the two. You could book more capacity on Claude, price pressure is unclear.3
Sam Altman, Elon Musk, and Demis Hassabis back Dario Amodei’s call for independent oversight. The focus is recursive self-improvement. Models that improve further without external help. This shifts the debate from voluntary pledges to mandatory review. If you work on frontier features, expect audits.4
Google introduces a new model that forecasts future sales from sales data, weather, and discount plans. The announcement gives no numbers on accuracy or availability. If you do forecasting, plan a comparison against your baseline once access is possible.5
Iris-mini and Iris-pro lead their classes as open-weight search agents. The ranking is based on the reported benchmarks. You can test and tune them locally or in your own cloud.6
ElevenLabs ships Music v2.5 in app and API, with Free and Pro tiers. The version is available. You can sketch ideas and hook the API into your workflow.7
OpenAI stays its course and claims a solution to a Millennium Prize Problem this week. That would be big, but external validation is pending. The piece frames this as flag-planting politics. Win instead of wait.8
Tools and releases
- Perplexity uses Astra for communications, code changes, and production monitoring, with far fewer manual interventions.9
- An open benchmarking harness for OpenAI models on Bedrock measures cost per correct answer, agent trajectory cost, and quality grades. It came out three days ago.10
- Moonshot AI targets 2 billion dollars in annual revenue. Three days ago OpenRouter reported up to 300 billion K3 tokens daily. Both without independent verification.11
Research
- Written reasoning steps line up with distinct internal activity patterns. That hints at new levers for analysis and control.12
- A two-hour university study finds: students do worse when AI is banned in class, compared to guided use.13
Sources
- GPT-6 Astra pilots a surveillance drone and runs a business on its own (the-decoder.com)
- GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks (the-decoder.com)
- Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO (the-decoder.com)
- Altman, Musk, and Hassabis back Amodei's call to add independent oversight (the-decoder.com)
- Google's new AI model predicts the future from sales data, weather, and discount schedules (the-decoder.com)
- Iris-mini and Iris-pro are the strongest open-weight search agents in their class (the-decoder.com)
- Elevenlabs makes Music v2.5 available via app and API with free and pro tier options (the-decoder.com)
- OpenAI just wants to win (theverge.com)
- Perplexity trusts GPT-6 Astra with end-to-end systems (openai.com)
- Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload (aws.amazon.com)
- Kimi-maker Moonshot AI targets $2B in annual revenue (techcrunch.com)
- AI models' written reasoning steps correspond to distinct internal patterns, a new study finds (the-decoder.com)
- Two-year university study finds banning AI from classrooms leaves students worse off (the-decoder.com)