Field Notes

OpenAI brings full-duplex voice, GPT-6 Astra leads in math, Deepseek saves agent RAM

Voice agents go real time, robots learn from one video, and math benchmarks shift.

The key headlines

OpenAI ships GPT-Live-1, a developer API for real-time voice dialogs. It supports full duplex. That means talking and listening at the same time. You build voice agents without push-to-talk and with natural interruption. The launch targets developers, not end users.1

Skild AI announces the robotics foundation model S1, built on Nvidia’s Physical AI. It is supposed to learn new, long task chains from a single video. Long-horizon means many steps over longer periods. Independent tests are missing, the claims come from the vendor. If that holds, you save data and time on changeovers.2

GPT-6 Astra tops ErdosBench for open math problems. OpenAI says the model should take speed out of research so humans can check and prove. The benchmark claims come from the vendor. If you work in math, use it as an idea source. The proofs stay with you.3

Deepseek presents V4.1-Flash, a multimodal model with 552 billion parameters. It targets agent scenarios and is supposed to cut memory use sharply. How large the advantage is remains unclear without neutral measurement. If the effect holds, you get more concurrent agents on the same hardware.4

Pocket FM reports a doubled run rate, the annualized revenue, to 500 million dollars. According to the company, 93 percent of audio hours come from AI, 99 percent of new content as well. Production is about 80 times cheaper. All numbers are vendor claims. If you plan audio formats, expect more AI ware in your users’ feeds.5

Tools and releases

Amazon SageMaker Inference introduces prefix-aware routing. It directs identical prompt starts to the same instance and, per benchmarks, cuts P50 time-to-first-token by up to 77 percent.6 TwelveLabs Marengo Embed 3.0 is generally available in Bedrock Knowledge Bases and enables managed semantic search in video, image, and audio.7 Universal Music announces with ElevenLabs an AI platform for remixes and mashups from licensed catalogs. Availability is still pending.8 Nvidia and Palantir link their platforms for AI-driven supply chains, starting with Nvidia’s own million-part operation.9 A configurable PII detector on Bedrock places the entities to detect into the prompt and beats a standard tool on five datasets without retraining.10 Meta launches Muse as a WhatsApp agent. It handles orders, emails, and price negotiations, all in chat.11

Research

The Agent Evaluation Metric breaks down multi-turn interactions at the turn level and marks the step that triggers the error trajectory.12 New results suggest that research agents learn compressible data models. That leaves little room for mere memorization.13

Sources

  1. OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time (the-decoder.com)
  2. Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video (blogs.nvidia.com)
  3. GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design (the-decoder.com)
  4. New Deepseek model V4.1-Flash cuts memory needs for AI agents (the-decoder.com)
  5. India’s Pocket FM doubles revenue run rate to $500M as AI powers 93% of audio content (techcrunch.com)
  6. Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference (aws.amazon.com)
  7. Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0 (aws.amazon.com)
  8. Universal Music is launching an AI music platform with ElevenLabs (theverge.com)
  9. Nvidia and Palantir team up to run supply chains with AI, starting with Nvidia's own million-part operation (the-decoder.com)
  10. Model-agnostic PII detection with LLMs (aws.amazon.com)
  11. Muse can shop, write emails, and negotiate prices for users, all through WhatsApp (the-decoder.com)
  12. Agent Evaluation Metric for multi-turn conversations (aws.amazon.com)
  13. Why don’t machine learning research agents overfit? (amazon.science)