Field Notes

Agents API and Live Voice launch, OpenAI scales storage, policy brake, new vision paper

New building blocks for agents and voice, very large infrastructure numbers, policy steps in, research shows a compact vision model.

The week in five points

  1. OpenAI opens the Agents API in a public beta. It is meant to give developers access to the infrastructure behind Codex and ChatGPT. If you build agents, you can offload backend work with it.1

  2. GPT-Live-1 brings natural full-duplex conversations to the API. It adds stronger instruction following, custom voices, and a telephony bridge. You can build real hotline flows or live assistants with it.2

  3. OpenAI describes the evolution from Habitat to a globally distributed storage platform. It serves, according to OpenAI, over 1 billion ChatGPT accounts and 22 million requests per second. The numbers come from the vendor.3

  4. OpenAI floats the idea of a joint AI slowdown in the US Congress. This is a sounding, not a decision. If your roadmap depends on fast model upgrades, plan for uncertainty.4

  5. Three days ago the paper on SenseNova-U1.5 appeared, an 8B-MoT multimodal model. It understands, reasons, and generates visual content in an encoder-free and VAE-free architecture. They name spatially coherent patch reconstruction to strengthen the visual interface. Independent checks are pending.5

Sources

  1. OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT (the-decoder.com)
  2. Build more natural voice experiences with GPT‑Live‑1 in the API (openai.com)
  3. Rapidly scaling online storage to serve over 1 billion ChatGPT users (openai.com)
  4. OpenAI floats a shared AI slowdown, takes it to Congress (the-decoder.com)
  5. Paper: SenseNova-U1.5: Towards Native Unified Visual Intelligence (huggingface.co)