OpenAI speeds up GPT-5.6, Meta opens Glimmer, Google cuts Gemini price, Qwen 3.8 launches
Speed, open weights, and security are shifting real-world practice with agents and code this week.
The week in five points
-
OpenAI introduces Ultrafast, a new inference mode—i.e., a faster execution mode for models. With it, GPT-5.6 runs up to 14× faster. The preview delivers up to 750 output tokens per second. The service uses Cerebras hardware and is available as an API tier.
-
OpenAI releases GPT-5.6-Cyber for defenders. The model is intended for vulnerability research and exploit validation. Usage runs via Daybreak Red and is intended only for authorized testing. You can structure workflows for security testing this way.
-
Meta releases Muse Glimmer with open weights. The dense 30B multimodal model targets local agents on consumer hardware. Quantization reduces the memory footprint below 20 GB. The license is Apache 2.0, a permissive license, and there is Unsloth support.
-
Google ships Gemini 3.7 Flash three weeks after 3.6. The focus is on coding, web development, knowledge work, and agentic workflows. The introductory price drops by 50 percent. Benchmarks show DeepSWE 65.3 percent and Code Arena Elo 1588.
-
Alibaba provides Qwen 3.8 with open weights under Apache 2.0. Qwen3.8-2.4T uses an MoE design—i.e., a mixture of experts—with 95B active parameters and can be served on NVIDIA GB300 NVL72. For local setups, Qwen3.8-27B with Unsloth works from 17 GB RAM. This covers both cluster and laptop scenarios.