Field Notes

OpenAI shows Jalapeño chip, Google launches Gemini Legal, Ukraine opens war dataset

Inference hardware, AI for law, and military data move forward, but solid comparisons often miss.

The key stories

OpenAI showed Jalapeño, its own inference chip. Internal benchmarks are said to beat Nvidia Blackwell and Rubin. The numbers come from the vendor, independent tests are missing. Nothing changes for you until hardware and services are available.1

OpenAI says Jalapeño delivers answers faster and more efficiently than rivals. Hardware lead Richard Ho says it is the “best of both worlds”. These are vendor claims without external confirmation.2

OpenAI lists first results: leading speed and efficiency in inference. Promised are higher throughput and lower latency for modern models. Without independent benchmarks this remains a performance promise.3

Ukraine opens a large labeled combat data set to UK companies. It is part of an AI weapons partnership with the UK government. If you work in defence in the UK, you can start an access request.4

Google launches Gemini Enterprise for Legal to automate contracts and research. The details come from Google’s announcement, solid real-world results are missing. If you work in a legal department, plan a tightly scoped pilot with real documents.5

Nvidia positions the Groq 3 LPX inference chip as four times faster than Cerebras. The comparison math is complex, and the numbers come from Nvidia. Without standardized benchmarks the comparison is hard.6

Tools and releases

Anthropic gives Claude a shared memory between Chat and Cowork, so you do not have to repeat projects and preferences.7

llama.cpp 0.3.0 brings the multimodal dots3-note, a new KV cache, MTP support for GLM-4.5-Air, and ggml 0.22.8

CUDA Python 1.0 goes stable, unifies the APIs, and opens full CUDA access from Python.9

Keenable comes out of stealth, raises 26 million dollars seed, and builds a web index for AI agents.10

Research

Quantization-Aware Healing shows that a compressed 4-bit model can beat its full-precision original. This is thanks to quantization. That means the deliberate reduction of numerical resolution during training.11

Sources

  1. OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks (the-decoder.com)
  2. OpenAI says its Jalapeño chip can power faster AI responses than the competition (theverge.com)
  3. Jalapeño’s first results show industry-leading speed and efficiency in AI inference (openai.com)
  4. Ukraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnership (the-decoder.com)
  5. Google launches Gemini for legal work to automate contracts and research (the-decoder.com)
  6. Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated (the-decoder.com)
  7. Claude Cowork finally remembers what you told the app in chat (techcrunch.com)
  8. llama.cpp v0.3.0 (github.com)
  9. CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access (developer.nvidia.com)
  10. Accel-backed Keenable is indexing the web for AI agents (techcrunch.com)
  11. Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original (huggingface.co)