Field Notes

Sol drives revenue, Mythos 5 scans, prices rise, agents replicate research

Frontier models move into production, hardware gets pricier, and agents inch closer to real work.

The key stories

OpenAI reports a clear revenue jump since launching GPT-5.6 Sol in early July. The figure comes from OpenAI and landed three days ago. If you plan integrations, test Sol against your existing models.1

Anthropic now uses its strongest model Claude Mythos 5 for cyber defense. The security scanner is reportedly running in production. The switch was reported three days ago. If you evaluate security scans, try Mythos-backed checks on your codebase.2

Inherent, founded by DeepMind alumni, introduces Faraday, an agent for research replication. The team says Faraday beat OpenAI and Anthropic at replicating papers. These are vendor tests. Independent benchmarks are missing. If you want to re-create paper results, keep an eye on Faraday.3

A memory shortage is pushing prices for Nvidia AI servers with Vera Rubin and Grace Blackwell up by about 15 percent, according to reports. That can delay procurement and shift budgets. If you plan capacity for 2026/27, expect higher unit prices.4

A new study finds few publicly documented plans from leading labs to contain rogue models. That points to gaps in operational readiness. If you work with frontier models, demand concrete emergency and shutdown plans.5

Tools and releases

OpenAI shows a preview of transparent backgrounds in GPT-Image-2. The feature is not generally available.6

The Agentic Data Operations Platform on Bedrock sketches agents that automate the bronze silver gold data pipeline cycle, with end-to-end governance.7

A pattern for query-aware context compression on Bedrock filters retrieval chunks with a smaller model before the main model answers. That lowers input costs.8

SOP-Bench brings an extensible agent benchmark for real business procedures instead of isolated proxy tasks.9

AgentX releases InferenceXv3 and a million-scale dataset, with 1M+ context length and 95%+ KV cache hit rate, an intermediate buffer to speed up decoding.10

NVIDIA DSX MaxLPS targets more performance per watt in AI factories under tight power limits.11

Research

Princeton and UC San Diego show that agents benefit from trained skills. They fail early on long tasks when they lack the right capabilities.12

New work shows that world models, meaning internal representations of the environment, predict wrong actions when they ignore human beliefs.13

Sources

  1. GPT-5.6 Sol drives OpenAI's revenue surge as it regains ground on Anthropic (the-decoder.com)
  2. Anthropic puts its most powerful model Claude Mythos 5 to work for cyber defense (the-decoder.com)
  3. Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research (techcrunch.com)
  4. Memory shortage reportedly drives Nvidia AI server prices up about 15 percent (the-decoder.com)
  5. Frontier AI labs still won’t say how they’d contain a rogue model (techcrunch.com)
  6. OpenAI's GPT-Image-2 can now generate images without a background (the-decoder.com)
  7. Agentic Data Operations Platform (ADOP): Data engineering into hours (aws.amazon.com)
  8. Reduce RAG costs on Amazon Bedrock with query-aware compression (aws.amazon.com)
  9. SOP-Bench: A new benchmark for evaluating AI agents on real business procedures (amazon.science)
  10. AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? (newsletter.semianalysis.com)
  11. Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS (developer.nvidia.com)
  12. Study explains why AI agents benefit from "skills" and when they fail (the-decoder.com)
  13. World models that ignore human beliefs predict the wrong actions, new research shows (the-decoder.com)