Field Notes

Agents go persistent, lab AI advances, benchmarks under review

Agents dig deeper into processes, and evaluation methods follow.

The week in five points

  1. Deepmind expands Co-Scientist. The AI plans experiments, operates lab gear, and drafts papers. So far there is only the vendor demo. An independent check is missing. If you work in a lab, set clear roles, approvals, and logfiles.1

  2. Deepmind tests a double-blind assessment for frontier models. A double-blind study where neither evaluators nor developers know what is being tested. The goal is sturdier results without leaking test sets. This is a test, not a standard. If you build benchmarks, hide model names and mix tasks systematically.2

  3. OpenAI is working on a Persistent Mode for its Codex agent. A persistent agent state with self-start and memory. The feature is in progress and not available. If you plan agents, expect background runs, budgets, escalations, and audit trails.3

  4. A US federal court declares the Pentagon blacklisting of Anthropic unlawful. That weakens the basis for political interference in procurements. If you work for US agencies, review contracts and risk clauses.4

  5. An Anthropic researcher shows a look into self-improving safety systems. Given ten benchmarks for misbehavior, scores rose on all without drops elsewhere. This is a single finding without peer review. If you build safety evals, expect metric chasing and vary your tests.5

Sources

  1. Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers (the-decoder.com)
  2. AI benchmarks have a trust problem and Google wants to fix it (the-decoder.com)
  3. Always-on and self-starting AI agents might be OpenAI's next big play (the-decoder.com)
  4. U.S. court rules Pentagon's blacklisting of Anthropic was unlawful (the-decoder.com)
  5. An Anthropic researcher just gave us a peek at self-improving AI (techcrunch.com)