Agents go persistent, lab AI advances, benchmarks under review
Agents dig deeper into processes, and evaluation methods follow.
The week in five points
-
Deepmind expands Co-Scientist. The AI plans experiments, operates lab gear, and drafts papers. So far there is only the vendor demo. An independent check is missing. If you work in a lab, set clear roles, approvals, and logfiles.1
-
Deepmind tests a double-blind assessment for frontier models. A double-blind study where neither evaluators nor developers know what is being tested. The goal is sturdier results without leaking test sets. This is a test, not a standard. If you build benchmarks, hide model names and mix tasks systematically.2
-
OpenAI is working on a Persistent Mode for its Codex agent. A persistent agent state with self-start and memory. The feature is in progress and not available. If you plan agents, expect background runs, budgets, escalations, and audit trails.3
-
A US federal court declares the Pentagon blacklisting of Anthropic unlawful. That weakens the basis for political interference in procurements. If you work for US agencies, review contracts and risk clauses.4
-
An Anthropic researcher shows a look into self-improving safety systems. Given ten benchmarks for misbehavior, scores rose on all without drops elsewhere. This is a single finding without peer review. If you build safety evals, expect metric chasing and vary your tests.5
Sources
- Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers (the-decoder.com)
- AI benchmarks have a trust problem and Google wants to fix it (the-decoder.com)
- Always-on and self-starting AI agents might be OpenAI's next big play (the-decoder.com)
- U.S. court rules Pentagon's blacklisting of Anthropic was unlawful (the-decoder.com)
- An Anthropic researcher just gave us a peek at self-improving AI (techcrunch.com)