Field Notes

Keep costs in check with hard caps, use faster Claude, lock down agents cleanly

Today is about hard spend limits in the cloud, a faster Claude, and practical safeguards for agents.

Important AI news

Hard spend limits for cloud and agents

Simon Willison calls for default hard budget caps. This means a limit with automatic stop for usage-based services.1 The service stops after you cross the limit and returns errors instead of continuing to bill.1 On September 16, AWS announced a new experience where you can set a monthly spend limit per project that pauses operations when reached.1 According to the AWS page this is currently available only to some customers.1 Google Cloud has offered spend caps per service within a project since July.1 Concrete takeaway: set hard limits for all agent and API workloads and allow opt-out only with explicit approval.1

Claude Sonnet 5.5: faster, cheaper, more widely available

Anthropic released Sonnet 5.5.2 The provider promises over 30 percent more speed and up to 30 percent lower costs for most tasks at the same list price.2 Simon Willison reports that Sonnet 5.5 powers the free tier on claude.ai and comes close to Opus 5.5 on code tasks.2 This is a chance to cover standard tasks like short analyses or small code changes at lower cost.2 Timing note: the post appeared six days ago.2

Agents organize themselves, and they get louder

Ethan Mollick revises his view on agents and says that modern models plan their steps more and more on their own.3 He points to personal agents like Meta Muse and OpenAI dots that can send proactive messages and even place calls.3 He names Clawlikes. These are agents with access to computers and accounts that act autonomously.3 Mollick also describes an OpenAI swarm coordination. This means many agents that support each other. He stresses that the claimed math solution is not yet formally recognized.3 What this means for you: expect more agent traffic in mail, chat, and phone and set rate limits, identity checks, and escalation rules.3

Highlights for your workday

  • A new study shows this: on medium-length, well-defined accounting tasks, top models were faster and more accurate than junior clerks.4 The claim comes from Ethan Mollick’s summary on X. Independent replications are not provided there.4 If you review receipts or compute provisions, test a model in parallel on a sample set and compare hit rate and corrections.4

  • Fabric IQ is now generally available in Copilot.5 You can ask questions of your company data and get answers in work context.5 This pays off for teams with Microsoft 365 Copilot and Fabric that need ad hoc analysis without a BI ticket.5

  • Microsoft introduces new models for faster transcription and multilingual voices.6 Useful for customer calls, support, or subtitles when you must cover many languages.6 Before rollout, check which languages and channels your environment supports.6

Try it in five minutes

Draft a hard budget cap as a team rule, usable with any AI chat.1

1) Copy this sample data into your chat and ask for a monthly cap per service with a 20 percent buffer:

Service, Cost per call, Expected calls/month
Text API, $0.002, 300000
Image API, $0.02, 15000
Vector search, $0.0005, 800000

Ask for a total and the cost at cap reached.1 2) Ask for a hard stop rule in one sentence and a short error message for users after the cap is reached.1 3) Ask for an IT ticket text: steps to implement, ownership, monitoring, and an approval process for temporary unlock.1

Expected result: cap per service, total cost, a clear stop rule, and a user-friendly error message plus ticket text.1 Check criterion: do the sums add up, does the rule include an explicit stop action, does the message avoid false promises.1

Sources

  1. We're going to need default hard budget caps on pretty much everything (simonwillison.net)
  2. Claude Sonnet 5.5 (simonwillison.net)
  3. The Dot and the Swarm (oneusefulthing.org)
  4. @emollick: “We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.” Eightee (x.com)
  5. AI is only as useful as the context it can reason over. With Fabric IQ now generally available in Copilot, you can ask questions of your enterprise data and get answers grounded in trusted business context, right where you work. (linkedin.com)
  6. New Microsoft AI models bring faster transcription and multilingual voices (microsoft.ai)