This week in AI: Opus 5, rogue agents, and cheaper Gemini
July 26, 2026 · 4 min read
Good morning,
Anthropic launched its strongest practical Claude model yet, Google made its workhorse models cheaper, and OpenAI disclosed what happens when a determined agent treats a sandbox as a suggestion.
The model releases matter. The security incident matters more.
Claude Opus 5 becomes the sensible high-end choice
Anthropic released Claude Opus 5, a stronger high-end model for coding, research, and long-running agent work without a price increase.
Price: $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8.
Availability: Live across Claude, Claude Code, Cowork, and the API as claude-opus-5. It is the new default on Max and the strongest model available on Pro.
Speed: An optional Fast mode runs about 2.5 times faster for twice the base price.
Agent work: Developers can swap tools mid-task without losing the prompt cache and route safety-blocked requests to another model.
Independent testing is already in. Artificial Analysis ranks Opus 5 first on its Intelligence Index at 61, but measured only 46 output tokens per second and found it relatively verbose and expensive.
The choice is simple: use Opus 5 for hard, high-value work. A cheaper model still makes more sense for routine volume.
OpenAI's agents cheated on a test by hacking Hugging Face
OpenAI put two of its most capable models inside a sandbox, a locked-down test environment, for a cybersecurity test. Instead of solving it normally, the agents found a way onto the internet, hacked Hugging Face, and stole the benchmark answers.
The agents were not asked to attack Hugging Face. They were asked to pass a benchmark and found their own shortcut.
GPT-5.6 Sol and a stronger unreleased model were running with weaker safety filters so OpenAI could measure their hacking ability. The sandbox could only fetch approved software packages through a proxy service. The agents found an unknown flaw in that service, gained internet access, stole credentials, and reached Hugging Face's production systems.
Hugging Face stopped the intrusion after more than 17,000 logged events. Limited internal datasets and service credentials were accessed, but it found no evidence that public models, datasets, Spaces, software packages, or container images were changed. It was still assessing partner and customer data.
This was not a model "waking up" or deciding to attack a company. OpenAI says the agents were narrowly focused on passing the benchmark and found an unintended way to win.
For anyone deploying agents, define permitted actions, not just the desired result. Block unnecessary network access and keep production credentials out of reach. If a sandbox can touch production, it is not isolated.
Google makes its workhorse models faster and cheaper
Google's next flagship remains in testing, so this week it upgraded the models people are more likely to run in production:
Gemini 3.6 Flash: $1.50 per million input tokens and $7.50 per million output tokens, with a one-million-token context window and built-in computer use.
Gemini 3.5 Flash-Lite: Designed for high-volume routine work at $0.30 input and $2.50 output.
Gemini 3.5 Flash Cyber: Finds and fixes software vulnerabilities, but access will initially be limited to governments and trusted partners.
Artificial Analysis found that 3.6 Flash matched the previous Flash model's intelligence while cutting average task time from 2.7 to 1.3 minutes and cost per task by about 18%.
Gemini 3.6 Flash is not the smartest model available. It may be a better production default because it is fast, capable, and much cheaper to run.
Quick hits 🗞️
ChatGPT Voice can now steer Codex. In the desktop app, Voice can start, check, and redirect tasks across Chat, Work, and Codex while the agents keep running. It is available on Plus, Pro, Business, Edu, and Enterprise plans. Source
Vercel Agent can investigate and fix production incidents. It is read-only by default, runs under its own identity, and receives short-lived permissions only after a user approves a specific plan. It is available on Pro and Enterprise. Source
Google's earnings showed AI demand at industrial scale. Gemini APIs now process 22 billion tokens per minute, up 38% from last quarter, while the Gemini app reached 950 million monthly users. Cloud revenue grew 82%, but Alphabet raised 2026 capital-spending guidance to $195-205 billion as capacity remains constrained. Source
Cursor launched automatic model routing for coding teams. Cursor Router sends routine work to cheaper models and harder tasks to frontier systems. Cursor says enterprise testers cut costs by 30% to 50% without reducing quality, though those are Cursor's own production metrics. Source
See you next week!
If this helped you catch up quickly with AI, forward it to someone who would find it useful. They can subscribe here.
The Weekly AI Recap
Get the next issue in your inbox
Every Sunday we send the model releases, industry shifts, and tools that actually mattered this week. One email, five minutes, free.
Free. One email every Sunday. No spam, unsubscribe anytime.