Claude brings bigger jobs into one conversation
Anthropic is merging Cowork into Claude chat, so you can ask a question or delegate a report from the same conversation. The rollout starts with Pro and Max across web, desktop and mobile over the coming weeks; other plans follow later.
New Docs and Slides let you edit Claude's output directly, present from the app, or export presentations as PowerPoint or PDF. These and the integrated Design tools are in beta on paid plans, with Enterprise activation controlled by admins.
Developers get redesigned Claude Code Projects: a coordinating agent divides a goal into parallel work sessions, shares project memory, reviews results and assembles the output. Separate copies of the code reduce interference, but overlapping edits can still conflict.
The Projects beta initially covers selected Pro and Max cloud-session users without existing web or desktop projects. Local execution is still coming, and parallel sessions consume your allowance faster.
Start with one recurring report or a coding task with separable parts. Keep human review before acting on the result or merging changes.
Voice AI gets better at talking and listening
Google's Gemini 3.8 Live models can keep a conversation going while calling tools in the background. A support agent can check an order without leaving the caller waiting in silence. The Extended Thinking version adds deeper reasoning for multi-step requests.
They're available through the Live API and AI Studio. Google's estimated audio rates are $0.005 per input minute and $0.018 per output minute: separate listening and speaking charges, not one all-in call price.
For turning recordings into text, Grok Voice Transcribe 2.0 keeps pricing at $0.10 per audio hour for batch processing and $0.20 for streaming. Speaker labels, timestamps and vocabulary hints are included.
xAI reports better recognition of noisy calls, multilingual speech and spoken identifiers. Its "twice as accurate" headline comes from internal evaluations. The launch also highlights Atlassian's use of Grok for Loom transcription.
These solve different jobs: Gemini handles a live conversation; Grok produces its written record. Test accents, names and account numbers before switching. Select Grok's version 2.0 explicitly during its rollout.
Quick hits 🗞️
-
Qwen adds long audio and video understanding. Qwen3.8-Omni-Flash offers a 1M-token context, text/image/audio/video input and tool calling through Alibaba Cloud. Output is text only; this is a hosted API, not an open-weight release. Source
-
An AI-connected account widened a security failure. Hacktron disclosed a forum vulnerability and login flaw that researchers say let them access OpenAI employee accounts and demonstrate internal-repository access through Codex. They report the OpenAI-side flaw was fixed in July; this is a new disclosure, not a new ongoing breach. Source
-
Figure tests robots in unfamiliar homes. Helix 2.5 tackled tidying, towel folding and bed making across 30 unseen homes. Figure reports 56% full-task success with its pretraining versus 9% without. Promising transfer, but still a company-run test of three trained behaviors. Source
-
A smaller local model: PrismML released Apache-2.0 Bonsai 2 27B. Its smallest language-model file is 5.95 GB; vision and working memory add overhead. It needs PrismML's compatible runtime, and performance-retention figures are vendor claims. Source
-
Google discloses an AI test crossing boundaries. Google told NBC that Gemini accessed three outside systems during May testing, mistaking them for test targets. Google says it stopped and caused no known damage. That account does not establish deliberate malicious intent. Source
See you next week!
If this helped you catch up quickly with AI, forward it to someone who would find it useful. They can subscribe here.