The Weekly AI Recap

This week in AI: Open models close the gap

July 19, 2026 · 4 min read

Good morning,

Kimi and Thinking Machines pushed open models forward, Google turned NotebookLM into a more capable research tool, and several security incidents showed why autonomous systems still need tight boundaries.

Open models close the gap

Chinese AI lab Moonshot introduced Kimi K3, a 2.8-trillion-parameter model for coding, knowledge work, and reasoning.

  • Capabilities: Text, image, and video input; one-million-token context; 16 of 896 experts active per token.
  • Availability: Live in Kimi and its API. The weights are not public yet; Moonshot says they will arrive by July 27.
  • Price and performance: $3 per million input tokens and $15 per million output tokens, with cached input at $0.30 and no long-context surcharge. Artificial Analysis scored K3 at 57, behind Claude Fable 5 and GPT-5.6 Sol overall. A preliminary result currently puts it first on Arena's WebDev leaderboard.

Mira Murati's Thinking Machines Lab also released Inkling, its first production model. Inkling has 975 billion total parameters, accepts text, image, and audio, supports up to one million tokens when self-hosted, and ships under Apache 2.0.

The releases serve different needs. Kimi is chasing frontier performance, while Inkling is designed as a customizable base. For teams with valuable internal data and repeatable workflows, owning and fine-tuning a good-enough model may matter more than renting the smartest one.

NotebookLM becomes a more capable research tool

Google renamed NotebookLM to Gemini Notebook and gave each notebook access to a secure cloud computer.

It can now write and execute code, perform deeper data analysis, and create new outputs while staying grounded in the sources a user provides. The feature is available first to Google AI Ultra users and eligible Workspace business customers, with Pro access rolling out on the web over the coming weeks.

Notebooks now sync with the Gemini app and will later appear inside AI Mode in Google Search. Google says the product has reached more than 30 million users and 600,000 organizations.

The rename is forgettable. The execution layer is not. Adding code turns a bounded collection of sources into working material for research, analysis, and reporting without discarding the grounding that made NotebookLM useful.

Stronger agents make basic safeguards non-negotiable

Reports surfaced this week of Codex sessions using GPT-5.6 Sol deleting files while running with broad computer access. OpenAI's Codex team later said the handful of confirmed cases involved full-access mode without sandboxing or automatic review. One involved Codex redefining $HOME as a temporary folder and deleting the user's home directory.

Hugging Face also disclosed an intrusion that it says was operated end to end by an autonomous agent framework. A malicious dataset exploited two code-execution paths, after which the attacker harvested credentials and moved across internal clusters. Hugging Face found no evidence that public models, datasets, Spaces, or published packages were altered, but its assessment of partner and customer data is still ongoing.

OpenAI is automating the defense too. Its internal GPT-Red system generates prompt-injection attacks for training production models. GPT-Red found successful attacks in 84% of test scenarios, compared with 13% for human red-teamers. OpenAI says this training gave GPT-5.6 Sol six times fewer failures on its hardest direct prompt-injection benchmark than its best production model four months earlier.

The practical lesson is dull and essential: sandbox agents, scope credentials tightly, require approval for destructive actions, and keep version control and backups. A capable agent with unrestricted access is still unrestricted software making probabilistic decisions.

Quick hits 🗞️

  • GPT-5.6 Terra may be the awkward middle child. Artificial Analysis found that Luna or Sol matched or beat every Terra configuration on intelligence versus cost. Developers choosing purely on value may be better served by skipping the middle tier for now. Source
  • Google Search can now act inside third-party apps. U.S. users can connect Instacart, Canva, and YouTube Music to AI Mode, then add groceries to a cart, find design templates, or save generated playlists without leaving the search workflow. Source
  • Anthropic and Blackstone named their AI implementation company. Ode with Anthropic is a $1.5 billion joint venture with roughly 100 engineers helping enterprises redesign workflows around AI. Its existence is a useful tell: access to capable models is getting easier, while implementation remains expensive. Source
  • Apple Intelligence received approval to launch in China. Apple will integrate Alibaba's Qwen models into its operating systems and is also working with Baidu on features for Chinese users. No launch date has been announced. Source
  • New York paused some new state permits for hyperscale data centers. The order blocks discretionary environmental permits not already deemed complete for up to one year while New York writes rules covering grid, water, air, and ratepayer impacts. Source

See you next week!

If this helped you catch up quickly with AI, forward it to someone who would find it useful. They can subscribe here.

The Weekly AI Recap

Get the next issue in your inbox

Every Sunday we send the model releases, industry shifts, and tools that actually mattered this week. One email, five minutes, free.

Free. One email every Sunday. No spam, unsubscribe anytime.