Happy Sunday, !
Welcome back to your weekly AI news roundup.
In case you missed it, here’s this week’s Thursday post:
Flashcards Are Good. “Gap” Cards Are Better.
TL;DR: Chatbots can identify gaps in your knowledge, so have them tailor flashcards to your blind spots.
Note: If you’re consistently missing out on my emails, remember to check your “Promotions” tab and mark whytryai@substack.com as a “Safe Sender.”
Here’s what happened in AI last week:
👩💻 AI releases
Aleph Alpha open-sourced Kolibri, a German-and-English AI model built in Europe for organizations that need to run AI on their own hardware.
Anthropic news:
Claude Code now has “mods”: small add-ons that change how it looks and behaves (and Claude can even make new ones for you).
Claude Sonnet 5.5 is a faster and cheaper upgrade to Sonnet 5 that nearly matches Opus 5.5 on everyday knowledge-work tasks.
Black Forest Labs released FLUX 3 Image that lets you move and edit elements by drawing boxes on a canvas while keeping the rest of the image intact.
ElevenLabs launched Eleven v4, its most expressive text-to-speech model yet, which can also clone a voice from just ten seconds of input audio.
Google launched Guided Vision in Gemini Live on Android, which helps blind and low-vision users find objects, match outfits, read fine print, and more.
Ideogram released Ideogram 4.5, an image editing model that keeps the picture consistent across multiple edits, even on high-resolution images.
Meta news:
Edits assistant in Instagram’s Edits app can study how your posts perform and suggest captions, hooks, and scripts based on what works for you.
Muse for Small Business is an AI agent that plugs into tools like Canva and Shopify to help business owners analyze campaigns, optimize ads, and more.
Muse Gadgets lets you connect Muse to homemade devices like e-ink displays and Raspberry Pi boards using free, open-source kits.
Microsoft news:
MAI-Transcribe-2-Streaming is a real-time transcription model that turns live speech into text in 60 languages.
MAI-Voice-2.1 is a realistic and expressive text-to-speech model that can instantly match a voice from a short clip and speaks 23 languages.
OpenAI news:
Dots are always-on agents with their own cloud computers that learn your preferences and can keep working on tasks while you’re away.
GPT-6.1 Sol nearly matches GPT-6 Astra performance on coding, computer use, and professional work tasks for a fifth of the price.
Sign in with ChatGPT lets users log in to third-party apps with their ChatGPT accounts and use their subscription credits and rate limits in certain apps.
Shopify news:
Canvas lets merchants design their whole store from a single workspace together with a “Sidekick” AI agent that makes changes in real time.
WebMCP checkout lets AI agents complete purchases on Shopify stores, including with Shop Pay.
SpaceXAI launched Team Bots: shared Grok assistants that learn from your team’s files and apps and remember context for everyone.
🔬 AI research
Google news:
Gemini 4 Argon is a frontier model that beats top-tier counterparts from other AI labs on many benchmarks (only for cybersecurity defenders for now).
Skills for Gemini and Workspace let users save and reuse custom AI instructions across tasks, rolling out in October and eventually replacing Gems.
Runway opened early access to Runway Ads that create visual ads, publish them to Google, Meta, and TikTok, and optimize them based on performance.
📖 AI resources
“Jagged Performance of Frontier Models” [REPORT]: comparison of how top AI models handle real agent tasks like web browsing, by Fig Labs.
“GLM-5.3 Cyber Capabilities” [REPORT]: Anthropic’s red-team analysis of inadequate safeguards on hacking-capable open-weight models.
“Open TTS Leaderboard” [BENCHMARK]: a Hugging Face leaderboard to compare open text-to-speech models on multilingual speech and voice cloning.
“State of Agent Skills” [REPORT]: Vercel’s look at what people are teaching AI agents through its skills.sh registry and which skills actually get installed.
🔀 AI random
Ataraxos, built by researchers from Carnegie Mellon, MIT, NYU, and Stanford, beat the best human Stratego player 15-1 after only $8K in training costs.
The FTC opened an investigation into Anthropic, OpenAI, and other AI companies over the risks their products may pose to consumers.
🤦♂️ AI fail of the week
Even with the self-juggling disappearing ghost balls, the real mystery here is somehow the choice of outfit.
🔐 All my paid goodies in one place
Take a peek at what’s behind the paywall:



