AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: xAI ships Grok 4.5, it claims Opus-class parity, then hits that same capability line with a leaner bar, shipping cheaper tokens.
SpaceXAI shipped Grok 4.5, and Musk is pitching it as an 'Opus-class' model that undercuts rival frontier models on price and efficiency. Treat the class claim as marketing until independent benchmarks land, but the pricing pressure is real either way — every aggressive price-performance release forces the other labs to respond, and API rate cards rarely move in only one direction. If you're running production inference, this is the moment to re-check your cost-per-task assumptions rather than your leaderboard assumptions. Benchmark it on your own workloads before believing anyone's chart.
HOW TO READ THIS Read top to bottom: OpenAI ships the GPT-Live model, which listens and speaks in the same moment instead of taking turns, and it now runs ChatGPT Voice.
OpenAI's new voice model generation can speak and listen at the same time — full duplex, not turn-taking — and it's already powering ChatGPT Voice. That's the specific unlock behind real-time translation and natural interruption, the two things that make current voice agents feel like walkie-talkies instead of conversations. If you've built voice UX on the assumption that the model must finish talking before it can hear you, that assumption just expired. Expect every voice-agent stack to be rearchitected around duplex within a year.
HOW TO READ THIS Read top to bottom: OpenAI's Codex node docks as a plugin into Claude Code's shell, runs live inside it, and the repo reaches 26,873 stars.
OpenAI published an official plugin that lets Claude Code users call Codex to review code or delegate tasks — and it's sitting at 26,000+ stars as one of GitHub's hottest repos today. Read that again: OpenAI shipped first-party interop into Anthropic's agent. Rival labs building into each other's toolchains is the strongest signal yet that multi-agent, multi-vendor stacks are the default architecture, not an experiment. If your agent strategy assumes a single-vendor stack, this is your cue to design for composition instead.
HOW TO READ THIS Read top to bottom: Cowork's desktop-only start, its expansion to phone and web, the Max-first rollout, then work following across devices.
Anthropic is rolling out Claude Cowork on mobile and web for the first time, starting with Max subscribers. The agentic-work platform escaping the desktop matters more than it sounds: it turns agent supervision into something you do from your phone between meetings, not from a dedicated terminal. Delegation-from-anywhere is the workflow shift that makes long-running agents practical for people who don't live in an IDE.
Microsoft's text-space optimizer that trains reusable natural-language skills for frozen LLM agents from trajectory data — prompt engineering, systematized.
Instant, concurrent, lightweight sandbox for running AI-agent code securely — the isolation layer every agent stack eventually needs.
Google Labs' library of Agent Skills for the Stitch MCP server, each following the open Agent Skills standard — the skills format is going cross-vendor.
Train a 64M-parameter LLM completely from scratch in about two hours — the best hands-on way to actually understand what's inside the big models.
Local, open-source AI app builder — a v0/Lovable/Replit alternative that keeps your code and prompts on your own machine.