AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: Sakana ships the small Fugu core, it issues commands to three larger models, their results converge into one orchestrated answer, and that yields a 73.7 SWE-Bench Pro score.
Sakana AI shipped Fugu and Fugu Ultra — a roughly 7B orchestrator that learns when and which frontier model to call, and will even recurse on itself when that's the better move. Fugu Ultra tops coding and reasoning benchmarks (73.7 on SWE-Bench Pro, 95.5 GPQA-Diamond), putting a 7B conductor in the same tier as Fable 5. The bet underneath is the one worth your attention: smart routing can beat raw scale. Audit your own stack for the queries you're paying frontier prices on that a router could handle cheaper — orchestration is becoming a first-class architecture layer, not a wrapper.
HOW TO READ THIS Read top to bottom: the source model, its compression, then the size cut it enables.
Tensor-network compression (CompactifAI) is now a shipping product, not a paper — open-source HyperNova 60B lands at roughly half its parent's size with 50–80% lower inference cost and retained accuracy, on the back of Multiverse's $215M raise. Compression just crossed from research curiosity to procurement line-item: same quality at half the footprint changes what you can run on-prem, at the edge, or on commodity hardware. If your inference bill scales with usage, a compressed model is now a credible default — not an experiment.
The tell isn't a startup compressing models — it's NVIDIA co-shipping one. Today Multiverse and NVIDIA released Pulsar 16B open-source (Apache 2.0, on Hugging Face): NVIDIA's own Nemotron-3 30B compressed to 16B with CompactifAI, built using NVIDIA's own Model Optimizer and Megatron Bridge libraries, no retraining. It holds 30B-class reasoning (AIME 2025 87.2, ~15 points ahead of gpt-oss-20B) at half the parameters and 43% more throughput on Blackwell. When the company that sells the GPUs helps you need fewer of them — and puts the weights in the open — the efficiency thesis stops being contrarian and becomes the roadmap.
HOW TO READ THIS Read top to bottom: Intel's efficiency climbed while its stock stayed flat, until the market suddenly closed the gap and shares jumped 250%.
Inference demand is making CPUs relevant again — Intel's ~250% 2026 run, NVIDIA baking adaptive compression into Rubin, and Google's FP4 inference all point the same way. Efficiency, not raw scale, is now the industry's roadmap: the winners optimize cost-per-token, not parameter count. 'Cheaper to serve' is becoming the competitive moat — and it's quietly reshaping which chips, and which vendors, matter.
HOW TO READ THIS Read top to bottom: the M4 Framework debuts, runs tokens through a 4-stage pipeline, cuts cost per token by 49% on a baseline scale, but the result still sits in unverified research status.
This isn't hindsight — I pioneered it. A year ago at AI4 2025 I led the Deloitte–Multiverse–Intel study that proved CPU inference can match GPU accuracy: 60% model compression and 49% lower cost-per-token on Llama 3. The multi-hardware, efficiency-first thesis I keynoted at GDS Atlanta is now the industry's mainstream bet — today's NVIDIA-and-Multiverse open-source headline is simply the market catching up to the architecture we shipped a year back.
An 'Agent OS' for spec-driven development — stop prompting, start specifying what you actually want built.
An integration layer for agents — one interface for the external tools and services your agents need to reach.
OpenAI-compatible proxy stacking 16 providers' free tiers (~1.7B tokens/month) behind a single /v1 endpoint.
Self-hostable bookmark-everything app — links, notes, images — with AI auto-tagging and full-text search.
Interconnects Nathan Lambert calls GLM-5.2 the step change that finally makes open-weight models viable for real agents.
The Sequence Last week in AI, recapped: a reported $60B Cursor deal, Google's talent drain, and Midjourney's body scanner.