AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Top to bottom: the arxiv paper, the cache shrinking to 4-bit, then a GPU packed with more agents.
UltraQuant compresses the key-value cache to 4 bits for context-heavy agents — the case where a long prefix is reused across many short turns and concurrency, not raw speed, decides GPU utilization. KV memory, not compute, is usually what caps how many sessions you can serve at once, so squeezing it ~4x lets you pack more concurrent agents onto hardware you already own, with no retraining. If you serve long-context agents, treat this as a concurrency-and-cost lever to benchmark before provisioning more GPUs — and measure quality at 4-bit on your own traffic, since aggressive cache quantization can quietly degrade long-range recall. The headline here isn't a faster model; it's cheaper density.
HOW TO READ THIS Read top to bottom: one skewed judge scores agents, its bias fans out through the network, so audit the evaluator.
A formal result that one skewed LLM-judge propagates and compounds across a multi-agent system — audit and isolate your evaluators before a single bad judge corrupts the whole network's decisions.
HOW TO READ THIS Read top to bottom: a quantum kernel circuit is rewritten as a tensor network, whose bond width sets simulation cost, making it classically simulable.
Proves entangling quantum kernels are matrix-product-operator factorizations of their Fourier tensors — pinning down exactly when a quantum kernel is classically simulable, and therefore where any real QML advantage could live.
HOW TO READ THIS Read top to bottom: an arxiv paper proposes swapping the exposed classical gradient link for a quantum-shielded channel during distributed all-reduce, yielding cheaper and more private training.
A quantum 'ring all-reduce' that makes gradient sync both more bandwidth-efficient and information-theoretically private — a speculative but concrete sketch of a fabric for large-scale training.
LLM research agent that surveys a topic and writes a full, cited report — a building block for grounded knowledge curation.
An ADE for running a fleet of parallel coding agents against your own subscription.
Review-first terminal diff viewer built for the agentic-coding loop — read every hunk before it lands.
754 cybersecurity skills for AI agents, mapped to MITRE ATT&CK, NIST CSF 2.0, ATLAS and more.
Garry Tan's exact Claude Code setup — 23 opinionated role-agents (CEO, designer, eng manager) as a starter team.
Interconnects Argues that banning open-weight models would be a strategic mistake — worth reading as policy pressure builds.
The Sequence Week in AI: a $60B Cursor deal, Google's talent drain, and Midjourney's body scanner.
Latent Space Flags GLM-5.2 passing the vibe check against GPT, with Z.ai forecasting an open Fable-class model by December.