AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: the arXiv paper releases open weights, the stack alternates Mamba scan rows with attention rows while a router lights only a few experts, and the result runs on your own servers inside an agent loop.
Nemotron 3 Ultra is a 550B-parameter open Mixture-of-Experts model — only ~55B active per token — that pairs Mamba's long-context efficiency with attention, pretrained on 20T tokens and tuned for agentic reasoning. The story isn't the parameter count; it's that a frontier-scale, tool-using base is now available under open weights and is genuinely self-hostable. For teams boxed in by hosted-API cost, latency, or data-residency rules, that shifts the build-vs-buy math for agents. Pull the weights and benchmark it on your own tool-use traces before assuming a closed API is your only path.
HOW TO READ THIS Read top to bottom: arxiv's paper, then old exact-prefix caching missing after any change, then the new editable cache fixing the divergent entry so reuse continues, ending in a fully-hit cache bank.
Edit and recompose prefill KV caches instead of discarding them whenever an input changes — breaking the exact-shared-prefix limit and lifting cache hit rates on templated or near-duplicate prompts.
HOW TO READ THIS Read top to bottom: the arXiv paper, its two rival algorithms, how each factors N, then the RSA risk it creates.
An empirical comparison of Shor's period-finding and Regev's newer lattice-sampling factoring — clearer footing on when RSA-style crypto breaks, and how soon post-quantum migration becomes urgent.
HOW TO READ THIS Read top to bottom: the paper, then qubits grouped into cosets, then fault tolerance, then fewer qubits.
A new family of two-block quantum LDPC codes built from group cosets, cutting the qubit overhead of error correction on the path to fault tolerance.
Open-source LLM engineering platform — evals, tracing, observability, and prompt management in one stack.
Open-source RAG engine that fuses deep document retrieval with agent capabilities.
ByteDance's open multimodal agent stack for building computer-use agents.
Official Chrome DevTools exposed as an MCP server so coding agents can drive and debug a real browser.
MCP server that indexes a repo into a persistent knowledge graph for code-aware agents.
The Sequence DeepMind's first real crack in next-token generation — early signal for life beyond the transformer.
Latent Space GLM-5.2 lands as the top frontend-coding model, alongside IndexShare for speculative decoding.
Interconnects The frame shifts from model safety to governing AGI-era institutions and policy.
SemiAnalysis Why RL training stalls when trainer and generator throughput don't match — and how to close the gap.