AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: Codex's full 372K window, a change made with no announcement that developers had to find themselves, and the resulting 272K window with the missing 100K left visibly empty.
OpenAI reduced Codex's model context window from 372K to 272K tokens — a roughly 27% cut that developers spotted before any announcement did, while codex sits at #1 on GitHub trending with over 100K stars. If your agent workflow assumes it can hold a whole repo in working memory, that assumption quietly broke, and the failure mode is silent: truncation and degraded recall, not an error message. This is the durable lesson about building on hosted models — context capacity is a vendor-side dial, not a contract, and it can move without a changelog entry. Re-measure your actual token usage against the new ceiling this week, and if you're near it, move to retrieval or explicit file selection rather than repo-dumping.
HOW TO READ THIS Read top to bottom: authors sued Anthropic, a judge approved the deal, the $1.5B fund pays authors out, making it the largest AI copyright case yet.
A judge granted final approval to Anthropic's $1.5 billion copyright settlement, closing one of the largest AI training-data cases to date. The case is resolved; the underlying legal question is not — no precedent was set on whether training on copyrighted work is fair use. What did get established is a number. Every lab training on copyrighted corpora now has a public benchmark for what this exposure costs to settle, and every enterprise legal team reviewing an AI vendor has a figure to anchor their diligence questions against.
HOW TO READ THIS Read top to bottom: millionco ships the tool, it lints agent-written React, flags a bad pattern, and racks up stars.
react-doctor is trending with a blunt pitch: your agent writes bad React, this catches it. Past 14K stars in TypeScript, it's less interesting as a tool than as a category signal. As agents write a growing share of the code, the review layer has to scale with them — and human review doesn't. Expect more deterministic tooling built specifically for the failure patterns of machine-generated output, which are systematic and therefore lintable in a way human mistakes never were. If you're running agents in a real codebase, the question worth asking is what your equivalent gate looks like.
HOW TO READ THIS Read top to bottom: the maker, the tiny agent it built, how it runs locally on a phone, and the result.
needle is a 26-million-parameter agentic model built for constrained hardware — small enough for phones, wearables, and embedded targets. That parameter count is the headline: it's three to four orders of magnitude below the frontier, which puts agentic behavior inside a power and memory budget that doesn't require a network round-trip. For anyone tracking edge deployment, this is a concrete data point that the agentic layer is being pushed down the stack, not just scaled up in the datacenter. Worth watching what capability actually survives at that size.
The Pythonic way to build MCP servers and clients — the default starting point if you're wiring tools into Claude.
Design principles for LLM software that's actually production-grade — the closest thing to an architecture rubric for agents.
An open-source terminal coding agent — the open-weights answer to Codex and Claude Code.
Very low-latency speech-to-text, intent recognition, and TTS for building voice agents and interfaces.
A structured math and CS curriculum for engineers moving from applying models to researching them.
Interconnects Nathan Lambert on Kimi K3 and what the latest open-weights release does to the closed-model gap.
Import AI Jack Clark on the open-vs-closed capability gap, Kimi K3, and Demis Hassabis's policy proposal.
The Sequence A week-in-review on China, model compression, and the open-model race.