AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: a huge model is squeezed into 8GB, ground through a bare CPU one slow token at a time, and still runs on an ordinary PC.
A developer who runs Kimi K3 on 32 H100s at work wrote a C99 inference engine that loads the same model on a single CPU with 8GB of RAM, and a related project on HN reports 0.50 tokens per second in 29GB. None of this is production-viable; half a token per second is a demonstration, not a deployment. What it does kill is the assumption that you need a GPU cluster to touch frontier-scale open weights at all, which is a different question from serving them. If you have written off local or air-gapped evaluation of a large open model because the hardware quote came back at eight figures, re-run that math: capability access and throughput are now separate line items.
HOW TO READ THIS Read top to bottom: Microsoft ships the MCP server, which lets one agent reach work items, repos, and pipelines, driving 1,935 stars.
Microsoft's official MCP server for Azure DevOps exposes work items, repos, and pipelines to coding agents, and it is climbing GitHub fast at 1,935 stars. The significance is not the feature list, it is that the ALM vendor shipped this itself instead of leaving it to a community wrapper. That is the step where MCP stops being a demo and starts being procurement. Review it accordingly: an official server with reach into work items and pipelines is a credentialed path into your SDLC, so scope its tokens like a service account, not like an editor plugin.
HOW TO READ THIS Read top to bottom: Reddit's CEO questions the deal as the stock slides, Reddit content flows into Google, Google's AI answer keeps users from clicking back to Reddit, and the DMCA suit stays open as a result.
With the stock falling, Reddit's CEO publicly asked whether Google's AI Overviews deliver a 'win-win' and hinted the licensing deal could end, while Reddit keeps its DMCA suit against Perplexity's alleged scraper alive. The data that grounds and trains these models is being repriced in public, by the parties holding it. If any part of your retrieval stack depends on a third-party content license, yours or your vendor's, that is a continuity risk with a price attached rather than a settled input. Worth asking vendors directly what retrieval quality looks like the day a major source walks.
HOW TO READ THIS Read top to bottom: the model works a problem, its step-by-step chain has a broken link, the answer lands correct anyway, and evals only ever check that final answer.
Quanta surveys mounting evidence that reasoning models often reach correct answers through chains of thought that do not actually justify them. If your eval grades only the final answer, a model that reasons badly and guesses well scores identically to one that reasons correctly, and you find out which you bought when the distribution shifts. The practical move is to stop treating chain-of-thought as an audit trail and start testing process: perturb the inputs, check whether the stated reasoning moves with them, and score faithfulness separately from accuracy. In regulated work, where the reasoning trace is often shown to a human reviewer as justification, this is a compliance question and not only a research one.
Persistent memory for coding agents, argued from benchmarks rather than vibes; trending because context windows keep growing and agents keep forgetting anyway.
LLM-driven extraction from unstructured documents, packaged for API deployment and ETL instead of notebooks — the unglamorous middle of most enterprise AI projects.
Community instructions, agents, and skills for Copilot, rising because agent configuration is becoming a versioned team artifact rather than a personal setting.
A catalogue of 500 agent use cases sorted by industry; useful less as code than as a map of what people are actually trying to automate.
Microsoft's 21-lesson GenAI course, still climbing because it remains the default answer to 'where do I send the team to start.'