AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: arXiv publishes the paper, the recipe is released open instead of locked like others, its data and training steps are shown in full, and other teams reproduce the same agent from it.
OpenThoughts-Agent released open data recipes for building broadly capable agents — the curation steps, data mix, and process that labs usually keep behind closed doors. Most agent training today is a black-box art: teams overfit to a single benchmark and can't say why their pipeline works. A documented, reproducible recipe means you can audit what went into an agent and reproduce results instead of trusting a leaderboard number. If you're training or fine-tuning agents, read the recipe and diff it against your own data pipeline — the gaps are usually where your reliability problems live.
HOW TO READ THIS Read top to bottom: a teacher model distills into a student, old trial-and-error runs are replaced by a predictive scaling-law curve, cutting compression cost.
This paper derives empirical scaling laws for distilling a large model into a smaller, task-specific one under real latency and cost budgets — not accuracy in the abstract. The payoff is predictive: you can estimate how small you can go before accuracy falls off, before spending the compute to find out. For anyone shipping at the edge or under a cost ceiling, that turns distillation from trial-and-error into a sizing calculation. Pull the curves and use them to set your floor ahead of the next distillation run.
HOW TO READ THIS Read top to bottom: raw unlabeled documents flow into DREAM, which drops the usual labeled-pair requirement and clusters embeddings directly, so a query still finds its match.
DREAM trains dense retrieval embeddings autoregressively, dropping the labeled positive/negative pairs that contrastive retrievers depend on. The labeled-pair bottleneck is real — building good training pairs is the slow, expensive part of standing up a custom retriever for RAG. If autoregressive training delivers competitive retrieval without that labeling cost, it lowers the barrier to a domain-tuned retriever. RAG builders should benchmark it against their current contrastive embedder on in-domain queries before committing to another labeling cycle.
HOW TO READ THIS Read top to bottom: a paper ships a method that reads the model's own gradients, whose spike pattern outs a hallucinated token without any separate checker model.
Grad Detect flags LLM hallucinations using the model's internal gradient signals, rather than a second model or repeated sampling to cross-check answers. External checkers and self-consistency are expensive — they multiply inference cost on every call. A gradient-based signal is a lighter-weight reliability layer you can run inline in high-stakes production. If you're gating outputs today with an LLM-as-judge or N-sample voting, evaluate this as a cheaper first-pass filter ahead of the expensive check.