AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: the White House orders OpenAI, GPT-5.6 hits a gate, only partners pass while the public is blocked, and the result is the first gov't-mandated AI delay.
The Trump administration asked OpenAI to delay the public launch of GPT-5.6, routing it to select partners first over security concerns — the first time a sitting US administration has directly intervened to slow a commercial model release. It happened, OpenAI appears to have complied, and that makes it a live precedent: government-gated releases are no longer hypothetical. Every major lab is now watching to see whether this becomes standing policy or stays a one-off request. If you're building on frontier APIs, start treating regulatory timing as a variable in your release planning alongside technical readiness.
HOW TO READ THIS Read top to bottom: one agent task loops, each step re-spends tokens, usage compounds, ending in a 56x internal research finding.
OpenAI's own Codex usage shows median output tokens grew 56× in Research, 32× in Customer Support, 27× in Engineering, and 13× in Legal since November 2025 — not analyst projections, but actual internal consumption from the lab building the models. The compounding is telling: token growth rates differ sharply by domain, which means agent loop depth and task complexity are already diverging across functions. For teams building agent infrastructure, this is the demand curve to plan against; orgs without headroom on token economics and rate-limit architecture are about to feel the ceiling.
HOW TO READ THIS Read top to bottom: the repo, its skill count, the six-way split, then the agents that gain it.
A GitHub repo maps 817 structured cybersecurity skills for AI agents across MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF, and the anti-fraud F3 framework — the most systematic attempt at a production security skill layer for agents published to date. Most teams shipping agents into production have either no security skill mapping or something bespoke and untested; this is the scaffold you'd otherwise spend months building. It covers 29 security domains, claims compatibility with Claude Code, Codex, Cursor, and Gemini CLI, and is Apache 2.0 — evaluate it against your agent's actual attack surface and adopt what fits.
HOW TO READ THIS Read top to bottom: Patronus AI banks $50M, builds eval infrastructure, then fires simulated attacks at an agent to check pass or fail.
Ex-Meta AI researchers raised $50M to build synthetic digital worlds that red-team agents before they touch real systems — automated adversarial QA at production scope. Agent evals are consolidating into their own infrastructure category: teams moving agents into production are finding that unit tests and manual review don't surface failure modes at the depth or breadth required. If you're taking agents from prototype to production, budget for a dedicated eval layer now; the cost looks small next to the first production incident a real red-team would have caught.
Open-source coding agent — a self-hostable alternative to Codex CLI and Claude Code for teams that want full control of the agent loop.
Open standard for agent skills via `npx skills` — composable, installable tool packages for AI coding agents.
Open context layer for data and AI — metadata management platform for wiring trusted data context into agentic pipelines.
AI-native cloud OS on Kubernetes managing the full application lifecycle — infrastructure substrate built for agent-first deployment.
Value investing research framework built on Claude Code + parallel multi-agent — Buffett/Munger/Duan methodology as an agentic research pipeline.