AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: OpenAI, the chip it built, how it's used, and the practical shift.
OpenAI and Broadcom unveiled Jalapeño, a chip purpose-built to run models rather than train them — the first time a frontier lab has designed its own serving silicon. Inference, not training, is where the recurring cost lives, so a lab that controls its serving stack controls its margins and stops paying the Nvidia tax on every token. If this pattern holds, the inference market fragments and pricing pressure flows downstream to everyone who rents capacity. If you're forecasting model-serving costs for 2027, stop assuming a single-vendor GPU curve and start modeling a multi-silicon world where your provider's hardware roadmap is a real variable.
SOURCE · OPENAIHOW TO READ THIS Top to bottom: who reported it, the headline number, the Un-0 method, then the unverified energy cut it would produce.
Databricks' former AI chief debuted Un-0, claiming it reproduces conventional AI workloads at one-thousandth of the energy. Treat the number as unproven until there's a reproducible benchmark — but the direction is what matters: as data-center power becomes the binding constraint (SemiAnalysis is now modeling 40GW+ of behind-the-meter capacity by 2028), efficiency starts to outweigh raw model quality at the margin. For anyone sizing infrastructure, the takeaway isn't 'switch today' — it's that energy-per-inference is becoming a first-class metric you should already be tracking next to latency and accuracy.
HOW TO READ THIS Read top to bottom: calesthio ships it, the coder turns into a studio, OpenMontage renders the edits, then stars roll in.
OpenMontage bills itself as the first open-source agentic video-production system — 12 pipelines, 52 tools, and 500+ agent skills wired into a coding assistant — and pulled 3,400+ stars in a single day. The signal isn't video specifically; it's that the agent-skills pattern is maturing into full end-to-end vertical workflows, not just code completion. If you're building internal agents, study how it decomposes a creative pipeline into discrete, composable skills — that architecture transfers to any domain where a human currently chains a dozen tools by hand.
HOW TO READ THIS Read top to bottom: the startup, its game-to-agent pipeline, the sim-to-real crossing, then the raise.
General Intuition raised $320M to train agents on millions of hours of video-game footage, betting that action data — not text — is the path to human-like intuition for embodied agents. It's one of the boldest swings yet at the physical-AI problem, and a clear vote that the next capability jump comes from new data modalities rather than bigger language models. Worth watching if you care about robotics or embodied systems: the open question is whether game-world skills transfer to messy physical environments, or stay trapped behind the sim-to-real gap.
Turns messy PDFs and Office docs into clean, LLM-ready markdown/JSON — the unglamorous ingestion layer most agentic RAG pipelines actually need.
A JavaScript in-page GUI agent that drives web interfaces from natural-language instructions — browser automation without brittle selectors.
Clone any website with a single command using AI coding agents — a fast scaffold for prototypes and redesigns.
An AI agent that evaluates and scores resumes — a concrete (and bias-sensitive) example of agents moving into HR workflows.
Field notes for going from vibe-coding to disciplined agentic engineering with Claude Code.
The Sequence Qwen pushes into robotics — another major model lab planting a flag in embodied, physical AI.