AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: AMD takes aim at Jetson, fuses Kria with Ryzen AI X100 into one module, closes the control loop, then reliability jumps 3.4x.
AMD launched the Kria AI Robotics Developer Platform alongside the Ryzen AI Embedded X100 Series — up to 16 Zen 5 cores, an RDNA 3.5 GPU, an NPU and FPGA fabric on one platform, rated for 8,000+ control decisions per second and sub-100ms vision-language-action reasoning. The comparative numbers are AMD's own — 3.4x better real-time control reliability and support for 2.3x more concurrent agents than NVIDIA's Jetson T5000 — so treat them as vendor benchmarks until someone independent reruns them. What is not in dispute is the structural change: Jetson has been the default robot brain with no credible open alternative, and pricing, allocation and roadmap leverage all follow from that. If you have robotics or edge-inference work landing in 2027, put the X100 on your evaluation list now and ask ODM partners about the Q4 2026 SOMs — the point of a second vendor is the quote you get from the first one.
HOW TO READ THIS Read top to bottom: NVIDIA releases weights, the 4B model ships, it outputs joint angles instead of pixel frames, and runs live at 15 Hz on edge hardware.
NVIDIA open-weighted a 4-billion-parameter omnimodal world model that emits robot joint angles and trajectories directly instead of predicted pixels — 32 actions per inference, real-time 15 Hz control on Jetson Thor, and it also runs on consumer GeForce RTX. Step-distilled variants cut sampling from 50 denoising steps to four for up to 25x faster inference, which is the part that actually matters for closed-loop control. It lands the same week AMD went after Jetson, and that is not a coincidence: open weights are how you hold a developer base when the hardware moat gets contested. If you have been waiting for physical-AI reasoning that doesn't need a datacenter round-trip, this is the first one you can download and benchmark yourself.
HOW TO READ THIS Read downward from the original model through layers packed at different precisions to the workstation result, with proportional filled bars comparing the reported sizes.
Unsloth's Dynamic 2.0 quantizations take the 754B-parameter GLM-5.2 from 1.51 TB in BF16 down to 217 GB at 1-bit and 238 GB at 2-bit, by assigning a different precision per layer instead of one uniform bit width. Uniform quantization is what made low-bit compression a quality cliff; per-layer allocation spends bits where the loss is sensitive and starves the layers where it isn't. The practical consequence is ownership: 217 GB is a workstation memory budget, not a rented endpoint. Run it against your own evals before trusting it — but at that size the experiment costs you a weekend rather than a procurement cycle.
HOW TO READ THIS Read top to bottom: a startup builds a chip, aims it at local 100B models, runs prefill across parallel lanes, then hits 1,416 tok/s.
Acrab unveiled GΞLIX 1, a 5nm edge AI SoC with a 20-core Arm CPU, multicore NPU acceleration and 273 GB/s of unified memory bandwidth, sized to run open models up to the 100-billion-parameter class locally — plus Agent Box, a one-time-purchase personal AI appliance built on it. The headline number, 1,416.8 tokens/sec prefill on Gemma 26B against 188.9 on a Mac Mini M4 Pro, comes from the company's own press release, not a third party; log it as a claim, not a result. Track it anyway, because purpose-built edge silicon decisively beating a general-purpose desktop chip is the precondition for no-token-fee local agents being a product instead of a hobby. Wait for independent benchmarks before you plan a budget around it.
Anti-AI-slop design skill for Claude Code, Cursor and Codex — opinionated defaults so generated UI stops looking generated.
MCP server for symbol-level GitHub code retrieval; claims 95%+ token savings versus dumping whole files into context.
20+ LLM implementations with recipes to pretrain, finetune and deploy at scale — the reference stack if you're bringing a quantized model in-house.
Agent-team framework where agents pick up tasks, message each other and review each other's work.
One hub for Claude Skills, agents, commands, hooks and plugins — check here before you build something that already exists.
SemiAnalysis Whether AMD's Advancing AI 2026 stack can actually crack the CUDA moat — the software half of today's lead story.