AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: the paper's claim, the arm-and-GPU setup that produced it, then the resulting size and speed.
TurboVLA drops the LLM-centric vision-language-action architecture — it encodes vision and language separately and fuses them only just before action prediction — and reports 32 Hz control at 31.2 ms latency on 0.9 GB of VRAM with 0.2B parameters. On a real AgileX Piper arm it scored 80–92.5% across four manipulation tasks and beat π0.5 on all four. The significance is not the benchmark, it is the deployment envelope: a policy that fits in under a gigabyte of VRAM runs on the robot, not in a datacenter, which removes network latency and per-inference cost from the operating model entirely. If you are scoping embodied or edge inference work, stop assuming the policy layer needs a hosted GPU tier and re-run your cost and latency budget against on-device numbers.
HOW TO READ THIS Read top to bottom: two ways of collecting demos merge into one open dataset, train one policy, and hit 85% on a real insertion task.
Simple AI's HiFi-UMI is a handheld gripper rig with 3 mm workspace accuracy and no external tracking, and policies post-trained only on its demonstrations deploy directly to a real robot — within ±3 points of teleoperation across three VLA and world-action backbones, with 85% success on precision insertion. The team open-sourced 2,000 hours of synchronized demonstrations. Teleoperation time on real hardware is the most expensive input in robot learning, and a handheld rig that a person can carry through a warehouse or a kitchen collapses the cost of collecting it. The takeaway for anyone budgeting a robotics data program: the bottleneck is shifting from robot hours to human hours, which is a very different procurement problem.
HOW TO READ THIS Top to bottom: nine VLMs are tested, each is fed an embodied navigation task and marked wrong, and the best of all nine still scores only 16.8%.
HumanCLAW harnesses a physics-simulated humanoid to a VLM that issues atomic skill commands, deliberately isolating decision-making from motor execution across 1,218 egocentric episodes in 41 indoor scenes. Nine frontier VLMs were tested, none solved it, and the best reached 16.8% — with the authors attributing the failure to missing embodied self-awareness rather than to perception. Read this against the lead story: the motor layer is getting cheap and fast while the layer that decides what to do next is still near the floor. If your robotics roadmap assumes a frontier model can be dropped in as the planner, this benchmark is the number to argue with before you commit the architecture.
HOW TO READ THIS Read top to bottom: one core model, its fan of end effectors, the mid-task tool swap between them, then the small weight slice retrained per tool.
Generalist extended GEN-1 across roughly 9,000 end-effector variations — five-finger hands, tongs, whisks, power screwdrivers — pretrained on more than half a million hours of real interaction data, with fine-tuning shifting only 2.5–11.4% of weights per new gripper. The demos include swapping tools mid-task without restarting. That low weight-delta is the number worth tracking: it implies most of what the model knows is hardware-agnostic, and the gripper is closer to a thin adapter than a new model. Vendor blog, not a peer-reviewed result, so treat the success rates as directional — but the architectural claim, a shared policy layer under heterogeneous hardware, is the one that changes how you plan a fleet.
Physical Intelligence's open release of the π0 / π0.5 model family — the exact baseline TurboVLA claims to beat, which makes it the reference implementation to reproduce the comparison against.
The de facto open stack for robot learning — datasets, policies, and low-cost hardware configs in one place — and the shortest path from a handheld-demo dataset like HiFi-UMI's to a trained policy.
An open 7B vision-language-action model with published training code; useful as the heavyweight control against this week's argument that 0.2B is enough for real-time manipulation.
The original UMI handheld-gripper data-collection system that HiFi-UMI builds on — worth reading first if you want to understand what the 3 mm accuracy claim is actually improving.
A fast open physics engine for robotics and embodied AI — the cheap way to run the kind of simulated-humanoid evaluation HumanCLAW uses before you put a policy on real hardware.