AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
Today's stories expose four engineering constraints: repository migrations, persistent memory, modular quantum scaling, and embedded QKD security analysis.
HOW TO READ THIS Read left to right: 520 runs across 20 whole-repository migrations faced a migration audit, fixed tests, and six-agent verification, with 28 passing every stage.
Deyao Hong and nine co-authors introduced SWE Refactor Bench to test whether coding agents can complete long-horizon, whole-repository stack migrations. The benchmark covers 20 migrations across four kinds of technical debt, extending evaluation beyond the bug-fixing tasks that dominate current agent benchmarks. It leads this digest because migrations are economically important engineering work, yet passing existing tests can conceal whether an agent performed the requested transformation at all.
The protocol first audits migration completeness, then runs a fixed behavioral test suite, and finally asks six independent coding agents to generate targeted tests for hidden differences. This design catches a failure mode the authors call Blindness, in which an agent preserves or copies the old implementation instead of completing the migration. Across 520 runs involving eight frontier models and 26 model-effort configurations, only 28 runs passed all three stages. Thirteen of the 20 tasks produced no accepted solution, while the best model, claude-opus-5, scored 47.0 out of 100.
For practitioners, the results argue against treating patch-level agent performance as evidence that repository-scale modernization can run unattended. The benchmark's important distinction is between migration completeness and behavioral correctness, which prior behavior-only evaluations do not measure separately. Agents or workflows that improve both capabilities could offer a competitive advantage by making neglected technical-debt programs cheaper and more repeatable. The evidence remains a first-version preprint built around 20 migrations, so its scores should be treated as a demanding baseline rather than a universal measure of production readiness.
HOW TO READ THIS Read top to bottom: one interaction seeds a topic-bound adversarial command, and a later related prompt recalls it through the agent and steers the response without direct access to the store.
Hanling Tian and seven co-authors developed InjecMEM, an attack against the persistent memory systems increasingly used by LLM agents for personalization and continuity. Their attack uses one interaction and requires no direct read or edit access to the memory store, yet aims to steer later answers on a target topic toward a chosen output. This paper was selected because persistent memory changes malicious input from a transient prompt-level risk into a threat that can influence future sessions.
InjecMEM combines a retriever-agnostic anchor containing high-recall topical cues with a short adversarial command. The command is learned through gradient-based coordinate search across synthetic prompt templates and insertion positions, with joint optimization across model backbones used to examine transfer. Across multiple memory systems and backbone models, the authors report reliable topic-conditioned retrieval and targeted generation, persistence under memory drift, and little effect on unrelated queries.
Teams deploying long-lived agents should therefore treat retrieved memories as untrusted input and apply controls at storage, retrieval, and execution boundaries. The specific contribution is an injection designed to persist through the memory pipeline without requiring control of the underlying store. Defenses that can identify or neutralize this pattern could become a practical differentiator for agent platforms used in finance, government, and other sensitive environments. The paper is a 29-page preprint accepted at COLM 2026, but the supplied evidence does not establish how well the attack transfers to every proprietary memory architecture or survives production defenses.
HOW TO READ THIS Follow the protected route across superconducting modules: latency and noise appear only at inter-chip boundaries, while the modular result is shown with modest estimated overhead against an ideal monolith.
The research team behind this preprint evaluated whether slow, noisy inter-chip operations make modular fault-tolerant superconducting quantum computers impractical. They released a hardware-grounded architectural co-design and resource-estimation protocol for surface-code processors assembled from manufacturable quantum processing units. The paper was selected because modular construction underpins many superconducting roadmaps, and the cost of crossing chip boundaries could determine whether those systems scale.
The design confines inter-chip latency and noise to module boundaries instead of allowing link constraints to become a system-wide orchestration bottleneck. The authors then estimate resources for RSA-2048 factorization using experimentally anchored parameters and realistic superconducting hardware constraints. Against a large ideal monolithic baseline, their calculations indicate only modest additional qubit and execution-time overhead, with the estimated penalty remaining nearly scale-invariant across a broad range of module capacities.
For hardware architects, this suggests that imperfect links may be managed through system design rather than solved entirely at the device level. The useful difference from simpler modularity claims is the combination of boundary-localized link effects, surface-code architecture, and workload-level resource estimates. If the assumptions hold in hardware, modular builders could gain a manufacturing and scaling advantage without paying an escalating logical-resource penalty. The evidence is still simulation and resource estimation in a preprint, not a demonstrated multi-module fault-tolerant computer, so the result depends on whether experimental systems match the modeled parameters.
HOW TO READ THIS Sender and receiver statistics fork into a bypassed semidefinite-program branch and an animated eigenvalue branch whose d²-scaling memory fits a 1 GB Raspberry Pi for real-time key rates.
The authors of this preprint developed a method that converts observed quantum key distribution data into a secure key-rate estimate using eigenvalue computations alone. It replaces the semidefinite or entropy-cone programs used by existing numerical approaches. The work was selected because operational QKD hardware needs security analysis that can run locally and continuously, rather than depend on expensive offline computation.
The method reformulates the calculation so that memory scales as d^2 instead of d^4 in the underlying Hilbert-space dimension. The team demonstrated real-time estimation on a Raspberry Pi with 1 GB of memory and a Cortex-A53 processor. According to the reported benchmarks, that implementation outperformed existing workstation-based approaches by several orders of magnitude.
This matters to practitioners building quantum-secured finance or government networks because security bookkeeping can potentially move directly into constrained QKD devices. The genuinely new element described here is achieving the estimate through eigenvalue calculations without semidefinite programming while sharply reducing memory demand. Embedded analysis could provide a competitive advantage through faster monitoring, simpler deployment, and less dependence on separate compute infrastructure. The result remains a preprint demonstration, and broader validation is needed across protocols, device conditions, and adversarial assumptions before it can support production security claims.