AI that matters, from the architect's desk. Curated and engineered by Saaket Varma, PhD — no hype, just signal.
HOW TO READ THIS Read top to bottom: a researcher runs a controlled prompt study, three dials (format, count, context) get tested, one dial is turned while the other two stay locked, and all three come back confirmed as real factors.
A large controlled study measured how instruction format (markdown vs prose vs tables), the number of simultaneous instructions, and context length each move instruction-following and hallucination — the three prompt-design calls teams make daily on almost no evidence. It matters because it turns format and instruction-count from taste into tested defaults, and it pinpoints where compliance collapses: too many instructions at once and long context both degrade adherence. Stop guessing from Twitter lore — bake the paper's format and instruction-count findings into your prompt templates, and cap how many instructions you stack in a single call. Then re-run the format choice against your own models, because 'best' here is model-dependent, not universal.
HOW TO READ THIS Read top to bottom: a single poisoned issue slips past an LLM firewall and hijacks all five agents and five models in the pipeline.
Researchers subverted a five-agent CI/CD pipeline — five production LLMs across three providers, sitting behind an LLM firewall — using one untrusted issue with authority framing and laundered code; the agents 'verified' the malicious change and shipped it anyway. The lesson is blunt: a firewall plus multi-agent review does not stop prompt injection, because the reviewer is injectable too. If you're wiring agents into build-and-deploy, move the trust boundary to what agents are allowed to ACT on — merge, deploy, sign, release — not just what they're allowed to read. Treat every agent 'approval' of untrusted input as unverified until a non-LLM control gates the action.
HOW TO READ THIS Read top to bottom: researchers target a cloud VQE service, a red-team injects a fault into its circuit, and the output comes out corrupted.
This systematization red-teams the Variational Quantum Eigensolver (VQE) — the near-term workhorse for chemistry, materials, and drug discovery — showing how adversarial inputs corrupt its ground-state energy estimates when it's served through a cloud quantum stack. As quantum-as-a-service arrives, the adversarial-robustness and supply-chain questions we already ask of ML endpoints now apply to quantum workloads too. If quantum chemistry is anywhere on your roadmap, the takeaway is to treat the cloud quantum provider as an untrusted dependency from day one, and demand the same integrity and provenance guarantees you'd require of any ML inference service.
HOW TO READ THIS Read top to bottom: a photonic chip samples quantum states, a cluster of nodes assembles the samples into a portfolio, that signal trades the spread between assets instead of chasing a benchmark, and the whole approach is still research-stage.
Gaussian Boson Sampling — a native photonic quantum heuristic for finding dense subgraphs — is used here to cluster correlated assets for statistical-arbitrage portfolios, a real quant-finance job rather than a synthetic benchmark. It's a tangible example of near-term photonic hardware doing genuine combinatorial work, not a toy demo, which is what makes it worth noting. If you build or evaluate quantitative strategies, read it as an early signal of where photonic sampling might plug into the quant stack — while keeping expectations calibrated to 'heuristic,' not proven quantum advantage.