Thursday, October 8, 2026 Newsletter Advertise
Breaking
Research news

Research Radar, October 8, 2026: 14 new AI and security papers to know

The most relevant new AI and security research from Nature & arXiv, with the authors' own summaries and links to the papers.

14 papers, Nature & arXiv

New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.

Peer-reviewed

New preprints (not yet peer-reviewed)

The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search

arXiv · Rongzhe Wei, Peizhi Niu, Xinjie Shen, Tony Tu, Yifan Li, Ruihan Wu, Eli Chien, Pin-Yu Chen, Olgica Milenkovic, Pan Li

Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the prompt-optimization paradigm: whether through traditional algorithmic search or recent agent-based workflows, the resulting prompts typically retain malicious semantic signals that modern guardrails are primed to detect.

Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

arXiv · Vikhyath Kothamasu, Virginia Smith, Chhavi Yadav

LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpler, benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill the malicious intent.

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

arXiv · Xingang Guo, Jing Gu, Brian Jang, Renxiong Wang, Utkarsh Tyagi, Daniel Quigley, Steven Li, David Yan, Daniel Yue Zhang, Darvin Yi, Forrest Huang, HiJae Kim, Tianyi Zhang, Jared Lichtarge, Jihua Huang, Le Xue, Manan Tomar, Qiuyi Richard Zhang, Ruofei Yu, Seth Neel, Yaning Hu, Marcella Valentine, Xinzhe Jiang, Daniel Evans, Chenguang Wang, Dustin Tran, Tong Zhao, Yinfei Yang, Yunzhong He

Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception.

CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

arXiv · Rafid Ahmed, Joseph Fioresi, Mubarak Shah, Yuzhang Shang

Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing.

Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus

arXiv · Jairo Gudi~no-Rosero, Juan Ignacio Zambrano, Umberto Grandi, C'esar A. Hidalgo

Large Language Models (LLMs) are gaining traction as a method to generate consensus statements and aggregate preferences in digital democracy experiments. Yet, participants can introduce critical vulnerabilities in LLM-based systems.

Post-Quantum Cryptography from Quantum Stabilizer Decoding

arXiv · Jonathan Z. Lu, Alexander Poremba, Yihui Quek, Akshar Ramkumar

Post-quantum cryptography currently rests on a small number of hardness assumptions, posing significant risks should any one of them be compromised. This vulnerability motivates the search for new and cryptographically versatile assumptions that make a convincing case for quantum hardness.

Routing-Aware Safety Alignment for Mixture-of-Experts Models

arXiv · Jiacheng Liang, Yuhui Wang, Tanqiu Jiang, Ting Wang

Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts.

The TechUpscale Brief

The day's cyber, AI and tech news in one short email, every weekday morning. Free. Unsubscribe anytime.

I'm most interested in

More Research