Tuesday, September 29, 2026 Newsletter Advertise
Breaking
Research

Research Radar, September 28, 2026: 11 new AI and security papers to know

The most relevant new AI and security research from Nature & arXiv, with the authors' own summaries and links to the papers.

11 papers , Nature & arXiv

New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.

Peer-reviewed

New preprints (not yet peer-reviewed)

BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering

arXiv · Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani, Xiang Li, Christopher Bain, Yuqi Zhou, Walker Peterson, Huidong Wang, Chenglang Hu, Jongchan Park, Xiao Cheng, Benjamin Swedlund, Sandra Murillo, Anjali Sivanandan, Shiyu Sun, Liang Lanfeng, Mohammad Tariqul Islam, Baju C. Joy, Ishaq N. Khan, Sreedhar S. Kumar, Gabriel Mercado-V'asquez, James V. Vizzard, Jonathan M. Matthews, Helen Huang, Xiaolu Guo, Ethan Nicklow, Guorui Chen, Ryan A. Neff, Surjendu Maity, Hyeonjin Park, Han-ho Joo, Katherine Dong, Yuyan Cai, Weihang Huang, Yichen Zou, Rui Yan, Raphael Figueroa, Artem Goncharov, Bella Rose Schremmer, Lian Elsa Linton, Keisuke Goda, Liang Gao, Ke Cheng, Leonardo Morsut, Jennifer L. Wilson, Jianping Fu, Lim Chwee Teck, Deblina Sarkar, Andreas Hierlemann, Savac{s} Tay, Alexander Hoffmann, Donald Richieri Griffin, Jun Chen, Shana O. Kelley, Shyni Varghese, Jinwoo Cheon, Wilbur A. Lam, James J. Moon, Wilson W. Wong, Samir Mitragotri, Dino Di Carlo

Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks.

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

arXiv · Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song

AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection.

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

arXiv · Hanzhang Ma, Ali Hariri, Tianxiang Shen, Bohua Zou, Qianjun Zheng, Ji Wang, Li Yi, Ning Jia, Yutao Liu, Haibo Chen, Lin Wang, Debayan Roy

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a combination of coarse-grained permission rules and LLM-based judgments about individual proposed actions.

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

arXiv · Pierre Beckmann, Marco Valentino, Andre Freitas

Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents.

From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs

arXiv · Jiawen Wang, Pritha Gupta, Eyke H"ullermeier, Xiaoxue Gao, Nancy F. Chen

Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically investigating their security risks is becoming increasingly important towards a trustworthy LLM era.

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

arXiv · Jiayi Chen, Guiling Wang

While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs.

Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains

arXiv · Toqeer Ali Syed, Asadullah Abdullah Khan

This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source integrity, dependency and SBOM analysis, CI configura tion auditing, artifact verification, and runtime policy evaluation, each supported by a large language model (LLM) that interprets artefacts, reasons over tool outputs, and produces structured security reports.

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

arXiv · Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu

A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface.

The TechUpscale Brief

The day's cyber, AI and tech news in one short email, every weekday morning. Free. Unsubscribe anytime.

I'm most interested in

More Research