Thursday, October 1, 2026 Newsletter Advertise
Breaking
Research

Research Radar, October 1, 2026: 14 new AI and security papers to know

The most relevant new AI and security research from Nature & arXiv, with the authors' own summaries and links to the papers.

14 papers, Nature & arXiv

New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.

Peer-reviewed

New preprints (not yet peer-reviewed)

ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning

arXiv · .Ibrahim Ethem Deveci, Funda Tan c{C}al{i}k, Bar{i}c{s} Deniz Sau{g}lam, Duygu Ataman

Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains.

PhantomEnvironments: Training LLM Agents in Fictional Worlds

arXiv · Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q. Weinberger

Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination.

VirusCascade: Hijacking Collaborative Reflection in LLM-Powered Recommender Agents

arXiv · Yurong Hao, Wen Zhou, Guowei Guan, Tiantong Wu, Fuyao Zhang, Wei Yang Bryan Lim

Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts.

Emergent alignment and the projectability of ethical personas

arXiv · Guillermo Del Pinal, Youngchan Lee, Calum McNamara, Alejandro Perez Carballo

Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate different characters/perspectives, which can be elicited and refined during post-training.

Aligned Data Can Induce Misalignment via Context Confusion

arXiv · Yavuz Bakman, Duygu Nur Yaldiz, Baris Askin, Swastik Roy, Morteza Ziyadi, Salman Avestimehr, Sai Praneeth Karimireddy

Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another.

Large Knowledge Model: A Knowledge Foundation for Agentic Science at Scale

arXiv · Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E

Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This requires a knowledge foundation that supports high-concurrency access, preserves traceable and reusable reasoning, and grows incrementally.

SURE: Framework for Safety to Construct Trustworthy AI

arXiv · Soeun Han, Jisoo Lee, Jeongyong Shim, Eunkyeong Lee, Eunmi Kim

Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains.

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

arXiv · Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process.

The TechUpscale Brief

The day's cyber, AI and tech news in one short email, every weekday morning. Free. Unsubscribe anytime.

I'm most interested in

More Research