New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.
Peer-reviewed
Tangermeme: a toolkit for understanding cis-regulatory logic using deep learning models
Retrofitted LLM can count the letter ‘i’s in ‘artificial intelligence’
Game theory driven multi-agent framework mitigates language model hallucination
Brain tumor segmentation using particle swarm optimized histogram equalization and a VGG19 based U-Net
OpenAI posts 700 maths preprints online: mathematicians are up in arms
Generalizable perturbation prediction
New preprints (not yet peer-reviewed)
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the prompt-optimization paradigm: whether through traditional algorithmic search or recent agent-based workflows, the resulting prompts typically retain malicious semantic signals that modern guardrails are primed to detect.
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpler, benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill the malicious intent.
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception.
CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents
Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing.
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories.
Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
Large Language Models (LLMs) are gaining traction as a method to generate consensus statements and aggregate preferences in digital democracy experiments. Yet, participants can introduce critical vulnerabilities in LLM-based systems.
Post-Quantum Cryptography from Quantum Stabilizer Decoding
Post-quantum cryptography currently rests on a small number of hardness assumptions, posing significant risks should any one of them be compromised. This vulnerability motivates the search for new and cryptographically versatile assumptions that make a convincing case for quantum hardness.
Routing-Aware Safety Alignment for Mixture-of-Experts Models
Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts.
