New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.
Peer-reviewed
Towards foundation-style models for energy-frontier heterogeneous neutrino detectors via self-supervised pretraining
AI-powered medical devices must be tested in real-world settings
AI-powered medical devices must be tested in real-world settings
New preprints (not yet peer-reviewed)
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning.
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Most existing LLM safety evaluation and defense methods are static: jailbreak vulnerabilities are assessed with fixed attacks, and guardrails are trained on fixed malicious-prompt datasets. In practice, adversaries continually evolve and expand the attack space.
Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to AI agents. By fine-tuning on simulated prompt injections, SecAlign, a leading open defense, reports LLMs with good test-time robustness and negligible benign utility drop.
PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents
While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user.
Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery
Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it.
PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
In this paper, we present PoisonCap: scalable temporal safety with strict use-after-free protection and initialisation safety for CHERI systems. Efficient memory safety is an increasing priority for programming languages, operating systems, and hardware designs, and CHERI is a leading hardware/software system that provides native spatial safety and a foundation for temporal memory safety.
Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money
AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it.
CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence
While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service restarts and host reboots.
