New research on AI and security, selected for relevance to language models, AI agents and cybersecurity. Summaries are the authors’ or publisher’s own words, linked to the original.
Peer-reviewed
Phylogeny-agnostic strain-level prediction of phage–host interactions from genomes using machine learning
Exploring out-of-distribution detection for sparse-view computed tomography with diffusion models
AI can widen science — but only if institutions stop rewarding the already measurable
ConTP reshapes transporter functional space to resolve substrate specificity beyond evolutionary proximity
New preprints (not yet peer-reviewed)
CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance
Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery, exploitation, and security analysis tasks, leaving the evaluation of attack chain provenance under realistic security logs insufficiently studied.
Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration
LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem.
ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt injection to inter-agent collusion.
AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents
Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user.
Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination
Large language models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, yet their ability to reliably execute distributed coordination protocols remains poorly understood. While AgentsNet, a benchmark framework for distributed coordination among LLM agents, enables such coordination, granting full reasoning autonomy often leads to inconsistent or degraded performance in complex domains.
A Benchmark for LLM's Understanding of Middle School and High School Science Topics
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula.
CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows
Crystal plasticity (CP) simulations predict the mechanical behavior of polycrystalline metals, yet their routine use is hindered by the manual effort of configuring heterogeneous tools, orchestrating multi-step data pipelines, and calibrating constitutive parameters against experiments. These bottlenecks impede productivity in systematic parameter studies, motivating interest in automated workflows.
SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing
Deep learning has improved inertial measurement unit (IMU) sensing for mobile and wearable applications. However, an IMU model trained in one domain often becomes unreliable when it is used with a new user, device, or body position.
