Live
Black Hat USADark ReadingBlack Hat AsiaAI BusinessAnthropic laat klanten extra betalen als ze Claude via OpenClaw willen gebruikenTweakers.netHackers Are Posting the Claude Code Leak With Bonus MalwareWired AIUnpacking the True Cost of Blockchain Indexing: More Than Just InfrastructureDEV CommunityThe coordinate space bug that four rewrites couldn't fixDEV CommunityThe Programmer's Fulcrum: 03 April, 2026DEV CommunityEnthusiast installs Win 3.1X on bare metal Ryzen 9 9900X and RTX 5060 Ti system using floppy disk drive — Asus motherboard’s ‘classic BIOS’ functionality was instrumental to the feattomshardware.comI Put VS Code, Claude, and a Terminal Inside a File Manager I built using React and Rust — Here's What HappenedDEV CommunityLooking for arXiv endorsement (cs.LG) – RL fine-tuning for VLMs (GRPO, MathVista)discuss.huggingface.coClaude Code at Enterprise Scale: Why You Need an AI GatewayDEV CommunityPowering Down Enterprises Tackle AI’s Soaring Energy CostsDev.to AIIs Micron the New Nvidia? - The Motley FoolGNews AI NVIDIAFrom Guesswork to Growth: AI-Driven Analytics for Grant WritingDev.to AIBlack Hat USADark ReadingBlack Hat AsiaAI BusinessAnthropic laat klanten extra betalen als ze Claude via OpenClaw willen gebruikenTweakers.netHackers Are Posting the Claude Code Leak With Bonus MalwareWired AIUnpacking the True Cost of Blockchain Indexing: More Than Just InfrastructureDEV CommunityThe coordinate space bug that four rewrites couldn't fixDEV CommunityThe Programmer's Fulcrum: 03 April, 2026DEV CommunityEnthusiast installs Win 3.1X on bare metal Ryzen 9 9900X and RTX 5060 Ti system using floppy disk drive — Asus motherboard’s ‘classic BIOS’ functionality was instrumental to the feattomshardware.comI Put VS Code, Claude, and a Terminal Inside a File Manager I built using React and Rust — Here's What HappenedDEV CommunityLooking for arXiv endorsement (cs.LG) – RL fine-tuning for VLMs (GRPO, MathVista)discuss.huggingface.coClaude Code at Enterprise Scale: Why You Need an AI GatewayDEV CommunityPowering Down Enterprises Tackle AI’s Soaring Energy CostsDev.to AIIs Micron the New Nvidia? - The Motley FoolGNews AI NVIDIAFrom Guesswork to Growth: AI-Driven Analytics for Grant WritingDev.to AI
AI NEWS HUBbyEIGENVECTOREigenvector

BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional Environments

arXivby [Submitted on 30 Jan 2026]March 30, 20262 min read2 views
Source Quiz
🧒Explain Like I'm 5Simple language

Hi there, little explorer! 🚀

Imagine you have a super-duper smart toy robot, like a helpful friend. This robot can do many things, like play games on a tablet or even tidy your room!

But sometimes, even smart robots can make funny mistakes, right? Like maybe it puts your socks in the fridge by accident! 😅

Scientists made a special game called BeSafe-Bench. It's like a big playground for these smart robots. They watch the robots play to see if they always do things safely and correctly.

Guess what? Many robots still make silly mistakes or do things that are not safe, even when they try to be helpful! So, the scientists are working hard to teach them to be super-duper safe helpers for everyone! 👍🤖

arXiv:2603.25747v1 Announce Type: new Abstract: The rapid evolution of Large Multimodal Models (LMMs) has enabled agents to perform complex digital and physical tasks, yet their deployment as autonomous decision-makers introduces substantial unintentional behavioral safety risks. However, the absence of a comprehensive safety benchmark remains a major bottleneck, as existing evaluations rely on low-fidelity environments, simulated APIs, or narrowly scoped tasks. To address this gap, we present BeSafe-Bench (BSB), a benchmark for exposing behavioral safety risks of situated agents in functional — Yuxuan Li, Yi Lin, Peng Wang, Shiming Liu, Xuetao Wei

View PDF HTML (experimental)

Abstract:The rapid evolution of Large Multimodal Models (LMMs) has enabled agents to perform complex digital and physical tasks, yet their deployment as autonomous decision-makers introduces substantial unintentional behavioral safety risks. However, the absence of a comprehensive safety benchmark remains a major bottleneck, as existing evaluations rely on low-fidelity environments, simulated APIs, or narrowly scoped tasks. To address this gap, we present BeSafe-Bench (BSB), a benchmark for exposing behavioral safety risks of situated agents in functional environments, covering four representative domains: Web, Mobile, Embodied VLM, and Embodied VLA. Using functional environments, we construct a diverse instruction space by augmenting tasks with nine categories of safety-critical risks, and adopt a hybrid evaluation framework that combines rule-based checks with LLM-as-a-judge reasoning to assess real environmental impacts. Evaluating 13 popular agents reveals a concerning trend: even the best-performing agent completes fewer than 40% of tasks while fully adhering to safety constraints, and strong task performance frequently coincides with severe safety violations. These findings underscore the urgent need for improved safety alignment before deploying agentic systems in real-world settings.

Subjects:

Artificial Intelligence (cs.AI)

Cite as: arXiv:2603.25747 [cs.AI]

(or arXiv:2603.25747v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2603.25747

arXiv-issued DOI via DataCite

Submission history

From: Xuetao Wei [view email] [v1] Fri, 30 Jan 2026 03:41:57 UTC (1,705 KB)

Was this article helpful?

Sign in to highlight and annotate this article

AI
Ask AI about this article
Powered by Eigenvector · full article context loaded
Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Knowledge Map

Knowledge Map
TopicsEntitiesSource
BeSafe-Benc…researchpaperarxivaiartificial-…arXiv

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 177 connections
Scroll to zoom · drag to pan · click to open

Discussion

Sign in to join the discussion

No comments yet — be the first to share your thoughts!

More in Research Papers