OpenAI 上线对齐失效报告网站,披露九起智能体失控事件
AIHOT ·
OpenAI 上线专门发布对齐失效报告的新网站,截至目前公布九起事件,大多数发生在强化学习训练阶段。
原文内容 · 暂无该语言译文按采集日期(UTC)汇总已发布资讯,原文日期单独标注。
返回资讯列表AIHOT ·
OpenAI 上线专门发布对齐失效报告的新网站,截至目前公布九起事件,大多数发生在强化学习训练阶段。
原文内容 · 暂无该语言译文Simon Willison ·
Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to…
原文内容 · 暂无该语言译文AIHOT ·
GitHub Security Lab 发布开源 seclab-taskflows 任务流,通过 gather_mobile_entry_point_info.yaml 和 classify_application_local.yaml 等提示词引导 LLM 审计 Android 应用,已发现并报告 24 个漏洞。
原文内容 · 暂无该语言译文AIHOT ·
Anthropic 发布 Claude Sonnet 5.5,在 Artificial Analysis Intelligence Index 得 56 分,仅比 Opus 5.5(max)低 2 分,max effort 下比 Sonnet 5 高 18 分。
原文内容 · 暂无该语言译文AIHOT ·
Perplexity 安全团队对运行 Perplexity Computer 的沙箱平台 SPACE 做了一个月红队测试,给 Opus 5、GPT-5.6 Sol、Kimi K3、Gemini 3.1 Pro 等 9 个模型 VM 内 root 权限,108 次运行中无一逃逸 VM 边界。
原文内容 · 暂无该语言译文arXiv cs.RO ·
Vision-language-action (VLA) policies can execute many tasks from standard initial states, yet a new instruction may fail after another task has altered the robot's physical sta…
原文内容 · 暂无该语言译文arXiv cs.CL ·
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, whil…
原文内容 · 暂无该语言译文arXiv cs.RO ·
In recent years, Human-Robot Collaboration (HRC) has taken on a central role in Industry 4.0 and collaborative robotics, demanding communication channels that are increasingly b…
原文内容 · 暂无该语言译文arXiv cs.CL ·
When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation ord…
原文内容 · 暂无该语言译文AIHOT ·
NVIDIA 推出 NVIDIA Open Agent Safety Platform,包含开源运行时 OpenShell、硬件层 Sentry 及 DOCA 技术。该平台旨在通过内核级隔离、零信任环境和带外(out-of-band)监控来防止 AI 智能体行为漂移和越权。OpenShell 将操作指令转化为可验证策略,BlueField DPU 在模…
原文内容 · 暂无该语言译文