AI News · 2026-09-15

按采集日期(UTC)汇总已发布资讯,原文日期单独标注。

返回资讯列表

GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures

arXiv cs.RO ·

A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outpe…

原文内容 · 暂无该语言译文

阅读原文 ↗

Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation

arXiv cs.RO ·

Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geomet…

原文内容 · 暂无该语言译文

阅读原文 ↗

Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks

arXiv cs.CL ·

We introduce Crypto Accounting Bench (CAB), a benchmark for assessing whether frontier and open-weight language models can reconstruct the complete journal entry that an organiz…

原文内容 · 暂无该语言译文

阅读原文 ↗

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

arXiv cs.CL ·

Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factu…

原文内容 · 暂无该语言译文

阅读原文 ↗

Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine

NVIDIA Technical Blog ·

Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE... Mixture of expe…

原文内容 · 暂无该语言译文

阅读原文 ↗

How Fyxer built an AI executive assistant people trust

OpenAI ·

Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.

原文内容 · 暂无该语言译文

阅读原文 ↗