GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
arXiv cs.RO ·
A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outpe…
Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation
arXiv cs.RO ·
Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geomet…
Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks
arXiv cs.CL ·
We introduce Crypto Accounting Bench (CAB), a benchmark for assessing whether frontier and open-weight language models can reconstruct the complete journal entry that an organiz…
Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
arXiv cs.CL ·
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factu…
Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
NVIDIA Technical Blog ·
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE... Mixture of expe…