Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published 6 days ago • 26
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 16 days ago • 76
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published Jul 7 • 38
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 41
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 20
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 20
BRAID Collection Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process • 1 item • Updated Jul 7 • 1
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published Jul 4 • 20
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 41
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 41
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe Paper • 2605.03677 • Published May 5 • 28
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning Paper • 2605.06326 • Published May 7 • 26