BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 11 days ago • 29
Emergent Collusion in Long-Horizon LLM Agent Interaction Paper • 2609.24967 • Published 6 days ago • 17
xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning Updated Dec 9, 2025 • 44 • 7
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 6 days ago • 130
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 9 days ago • 34
wnkh/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-RLHF-v6 Viewer • Updated Jul 10, 2025 • 100 • 40 • 3
wnkh/vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-original-only Viewer • Updated Oct 18, 2025 • 244 • 33 • 4
LangAGI-Lab/magpie-reasoning-v1-10k-step-by-step-rationale Viewer • Updated Jan 31, 2025 • 10k • 59 • 3
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 12 days ago • 788
CreitinGameplays/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken-mistral Viewer • Updated Feb 9, 2025 • 10k • 35 • 2
CreitinGameplays/magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-changedtoken Viewer • Updated Feb 9, 2025 • 10k • 42 • 4
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation Paper • 2609.13770 • Published 15 days ago • 9
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs Paper • 2609.10895 • Published 18 days ago • 54
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 17 days ago • 172
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 23 days ago • 44
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 19 days ago • 28