-
More Agents Is All You Need
Paper • 2402.05120 • Published • 59 -
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Paper • 2608.04828 • Published • 1 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 148 -
FrontierChallenge: Evaluating Scientific Workflow Completion
Paper • 2608.24979 • Published • 140
Collections
Discover the best community collections!
Collections including paper arxiv:2608.01964
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Paper • 2608.11924 • Published • 289 -
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Paper • 2608.13417 • Published • 57 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
Mental World Modeling
Paper • 2607.27201 • Published • 108 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 100 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 60
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
DAPD: Dual-Anchored Policy Distillation
Paper • 2608.01735 • Published • 153 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 60 -
UEmbed: Unified Sparse and Dense Multimodal Embeddings
Paper • 2608.02583 • Published • 51
-
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762 -
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Paper • 2608.14391 • Published • 280 -
EnvHarness: Awakening Static Worlds for Agent Learning
Paper • 2608.19880 • Published • 270 -
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Paper • 2608.00677 • Published • 263
-
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Paper • 2607.00341 • Published -
J-CoT: Chain-of-Thought in J-Space
Paper • 2607.21981 • Published • 1 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Paper • 2607.12395 • Published • 100 -
Meshy T2: Fast Native Mesh Generation with Flow Matching
Paper • 2607.28675 • Published • 59 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183
-
More Agents Is All You Need
Paper • 2402.05120 • Published • 59 -
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Paper • 2608.04828 • Published • 1 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 148 -
FrontierChallenge: Evaluating Scientific Workflow Completion
Paper • 2608.24979 • Published • 140
-
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762 -
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Paper • 2608.14391 • Published • 280 -
EnvHarness: Awakening Static Worlds for Agent Learning
Paper • 2608.19880 • Published • 270 -
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Paper • 2608.00677 • Published • 263
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Paper • 2608.11924 • Published • 289 -
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Paper • 2608.13417 • Published • 57 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762
-
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Paper • 2607.00341 • Published -
J-CoT: Chain-of-Thought in J-Space
Paper • 2607.21981 • Published • 1 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 762
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
Mental World Modeling
Paper • 2607.27201 • Published • 108 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 100 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 60
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183 -
DAPD: Dual-Anchored Policy Distillation
Paper • 2608.01735 • Published • 153 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 60 -
UEmbed: Unified Sparse and Dense Multimodal Embeddings
Paper • 2608.02583 • Published • 51
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Paper • 2607.12395 • Published • 100 -
Meshy T2: Fast Native Mesh Generation with Flow Matching
Paper • 2607.28675 • Published • 59 -
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 183