-
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Paper • 2504.00891 • Published • 14 -
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
Paper • 2506.09942 • Published • 5 -
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
Paper • 2505.15801 • Published • 17 -
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
Paper • 2510.04081 • Published • 23
Collections
Discover the best community collections!
Collections including paper arxiv:2512.05439
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 296 -
RAG-Anything: All-in-One RAG Framework
Paper • 2510.12323 • Published • 57 -
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Paper • 2310.11511 • Published • 78 -
SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
Paper • 2308.00436 • Published • 23
-
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Paper • 2410.02740 • Published • 54 -
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
Paper • 2410.01215 • Published • 39 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 121 -
EuroLLM: Multilingual Language Models for Europe
Paper • 2409.16235 • Published • 29
-
facebook/w2v-bert-2.0
Feature Extraction • 0.6B • Updated • 2.45M • 201 -
facebook/metaclip-h14-fullcc2.5b
Zero-Shot Image Classification • 1.0B • Updated • 22k • 49 -
openai/clip-vit-large-patch14
Zero-Shot Image Classification • 0.4B • Updated • 7.87M • 1.96k -
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 707k • 1.45k
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 138 • 98 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 38 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 97 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 88
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Paper • 2504.00891 • Published • 14 -
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
Paper • 2506.09942 • Published • 5 -
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
Paper • 2505.15801 • Published • 17 -
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
Paper • 2510.04081 • Published • 23
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 296 -
RAG-Anything: All-in-One RAG Framework
Paper • 2510.12323 • Published • 57 -
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Paper • 2310.11511 • Published • 78 -
SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
Paper • 2308.00436 • Published • 23
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 138 • 98 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 38 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 97 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 88
-
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Paper • 2410.02740 • Published • 54 -
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
Paper • 2410.01215 • Published • 39 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 121 -
EuroLLM: Multilingual Language Models for Europe
Paper • 2409.16235 • Published • 29
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
facebook/w2v-bert-2.0
Feature Extraction • 0.6B • Updated • 2.45M • 201 -
facebook/metaclip-h14-fullcc2.5b
Zero-Shot Image Classification • 1.0B • Updated • 22k • 49 -
openai/clip-vit-large-patch14
Zero-Shot Image Classification • 0.4B • Updated • 7.87M • 1.96k -
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 707k • 1.45k