KABI
dongguanting
AI & ML interests
Reasoning and Alignment for Large Language Models
Recent Activity
updated a collection about 9 hours ago
Agent-Reflex updated a collection about 9 hours ago
Agent-Reflex updated a collection about 9 hours ago
Agent-ReflexOrganizations
AEPO
The official datasets and model checkpoints of AEPO
-
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 109 -
dongguanting/Qwen3-8B-AEPO-DeepSearch
Text Generation • 8B • Updated • 23 • • 2 -
dongguanting/QwQ-32B-AEPO-DeepSearch
Text Generation • 33B • Updated • 25 • 2 -
dongguanting/Qwen3-14B-AEPO-DeepSearch
Robotics • 15B • Updated • 17 • 1
Tool-Star
Tool-Star is a reinforcement learning-based framework designed to empower LLMs to autonomously invoke multiple external tools during stepwise reasonin
-
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Paper • 2505.16410 • Published • 60 -
dongguanting/Tool-Star-SFT-54K
Viewer • Updated • 54k • 242 • 11 -
dongguanting/Multi-Tool-RL-10K
Viewer • Updated • 10k • 62 • 5 -
dongguanting/Tool-Star-Qwen-7B
Text Generation • 8B • Updated • 191 • 2
Agent-Reflex
AEPO
The official datasets and model checkpoints of AEPO
-
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 109 -
dongguanting/Qwen3-8B-AEPO-DeepSearch
Text Generation • 8B • Updated • 23 • • 2 -
dongguanting/QwQ-32B-AEPO-DeepSearch
Text Generation • 33B • Updated • 25 • 2 -
dongguanting/Qwen3-14B-AEPO-DeepSearch
Robotics • 15B • Updated • 17 • 1
ARPO
The official datasets and model checkpoints of ARPO
Tool-Star
Tool-Star is a reinforcement learning-based framework designed to empower LLMs to autonomously invoke multiple external tools during stepwise reasonin
-
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Paper • 2505.16410 • Published • 60 -
dongguanting/Tool-Star-SFT-54K
Viewer • Updated • 54k • 242 • 11 -
dongguanting/Multi-Tool-RL-10K
Viewer • Updated • 10k • 62 • 5 -
dongguanting/Tool-Star-Qwen-7B
Text Generation • 8B • Updated • 191 • 2
RAG-Critic