Correct Reasoning Paths Visit Shared Decision Pivots AI updates on arXiv.org

_ October 28, 2025_ Tech Jacks Solutions_ 0 Comments

arXiv:2509.21549v2 Announce Type: replace
Abstract: Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we introduce the idea of decision pivots-minimal, verifiable checkpoints that any correct reasoning path must visit. We hypothesize that correct reasoning, though stylistically diverse, converge on the same pivot set, while incorrect ones violate at least one pivot. Leveraging this property, we propose a self-training pipeline that (i) samples diverse reasoning paths and mines shared decision pivots, (ii) compresses each trace into pivot-focused short-path reasoning using an auxiliary verifier, and (iii) post-trains the model using its self-generated outputs. The proposed method aligns reasoning without ground truth reasoning data or external metrics. Experiments on standard benchmarks such as LogiQA, MedQA, and MATH500 show the effectiveness of our method. Read More

Author

Gallery

Contacts

Correct Reasoning Paths Visit Shared Decision Pivots AI updates on arXiv.org

Tech Jacks Solutions

Leave a comment Cancel reply

Our Address

Our Mailbox

Our Phone

Gallery

Contacts

Correct Reasoning Paths Visit Shared Decision Pivots AI updates on arXiv.org

Tech Jacks Solutions

Can Large Language Models Unlock Novel Scientific Research Ideas? AI updates on arXiv.org

Faster Reinforcement Learning by Freezing Slow States AI updates on arXiv.org

Leave a comment Cancel reply

Our Address

Our Mailbox

Our Phone