arXiv · 2603.18363
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Abstract
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting $α$-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution ($α> 1$) to intensify logical reasoning, or flattening it ($α< 1$) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang. 2026-05-25. PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching. https://arxiv.org/abs/2603.18363
Cite the original work for its findings. Save a collection to share your selection of sources.