Search arXiv⌕ Search

arXiv · 2610.04307

ML-OPF-Bench: Benchmarking Machine Learning for Optimal Power Flow

Abstract

Machine Learning (ML) methods promise a fast solution process for Optimal Power Flow (OPF). While inconsistent test cases, implementations, and evaluation metrics across existing studies make it challenging to determine which algorithmic advances are most critical for real-world deployment. To this end, we propose ML-OPF-Bench, a unified benchmark for AC- and DC-OPF that evaluates representative ML algorithms under a consistent pipeline, stress-tests them across system sizes, distribution shifts, and resource budgets, and ranks them with a multi-objective framework. We find that prediction accuracy alone is not a reliable indicator of operational feasibility. Under heavily loaded, congested conditions, even the strongest in-distribution performers lose their advantage, while feasibility is maintained largely by post-processing that enforces the target constraints rather than by the underlying pure ML predictor. Data scaling shows that prediction accuracy and constraint violations follow different trajectories, whereas compute scaling shows that returns diminish and that larger models do not consistently perform better. These results expose critical trade-offs among ML methods' speed, accuracy, and feasibility, and offer practical guidance for future ML-OPF design. We open-source the benchmark as an extensible Python package for integrating new learning-based OPF algorithms and evaluating them under the same standard as the existing baselines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xinyi Liu, Xuan He, Danny H. K. Tsang, Yize Chen. 2026-10-03. ML-OPF-Bench: Benchmarking Machine Learning for Optimal Power Flow. https://arxiv.org/abs/2610.04307

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model

Despite advances in policy pretraining, embodied AI systems can plateau during task-specific fine-tuning as uniform scenario collection encounters fewer of the remaining failures. We address this problem with a policy-level recursive self-improvement loop: execution outcomes train a criticality world model, whose risk scores guide scenario collection for specialist training. To account for the resulting shift in scenario frequencies, weighted resampling approximately corrects the collection bias. Because specialized updates can weaken nominal behavior, a risk-based gate selects between the specialist and a frozen nominal policy at inference. We extend this procedure to reinforcement learning and behavior cloning. A sampling analysis characterizes the correction, while a controlled study examines the trade-off between critical and nominal coverage. Across five embodied domains, the composed systems reduce failure rates by 59--67% for locomotion and manipulation and 8--25% for VLA benchmarks relative to their respective baselines.

eess.SY↗

Closed-Loop Refinement and Execution for Learned Driving Planners

Learning-based driving planners are usually trained and evaluated in open loop against logged trajectories. In closed loop, a trajectory with small displacement error can still stall the vehicle, steer it into a conflict with surrounding agents, or be executed with abrupt braking. We introduce Closed-Loop Refinement and Execution (CLRE), a hierarchical receding-horizon control framework designed to mitigate these failure modes while leaving the upstream planner frozen and adding no new learned model. The upper layer treats the nominal trajectory as a reference and solves a finite-horizon optimal control problem that trades route progress against interaction with predicted agents. Solving it from several initializations gives a candidate set, and a prediction-conditioned oriented-bounding-box (OBB) feasibility test retains only candidates whose minimum predicted OBB clearance over the horizon meets a threshold. The lower layer executes the lowest-cost survivor, or a route-centerline backup when none remains, through the tracking controller supplied with the planner, augmented by a range-based speed bound and a saturated proportional braking law. In closed-loop simulation on the 220-route Bench2Drive validation set with VAD as the upstream planner, CLRE raises the driving score from 42.26 to 55.89 and route completion from 55.69 to 71.51, and reduces collision events from 117 to 97.

eess.SY↗

A Dynamic Generalized Kalman Consensus Filter for Switching Sensor Networks

Distributed state estimation is critical for applications such as surveillance, autonomous navigation, and wide-area monitoring, where sensor agents must cooperatively track targets using only local measurements and neighbor-to-neighbor communication. Existing distributed filters have been shown to achieve accurate estimation even under sparse inter-agent communication and limited sensing ranges. However, many of these methods rely on consensus parameters that depend on global properties of the communication graph, such as the maximum degree of the graph, and are therefore sensitive to changes in network topology. This limitation is particularly significant in sensor networks with mobile agents, where communication links change over time. This paper presents a Dynamic Generalized Kalman Consensus Filter for target tracking in sensor networks with switching communication topologies. The proposed algorithm computes information-based consensus weights using only locally available quantities, eliminating the need for global network parameters. Numerical simulations demonstrate that the proposed algorithm maintains estimation accuracy under switching network topologies and outperforms existing distributed filters in the given tracking problem.

eess.SY↗