Search arXivSearch

arXiv subjects

Dongming Wang

Publications and source records attributed to Dongming Wang.

3 recordsLinked to original sources

Distributed primal-dual algorithm for constrained multi-agent reinforcement learning under coupled policies

This paper investigates constrained multi-agent reinforcement learning (CMARL) in coupled environments, where agents collaboratively maximize the sum of local objectives while satisfying individual safety constraints. Existing studies face two limitations: (1) most rely on independent policies that fail to capture complex interactions in coupled environments; and (2) agents require the global Lagrange multipliers, which are sensitive learned variables whose global sharing risks exposing private agent-specific information. To overcome these issues, we propose a framework where agents adopt coupled policies that depend on both the states and policy parameters of their $κ_p$-hop neighbors, where $κ_p>0$ denotes the coupling distance, and develop a distributed and scalable primal-dual (DSPD) algorithm wherein each agent accesses only information within a prescribed local neighborhood. In the proposed algorithm, agents exchange sensitive parameters only with immediate neighbors over a separate time-varying network, while maintaining local estimates to execute the coupled policy. We establish that the proposed algorithm achieves $ε$-policy stationary convergence with approximation error $\mathcal{O}(γ^{\frac{κ+1}{κ_{p}}})$, where $κ>0$ is the truncated distance and $γ\in(0,1)$ is discount factor. Simulations on a wireless access-control network demonstrate that the proposed algorithm outperforms existing state-of-the-art algorithms, validating its effectiveness.

cs.MA

Geodesic strong convexity does not imply forward invariance under gradient flow on SO(3): a certified counterexample

Let $mathcal{C}=\overline{\mathcal{B}}_ρ(R_c)$ be a geodesic ball of radius $ρ<π/2$ in SO(3) with the bi-invariant metric, and let $f$ be geodesically strongly convex on $\mathcal{C}$ with an interior minimizer. It is tempting to expect the gradient flow $\dot R=R(-\nabla f)^\wedge$ to keep $\mathcal{C}$ forward invariant: the flow is attracted to an interior point, and strong convexity appears to leave no room for outward motion. We show this expectation is false by an explicit, fully certified construction with $ρ=0.3$: a cost, quadratic in the principal logarithmic chart with off-diagonal coupling $0.7$, whose geodesic Hessian satisfies $\Hess f\succeqμI_3$ on all of $\mathcal{C}$ with a machine-certified modulus $μ\geq0.172$, rigorous ball arithmetic over exact rational inputs, yet whose descent velocity at a boundary point has the exact rational outward radial component $21/500$. A continuity corollary of the exact rate certifies that the flow exits the ball; numerical integration puts the peak excursion near $0.3143$ before convergence to the minimizer. The mechanism is elementary: strong convexity constrains the projection of the gradient onto the minimizer direction, not onto the inward radial direction. Code reproducing every certified constant and figure accompanies the note.

eess.SY

Securing Cooperative Sensing in UAV Swarms Against Conformity-Driven Byzantine Attacks

In integrated sensing and communication (ISAC) enabled 6G unmanned aerial vehicle (UAV) swarm networks, the widely adopted imitation-based conformity cooperation mechanism can be exploited by Byzantine attackers to fabricate false consensus, causing the effective error probability of normal UAVs to evolve dynamically and far exceed their inherent sensing errors, which invalidates conventional fusion methods built on the independence assumption. This paper proposes a conformity-aware Byzantine-resilient fusion framework that couples evolutionary game theory with maximum a posteriori (MAP) estimation. First, the strategy updates of normal UAVs are characterized by bounded-rational opinion dynamics, and the evolution dynamics of the misinformation ratio together with its evolutionarily stable state (ESS) are derived under death birth updating. Three theoretical results are then established: under heterogeneous per-node sensing errors, the zeroth-order ESS depends on the error distribution only through its mean; a closed-form first-order weak-selection correction to the ESS is obtained, together with an exact mean-field fixed point valid for arbitrary selection intensity; and it is revealed that swarm level misinformation can overwhelm the majority if and only if the attack probability exceeds one half, with this threshold independent of both the sensing error and the malicious ratio. Embedding the predicted error dynamics into a per-node MAP rule, the resulting fusion mechanism achieves nearly 100% situation-inference accuracy under different network topologies, attack intensities, network scales, and sensing-error distributions, and maintains accuracy above 99% under +-20% parameter mismatch. In contrast, majority voting, reputation weighting, and independent fusion collapse completely once the majority-flip threshold is crossed.

eess.SY