SearcharxivSearch

arXiv · 2511.17653

MARL-CC: A Mathematical Framework forMulti-Agent Reinforcement Learning in ConnectedAutonomous Vehicles: Addressing Nonlinearity,Partial Observability, and Credit Assignment forOptimal Control

Abstract

Multi-Agent Reinforcement Learning (MARL) has emerged as a powerfulparadigm for cooperative decision-making in connected autonomous vehicles(CAVs); however, existing approaches often fail to guarantee stability, optimality,and interpretability in systems characterized by nonlinear dynamics,partial observability, and complex inter-agent coupling. This study addressesthese foundational challenges by introducing MARL-CC, a unified MathematicalFramework for Multi-Agent Reinforcement Learning with Control Coordination.The proposed framework integrates differential geometric control, Bayesian inference,and Shapley-value-based credit assignment within a coherent optimizationarchitecture, ensuring bounded policy updates, decentralized belief estimation,and equitable reward distribution. Theoretical analyses establish convergence andstability guarantees under stochastic disturbances and communication delays.Empirical evaluations across simulation and real-world testbeds demonstrate upto a 40% improvement in convergence rate and enhanced cooperative efficiencyover leading baselines, including PPO, DDPG, and QMIX.These results signify a decisive advance in control-oriented reinforcement learning,bridging the gap between mathematical rigor and practical autonomy.The MARL-CC framework provides a scalable foundation for intelligent transportation,UAV coordination, and distributed robotics, paving the way toward interpretable, safe, and adaptive multi-agent systems. All codes and experimentalconfigurations are publicly available on GitHub to support reproducibilityand future research.

Explore related subjects

Keep this discovery

BibTeXRIS

Mazyar Taghavi, Javad Vahidi. 2025-11-20. MARL-CC: A Mathematical Framework forMulti-Agent Reinforcement Learning in ConnectedAutonomous Vehicles: Addressing Nonlinearity,Partial Observability, and Credit Assignment forOptimal Control. https://doi.org/10.21203/rs.3.rs-7996305%2Fv1

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Average Chord Lengths in a Triangle

Let $P$ be a point inside a triangle $T$. We consider the average length of the chords of $T$ through $P$, where the direction of the chord is chosen uniformly. An elementary formula is obtained in terms of the distances from $P$ to the sides and vertices of the triangle. Several classical triangle centers give especially simple specializations. For example, if $I$ is the incenter, then \[ M_T(I)=\frac{2r}{\pi} \log\left(\cot\frac A4\cot\frac B4\cot\frac C4\right). \] Our main result is the sharp inequality \[ M_T(P)\le \frac{p}{\pi\sqrt3}\log(2+\sqrt3), \] valid simultaneously for every triangle of perimeter $p$ and every interior point $P$. Thus, among all such pairs $(T,P)$, the largest possible average chord length occurs only when $T$ is equilateral and $P$ is its center. The proof is an elementary symmetrization argument. We close with brief remarks relating the problem to the radial center of a convex body, the electrostatic potential center of a triangle, and dual quermassintegrals.

math.GM

A Proof of Liu's Conjecture on the Fundamental Triangle Inequality

Let $a,b,c$ be the side lengths of a triangle, and let $R$ and $r$ denote its circumradius and inradius, respectively. We prove a conjecture of Liu stating that \[\sum_{\mathrm{cyc}} \left(\frac{a(b+c-a)}{bc}\right)^k \geq 2+\left(\frac{2r}{R}\right)^k,~~k>1, \] with the reverse inequality for $0<k<1$. The proof reduces the problem to three positive variables with fixed sum and product. We also determine the equality cases.

math.GM