SearcharxivSearch

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 289 records · Page 16Linked to original sources

When Do Type-Specific Wages Buffer Distributional Incidence in TANK?

When do relative wages buffer the unequal incidence of aggregate shocks? I derive a consumption-gap decomposition and a present-value condition for partial offset in a TANK model. An extension separates wage-setting demand elasticity from substitution between labor segments and allows each segment to contain both financial types. With a zero inherited wage gap and a same-sign discounted wedge, substitution above one gives offsetting earnings reallocation; substitution below one gives amplification. The channel disappears when financial types have identical segment exposure. Numerical experiments assess these mechanisms, shock persistence, policy feedback, and aggregate-IRF matching. In the nested perfect-alignment monetary benchmark, the peak consumption gap is about two-fifths smaller under type-specific wages than under the common-wage closure. These are conditional model comparisons, not empirical effect estimates or welfare rankings.

econ.GN

DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo

Achieving human-level manipulation requires dexterous robotic hands capable of complex object interactions. Advancing such capabilities further demands standardized benchmarks for systematic evaluation. However, prior work does not focus on task-oriented dexterous manipulation. Existing benchmarks also often lack high quality human demonstrations and an easy-to-use evaluation pipeline. In this paper, we present DexJoCo, a benchmark and toolkit for task-oriented dexterous manipulation, comprising 11 functionally grounded tasks that evaluate tool-use, bimanual coordination, long-horizon execution, and reasoning. We develop a low-cost data collection system and collect 1.1K trajectories across these tasks, with support for domain randomization to assess robustness. We benchmark modern models under diverse settings, including visual and dynamics randomization, multi-task training, and action-head adaptation. Through extensive empirical analysis, we identify several important insights and common limitations of current policies in dexterous manipulation, highlighting key challenges for future research in dexterous hand robot learning. Code is available at: https://dexjoco.github.io

cs.RO

Functionalization via Structure Completion and Motion Rectification

Acquisition and creation of 3D assets have been largely view- or appearance-driven. As a result, existing digital 3D models often lack the requisite structural components to function as intended, such as joints, supports, interiors, or interaction elements. At the same time, even human-annotated motions are frequently error-prone, leading to physically implausible behavior. We introduce object functionalization, a novel task aimed at transforming visually plausible but non-functional 3D models into functional and physically operable ones. We formulate functionalization as a graph completion problem over a new functional graph representation, where labeled nodes represent object parts, labeled edges encode functional and contact relations, and movable nodes carry motion attributes, so that structural functional deficiencies manifest as missing nodes or incorrect edges. We develop a neural Graph Functionalizer (GraFu) to complete an incomplete graph representing a non-functional 3D object. The completed graph then drives a geometry realization stage that instantiates predicted connectors and structural elements in 3D, with the compelling side effect of rectifying erroneous human-annotated and predicted motions. To support training and evaluation, focusing on furniture as a rich and challenging target category, we introduce FurFun-233, a dataset of 233 paired non-functional and functionalized furniture models. On PartNet-Mobility ("zero-shot") and HSSD test sets, our method matches state-of-the-art methods in motion prediction accuracy while substantially improving functionality in terms of collision and connectivity. Project page: https://mingrui-zhao.github.io/Functionalization/

cs.CV

The Ring of Differential Operators on a Nodal Curve is not a Bialgebroid

In a previous article, we showed that local projectivity is a sufficient condition for the existence of a bialgebroid structure on the ring of differential operators on an affine variety. In this note, we show using elementary methods that the ring of differential operators on a nodal curve is neither locally projective nor does it admit a bialgebroid structure.

math.QA

Funnel control with input filter for nonlinear systems with arbitrary relative degree

This paper addresses output reference tracking with prescribed transient performance for unknown nonlinear multi-input multi-output systems of arbitrary relative degree. We propose a novel derivative-free extension of funnel control based on a set of auxiliary filter variables that estimate the output derivatives. The resulting controller ensures that the tracking error evolves within prescribed performance bounds, while avoiding differentiation of the output signal and maintaining a simple structure with only a small number of tuning parameters. The effectiveness of the proposed approach is validated by a numerical example.

math.OC

The quantum Almeida-Thouless line in the self-overlap-corrected quantum Sherrington-Kirkpatrick model

We present a complete analysis of the glass transition in the self-overlap-corrected Sherrington\--Kirkpatrick (SK) model in a transverse magnetic field, also referred to as the quantum SK (QSK) model. In particular, we determine the phase boundary separating the glassy and paramagnetic phases explicitly. Such an analytic characterization of the glass transition is not expected for the true QSK and even unknown for most vector glass models. Despite being of independent interest, the analysis of the self-overlap corrected QSK model serves as important ingredient in the characterization of paramagnetic behavior in the real QSK. The proof is based on a simplified Parisi variational principle for the quantum pressure, which only involves classical Parisi order parameters. As part of the proof, we also analyze the pressure of the self-overlap-constrained quantum SK model and its Parisi description, as well as the pressure of generalized quantum Hopfield models.

math-ph

A positive solution to the $L^p$ projection centroid conjecture

In a classical paper [21] in 2000, Lutwak-Yang-Zhang established the $L^p$ analog of the Petty projection inequality and the $L^p$ analog of the Busemann-Petty centroid inequality. In Section 7 of [21], Lutwak-Yang-Zhang proposed the important $L^p$ projection centroid conjecture. We give a positive solution to the $L^p$ projection centroid conjecture in this work.

math.MG

A Matrix-Theoretic Exact Formula for Counting Primes in Intervals Between Consecutive Odd Squares

Matrix $B=(b_{ij})$ with $b_{ij}=(2j+1)(2j+2i-1)$ was introduced in \cite{Shi2024} as an additive sieve for odd primes. In this paper we introduce the minimal-anchor function $\pmin(d)$, the least odd prime $p$ such that $p+d$ is prime (sequence A020483 of the OEIS at index $d/2$), whose finiteness for all even $d$ is exactly the weak Polignac (Maillet) conjecture, i.e.\ the statement that every row of $B$ contains a semiprime. We prove that for each fixed $z$ the set $\{d\ \text{even}:\pmin(d)\le z\}$ has density zero, with the asymptotic $(π(z)-1)X/\log X$; consequently no fixed finite set of anchor primes can cover a positive proportion of the rows. Density-one coverage of the rows nevertheless holds, by a classical theorem of Lavrik which we restate in the matrix-$B$ framework: almost every row contains the number of semiprimes predicted by the Hardy--Littlewood conjecture. We formulate quantitative conjectures on $\pmin$ and support them with numerical data. Every-row coverage (= weak Polignac) is explicitly left open. We prove unconditional lower bounds for $\pmin$: for every $ψ\to0$, $\pmin(d)>ψ(d)\log d\log\log d$ for almost all even $d$, which is the conjectured typical order; and $\max_{d\le X}\pmin(d)\ge(\frac12+o(1))\log X\log\log X$. The same counting gives the corresponding lower bounds for the least prime in a Goldbach partition.

math.NT

Rate-Splitting--Inspired Bistatic OFDM-ISAC

Achieving effective uplink bistatic ISAC over an OFDM waveform gives rise to challenging interference structures. These are mostly due to unequal direct- and echo-path contributions and Doppler-induced ICI, rendering orthogonal resource separation and fixed SIC strategies inadequate. To address this problem, we propose a RS-inspired framework where the transmitter splits each communication message into a robust and a supplementary stream, which are jointly superposed over a sensing signal. Furthermore, we present the design of a staged sensing-communication receiver. Based on this framework, we derive tractable per-subcarrier SINR expressions and establish the relation between sensing accuracy and communication reliability based on the Fisher information. Building on these, we formulate a joint power-allocation problem for SE maximization under sensing-performance and power constraints. The resulting non-convex formulation is solved using convex surrogates and fractional programming. Numerical results demonstrate that, compared to NOMA-inspired baselines, the proposed framework provides more effective IFI management and improved robustness to Doppler-induced ICI.

eess.SP

VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding

Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Embodied online occupancy prediction remains predominantly Gaussian-centric, despite the wide use of voxel representations for frame-wise scene completion. We present VEOcc, to the best of our knowledge, the first voxel-centric framework for online embodied semantic occupancy prediction. It incrementally maintains a sparse global semantic voxel map from monocular observations and enables open-ended expansion without predefined scene bounds. To robustly integrate noisy multi-view predictions, we further introduce a Spatio-Temporal-Aware Online Update Strategy comprising Cross-Temporal Logit Aggregation (TLA) for short-term temporal consistency, Reliability-Aware Confidence Modulation (RCM) for spatial uncertainty calibration, and Confidence-Driven Incremental State Update (CSU) for robust global state assimilation. Extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet demonstrate state-of-the-art performance among models of comparable scale in both local and embodied settings. Moreover, onboard deployment on a mobile robot validates practical online operation and long-horizon scalability, while results on self-collected handheld sequences demonstrate zero-shot generalization to unseen real-world environments. Code and supplementary visualizations are available on our project page: https://wryzju.github.io/VEOcc/.

cs.CV

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approaching saturation. This single-user assumption sidesteps what makes group planning hard for an agent: discovering private preferences across multiple users, surfacing conflicts, and balancing utility against fairness. To bring the task back to its multi-user reality, we introduce \textbf{\textit{GroupTravelBench}}, the first benchmark for \textbf{multi-user, multi-turn} travel planning. Built from real user profiles, POI data, and ticket prices, it comprises 650 tasks across three difficulty levels, each running in a synchronous group-chat sandbox with cached tool data for reproducible offline evaluation. Beyond the multi-step reasoning and tool use that single-user benchmarks already test, GroupTravelBench probes three group-specific capabilities: \textit{(i) elicitation} of private preferences through multi-turn dialogue; \textit{(ii) coordination} of inter-user conflicts via compromise or subgrouping; and \textit{(iii) planning} that balances group utility against fairness. We pair this with a complementary evaluation framework combining rule-based outcome metrics and LLM-judge process metrics. Across a wide range of frontier models, even the strongest agents fall short on all four rule-based outcome metrics, with plan validity below 12\%, suggesting that group-level outcome quality is a key open challenge for LLM travel-planning agents.

cs.CL

Fixed Point Rigidity of the Operator $Γ_pΠ_p^\ast$ and the LYZ Conjecture

We characterize the fixed points of the operator $Γ_pΠ_p^\ast$ for $n\geq 3$ and $1 0$ if and only if $K$ is an origin-centered ellipsoid, thereby settling the Lutwak--Yang--Zhang fixed-point conjecture in this range. Our proof is based on a variational analysis along linear reflection shadow systems. To address the nonlinear structure of the $L_p$ setting, we introduce the $L_p$-Projection Rolodex, which provides a dimensional reduction of the volume of the polar $L_p$-projection body to weighted lower-dimensional sectional functionals. A suitable change of variables, together with Ball's harmonic Prékopa--Leindler inequality, yields the convexity needed along the shadow system. Under the fixed-point condition, a first-variation identity then forces $\operatorname{vol}_n(Π_p^\ast K_t)$ to remain constant throughout the deformation. The rigidity statement follows from the equality characterization under Steiner symmetrization.

math.FA

DDGAD: Disagreement-Driven Graph Anomaly Detection via Adapt-Then-Combine

Graph anomaly detection (GAD) commonly relies on message passing to jointly encode node attributes and neighborhood context. However, once the two are mixed, an abnormal post-encoding state may reflect either an intrinsic node deviation or incompatible contextual influence, making its source ambiguous. We propose Disagreement-Driven Graph Anomaly Detection (DDGAD), which treats persistent incompatibility between node-wise and contextual estimates as anomaly evidence. Inspired by Adapt-Then-Combine (ATC), DDGAD reverses its consensus objective: Adapt produces a node-wise estimate without new same-step neighbor aggregation, while Combine forms a neighborhood-dependent contextual estimate, and their pre-consensus disagreement is accumulated across ATC steps for detection. We further characterize this signal from graph-spectral and source-response perspectives and derive sufficient conditions for anomaly--normal separation under contextual mixing. Experiments on six benchmarks show the highest average AUROC among the evaluated methods, while controlled interventions and Adapt-operator controls further support persistent disagreement as an effective detection signal.

cs.LG

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term introduces curvature-sensitive corrections while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.

cs.LG

A reparametrization invariant nonabelian surface holonomy

We introduce a nonabelian surface holonomy that is constructed from a one-form gauge potential that takes values in a loop algebra of the $U(N)$ gauge group. The surface holonomy parallel transports a nonabelian string. Although it is not manifest in our formulation, we will see that our nonabelian surface holonomy is invariant under reparametrizations of the surface.

hep-th

Logarithmic oscillatory multipliers and log-subdyadic square functions

We develop square-function estimates for Fourier multipliers whose local oscillation scale is \[ ρ(R)=\frac{R}{(\log R)^{γ-1}}, \qquad γ>1. \] This scale lies strictly between the dyadic scale and every fixed power-subdyadic scale at high frequency. For high-frequency symbols satisfying a localized Sobolev condition on balls of radius comparable to $ρ(R)$, we prove a pointwise square-function estimate and a weighted $L^2$ multiplier inequality. After adjoining a smooth compactly supported low-frequency part, we derive unweighted $L^p$ bounds. The weighted estimate is governed by a logarithmic geometric maximal operator which is strongly bounded above the critical $L^r$ threshold, satisfies weak type at the critical equality, and fails even weak type below it. As a model application, consider \[ L(ξ)=\frac12\log(e^2+|ξ|^2), \qquad m_{γ,β}(ξ)=L(ξ)^{-β}e^{iL(ξ)^γ}. \] For $p=2$, the associated multiplier is bounded on $L^2$ for every $β\geq0$. For $1 d(γ-1)\left|\frac12-\frac1p\right|. \] At the critical equality we obtain the corresponding Lorentz endpoint estimates.

math.FA

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over single agent systems, they lead to huge computational overheads because of heavy communication between agents. Previous research has made efforts to train a sparse multi-agent graph or fine-tune a planner to orchestrate the workflow better. However, such extra training processes introduce computational costs and limit MAS to specific domains, therefore compromising their generalizability. In this paper, we propose CONCAT, a training-free multi-agent collaboration framework based on CONsensus and Confidence-driven Ad hoc Teaming to efficiently organize agent interactions. Specifically, agents are clustered based on their initial answers, and leaders of each cluster are selected based on the agents' confidence. Then, a heuristic function based on the Theory of Mind is designed to predict the collaboration benefits between every two leaders according to their answers and confidence. Finally, an ad hoc multi-agent network is organized after evicting a percentage of communications based on the predicted benefits. Experiments across three LLMs and three benchmarks show that CONCAT achieves up to 2.02x higher efficiency (accuracy/latency ratio) than LLM-Debate and outperforms training-aware methods such as AgentDropout, while reducing average latency by 50.1% on Qwen2.5-14B-Instruct, without any task-specific training.

cs.MA

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models require controllable, causal, and low-latency rollout, which in practice demands a full pipeline spanning data construction, controllable fine-tuning, autoregressive training, few-step distillation, and streaming inference. In this work, we present minWM, a full-stack open-source framework for building real-time interactive video world models. minWM provides an end-to-end pipeline that converts existing bidirectional T2V/TI2V video foundation models into camera-controllable few-step autoregressive world models. Specifically, minWM first fine-tunes a bidirectional video diffusion model with camera control, and then applies the Causal Forcing / Causal Forcing++ pipeline, including AR diffusion training, causal ODE or causal consistency distillation, and asymmetric DMD, to distill it into a few-step autoregressive generator for low-latency rollout. The framework is modular and architecture-extensible: we instantiate it on representative open backbones, including Wan2.1-T2V-1.3B and HY1.5-TI2V-8B, covering both cross-attention-based condition injection and MMDiT-style architectures. minWM also supports adapting existing video world models, such as HY-WorldPlay, to new data distributions, training recipes, and latency targets. Beyond releasing runnable scripts, checkpoints, documentation, and inference code, we provide practical ablations on camera trajectory quality, controllability training steps, and minimal batch-size requirements. We hope minWM serves as a reproducible and extensible recipe for building and adapting real-time interactive video world models. Project Page: [https://github.com/shengshu-ai/minWM](https://github.com/shengshu-ai/minWM)

cs.CV