SearcharxivSearch

arXiv subjects

Gao Wang

Publications and source records attributed to Gao Wang.

At least 19 recordsLinked to original sources

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.

cs.CV

Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms

Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It has broad applications in autonomous driving, service robotics, warehouse logistics, human-robot collaboration, crowd navigation, and multi-robot systems. This survey reviews representative works published primarily between 2015 and 2025, with a particular focus on how recent learning-based advances extend, complement, or interact with classical planning foundations. We first revisit classical planning methods as algorithmic foundations and reference frameworks for learning-based extensions. We then propose a role-of-learning taxonomy that categorizes existing methods according to how learning participates in the planning pipeline, including direct policy learning, learning-augmented classical planning, hybrid planning, and training enhancement methods. For each category, we summarize the main problem settings, representative algorithms, key ideas, integration mechanisms, strengths, and limitations. We further analyze how observation representations, prediction uncertainty, interaction modeling, planner integration, safety constraints, and training strategies shape learning-based motion planning in dynamic environments. Finally, we discuss open challenges and future directions, including sim-to-real gap, safe and certifiable planning, dense crowd navigation, perception-planning coupling, and embodied AI.

cs.RO

Coverage Path Planning: Classical Foundations, Recent Advances, and Future Directions

Coverage path planning (CPP) is a fundamental problem in robot motion planning, whose aim is to produce robot trajectories that provide complete coverage of target workspaces while minimizing task-specific objectives such as path length, overlap, number of turns, and energy consumption. CPP has widespread applications in cleaning, inspection, mapping, agriculture, manufacturing, surveillance, demining, and environmental monitoring. Although classical CPP has been extensively studied, recent advances have extended CPP beyond single-robot settings to multi-robot systems, complex 3D environments, constrained platforms, learning-based coverage planning, and visual coverage tasks. This paper presents a comprehensive survey of 125 representative works published primarily between 2015 and 2026, while presenting the evolution of recent developments in light of the classical CPP methods published before 2015. The CPP methods are organized into six main categories: single-robot CPP, multi-robot CPP, 3D CPP, constrained CPP, learning-based CPP, and visual CPP. For each category, the review summarizes the main planning formulations, representative algorithms, strengths, and limitations. In addition, the review analyzes how environmental knowledge, workspace geometry, robot constraints, sensing objectives, and coordination requirements shape the CPP problem. The survey further discusses open challenges in scalable online planning, multi-robot coordination, 3D and visual coverage, unified platform-constrained and resource-aware coverage, and learning-enhanced coverage. Thus, the survey provides a structured overview of recent CPP developments and future research directions.

cs.RO

Motion Planning in Dynamic Environments: A Survey from Classical to Modern Methods

Motion planning in dynamic environments requires robots to continuously adapt their paths in response to environmental changes for safe and uninterrupted navigation. While many surveys have reviewed planning in static settings, systematic reviews focused on dynamic environments remain limited. This paper presents a comprehensive survey of 138 works, primarily published between 2015 and 2025, spanning both classical and learning-based approaches. The motion planning methods are grouped into five categories based on the concepts of sampling, graph search, model predictive control, learning, and additional classical local planning approaches, including velocity obstacles, potential fields and dynamic windows. The learning techniques include supervised learning and reinforcement learning. We also discuss the role of dynamic perception in motion planning, covering techniques for detecting and modeling moving obstacles using cameras, LiDAR, and event-based sensors. The survey analyzes the principles, strengths, and limitations of each method, with particular attention to challenges unique to dynamic environments, such as prediction uncertainty, human-robot interaction, and the freezing robot problem. The survey provides researchers with a structured understanding of motion planning methods in dynamic environments.

cs.RO

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive single-frame editing and semantic motion description. By harnessing the robust generative priors of image-to-video diffusion models, SimInsert propagates edits temporally, strictly preserving background invariance while enabling plausible, text-driven interactions between the inserted object and the dynamic environment. Our approach hinges on non-invasive guidance mechanisms that enforce structural consistency, facilitate seamless boundary fusion, and counteract the fidelity drift that typically accumulates during the denoising trajectory. Extensive quantitative experiments validate our efficacy: SimInsert surpasses state-of-the-art methods with an 18.8\% gain in PSNR, 20.1\% in SSIM, and a 44.1\% decrease in LPIPS, offering a streamlined solution for high-fidelity video editing.

cs.CV

Symbiotic Brain-Machine Drawing via Visual Brain-Computer Interfaces

Brain-computer interfaces (BCIs) are evolving from research prototypes into clinical, assistive, and performance enhancement technologies. Despite the rapid rise and promise of implantable technologies, there is a need for better and more capable wearable and non-invasive approaches whilst also minimising hardware requirements. We present a non-invasive BCI for mind-drawing that iteratively infers a subject's internal visual intent by adaptively presenting visual stimuli (probes) on a screen encoded at different flicker-frequencies and analyses the steady-state visual evoked potentials (SSVEPs). A Gabor-inspired or machine-learned policies dynamically update the spatial placement of the visual probes on the screen to explore the image space and reconstruct simple imagined shapes within approximately two minutes or less using just single-channel EEG data. Additionally, by leveraging stable diffusion models, reconstructed mental images can be transformed into realistic and detailed visual representations. Whilst we expect that similar results might be achievable with e.g. eye-tracking techniques, our work shows that symbiotic human-AI interaction can significantly increase BCI bit-rates by more than a factor 5x, providing a platform for future development of AI-augmented BCI.

q-bio.NC

A Luminance-Aware Multi-Scale Network for Polarization Image Fusion with a Multi-Scene Dataset

Polarization image fusion combines S0 and DOLP images to reveal surface roughness and material properties through complementary texture features, which has important applications in camouflage recognition, tissue pathology analysis, surface defect detection and other fields. To intergrate coL-Splementary information from different polarized images in complex luminance environment, we propose a luminance-aware multi-scale network (MLSN). In the encoder stage, we propose a multi-scale spatial weight matrix through a brightness-branch , which dynamically weighted inject the luminance into the feature maps, solving the problem of inherent contrast difference in polarized images. The global-local feature fusion mechanism is designed at the bottleneck layer to perform windowed self-attention computation, to balance the global context and local details through residual linking in the feature dimension restructuring stage. In the decoder stage, to further improve the adaptability to complex lighting, we propose a Brightness-Enhancement module, establishing the mapping relationship between luminance distribution and texture features, realizing the nonlinear luminance correction of the fusion result. We also present MSP, an 1000 pairs of polarized images that covers 17 types of indoor and outdoor complex lighting scenes. MSP provides four-direction polarization raw maps, solving the scarcity of high-quality datasets in polarization image fusion. Extensive experiment on MSP, PIF and GAND datasets verify that the proposed MLSN outperms the state-of-the-art methods in subjective and objective evaluations, and the MS-SSIM and SD metircs are higher than the average values of other methods by 8.57%, 60.64%, 10.26%, 63.53%, 22.21%, and 54.31%, respectively. The source code and dataset is avalable at https://github.com/1hzf/MLS-UNet.

cs.CV

Digitization Can Stall Swarm Transport: Commensurability Locking in Quantized-Sensing Chains

We present a minimal model for autonomous robotic swarms in one- and higher-dimensional spaces, where identical, field-driven agents interact pairwise to self-organize spacing and independently follow local gradients sensed through quantized digital sensors. We show that the collective response of a multi-agent train amplifies sensitivity to weak gradients beyond what is achievable by a single agent. We discover a fractional transport phenomenon in which, under a uniform gradient, collective motion freezes abruptly whenever the ratio of intra-agent sensor separation to inter-agent spacing satisfies a number-theoretic commensurability condition. This commensurability locking persists even as the number of agents tends to infinity. We find that this condition is exactly solvable on the rationals -- a dense subset of real numbers -- providing analytic, testable predictions for when transport stalls. Our findings establish a surprising bridge between number theory and emergent transport in swarm robotics, informing design principles with implications for collective migration, analog computation, and even the exploration of number-theoretic structure via physical experimentation.

cond-mat.soft

GPU-Accelerated Monte Carlo Simulation and Experimental Study of Radiative Transfer in Multiple Scattering Media

Addressing the problem of photon multiple scattering interference caused by turbid media in optical measurements, biomedical imaging, environmental monitoring and other fields, existing Monte Carlo light scattering simulations widely adopt the Henyey-Greenstein (H-G) phase function approximation model. However, traditional computational resource limitations and high numerical complexity have constrained the application of precise scattering models. Moreover, the single-parameter anisotropy factor assumption neglects higher-order scattering effects and backscattering intensity, failing to accurately characterize the multi-order scattering properties of complex media. To address these issues, we propose a GPU-accelerated Monte Carlo-Rigorous Mie scattering transport model for complex scattering environments. The model employs rigorous Mie scattering theory to replace the H-G approximation, achieving efficient parallel processing of phase function sampling and complex scattering processes through pre-computed cumulative distribution function optimization and deep integration with CUDA parallel architecture. To validate the model accuracy, a standard scattering experimental platform based on 5{\mu}m polystyrene microspheres was established, with multiple optical depth experimental conditions designed, and spatial registration techniques employed to achieve precise alignment between simulation and experimental images. The research results quantitatively demonstrate the systematic accuracy advantages of rigorous Mie scattering phase functions over H-G approximation in simulating lateral scattering light intensity distributions, providing reliable theoretical foundations and technical support for high-precision optical applications in complex scattering environments.

physics.optics

Intrinsic pressure as a convenient mechanical framework for dry active matter

The identification of local pressure in active matter systems remains a subject of considerable debate. Through theoretical calculations and extensive simulations of various active systems, we demonstrate that intrinsic pressure (defined in the same way as in passive systems) is an ideal candidate for local pressure of dry active matter, while the self-propelling forces on the active particles are considered as effective external forces originating from the environment. Such a framework is universal and especially convenient for analyzing mechanics of dry active systems, and it recovers the conventional scenario of mechanical equilibrium well-known in passive systems. Thus, our work is of fundamental importance to further explore mechanics and thermodynamics of complex active systems.

cond-mat.soft

Active Hyperuniform Networks of Chiral Magnetic Micro-Robotic Spinners

Disorder hyperuniform (DHU) systems possess a hidden long-range order manifested as the complete suppression of normalized large-scale density fluctuations like crystals, which endows them with many unique properties. Here, we demonstrate a new organization mechanism for achieving stable DHU structures in active-particle systems via investigating the self-assembly of robotic spinners with three-fold symmetric magnetic binding sites up to a heretofore experimentally unattained system size, i.e., with $\sim 1000$ robots. The spinners can self-organize into a wide spectrum of actively rotating three-coordinated network structures, among which a set of stable DHU networks robustly emerge. These DHU networks are topological transformations of a honeycomb network by continuously introducing the Stone-Wales defects, which are resulted from the competition between tunable magnetic binding and local twist due to active rotation of the robots. Our results reveal novel mechanisms for emergent DHU states in active systems and achieving novel DHU materials with desirable properties.

cond-mat.soft

Informational Memory Shapes Collective Behavior in Intelligent Swarms

We present an experimental and theoretical study of 2-D swarms in which collective behavior emerges from both direct local mechanical coupling between agents and from the exchange and processing of information between agents. Each agent, an air-table drone endowed with internal memory and a binary decision variable, updates its state by integrating a time series of memories of local past collisions. This internal computation transforms the drone swarm into a dynamical information network in which history-dependent feedback drives spontaneous complete spin polarization, pitchfork bifurcated spin collectives, and chaotic switching between collective states. By tuning the depth of memory and the decision algorithm, we uncover a memory-induced phase transition that breaks spin symmetry at the population level. A minimal theoretical model maps these dynamics onto an effective potential landscape sculpted by informational feedback, revealing how temporally correlated computation can replace instantaneous forces as the driver of collective organization, informed by experiments. These results position physically interacting drone swarms as a model system for exploring the physics of informational drone ensembles whose emergent behavior arises from the interplay between physical interaction and information processing.

physics.soc-ph

L1 Adaptive Resonance Ratio Control for Series Elastic Actuator with Guaranteed Transient Performance

To eliminate the static error, overshoot, and vibration of the series elastic actuator (SEA) position control, the resonance ratio control (RRC) algorithm is improved based on L1 adaptive control(L1AC)method. Based on the analysis of the factors affecting the control performance of SEA, the algorithm schema is proposed, the stability is proved, and the main control parameters are analyzed. The algorithm schema is further improved with gravity compensation, and the predicted error and reference error is reduced to guarantee transient performance. Finally, the effectiveness of the algorithm is validated by simulation and platform experiments. The simulation and experiment results show that the algorithm has good adaptability, can improve transient control performance, and can handle effectively the static error, overshoot, and vibration. In addition, when a link-side collision occurs, the algorithm automatically reduces the link speed and limits the motor current, thus protecting the humans and SEA itself, due to the low pass filter characterization of L1AC to disturbance.

cs.RO

LCAUnet: A skin lesion segmentation network with enhanced edge and body fusion

Accurate segmentation of skin lesions in dermatoscopic images is crucial for the early diagnosis of skin cancer and improving the survival rate of patients. However, it is still a challenging task due to the irregularity of lesion areas, the fuzziness of boundaries, and other complex interference factors. In this paper, a novel LCAUnet is proposed to improve the ability of complementary representation with fusion of edge and body features, which are often paid little attentions in traditional methods. First, two separate branches are set for edge and body segmentation with CNNs and Transformer based architecture respectively. Then, LCAF module is utilized to fuse feature maps of edge and body of the same level by local cross-attention operation in encoder stage. Furthermore, PGMF module is embedded for feature integration with prior guided multi-scale adaption. Comprehensive experiments on public available dataset ISIC 2017, ISIC 2018, and PH2 demonstrate that LCAUnet outperforms most state-of-the-art methods. The ablation studies also verify the effectiveness of the proposed fusion techniques.

eess.IV

Inertial Spinner Swarm Experiments: Spin Pumping, Entropy Oscillations and Spin Frustration

We present here an inertial active spinning swarm consisting of mixtures of opposite handedness torque driven spinners floating on an air bed with low damping. Depending on the relative spin sign, spinners can act as their own anti-particles and annihilate their spins. Rotational energy can become highly focused, with minority fraction spinners pumped to very high levels of spin angular momentum. Spinner handedness also matters at high spinner densities but not low densities: oscillations in the mixing spatial entropy of spinners over time emerge if there is a net spin imbalance from collective rotations. Geometrically confined spinners can lock themselves into frustrated spin states.

cond-mat.soft

Computational imaging with the human brain

Brain-computer interfaces (BCIs) are enabling a range of new possibilities and routes for augmenting human capability. Here, we propose BCIs as a route towards forms of computation, i.e. computational imaging, that blend the brain with external silicon processing. We demonstrate ghost imaging of a hidden scene using the human visual system that is combined with an adaptive computational imaging scheme. This is achieved through a projection pattern `carving' technique that relies on real-time feedback from the brain to modify patterns at the light projector, thus enabling more efficient and higher resolution imaging. This brain-computer connectivity demonstrates a form of augmented human computation that could in the future extend the sensing range of human vision and provide new approaches to the study of the neurophysics of human perception. As an example, we illustrate a simple experiment whereby image reconstruction quality is affected by simultaneous conscious processing and readout of the perceived light intensities.

cs.CV

Nth-order nonlinear intensity fluctuation amplifier

Stronger light intensity fluctuations are pursued by related applications such as optical resolution, image enhancement, and beam positioning. In this paper, an Nth-order light intensity fluctuation amplifier is proposed, which was demonstrated by a four-wave mixing process with different statistical distribution coupling lights. Firstly, its amplification mechanism is revealed both theoretically and experimentally. The ratio $R$ of statistical distributions and the degree of second-order coherence ${g^{(2)}}(0)$ of beams are used to characterize the affected modulations and the increased light intensity fluctuations through the four-wave mixing process. The results show that the amplification of light intensity fluctuations is caused by not only the fluctuating light fields of incident coupling beams, but also the fluctuating nonlinear coefficient of interaction. At last, we highlight the potentiality of applying such amplifier to other N-order nonlinear optical effects.

physics.optics

Classification and Generation of Light Sources Using Gamma Fitting

In general, the typical approach to discriminate antibunching, bunching or superbunching categories make use of calculating the second-order coherence function ${g^{(2)}}(τ)$ of light. Although the classical light sources correspond to the specific degree of second-order coherence ${g^{(2)}}(0)$, it does not alone constitute a distinguishable metric to characterize and determine light sources. Here we propose a new mechanism to directly classify and generate antibunching, bunching or superbunching categories of light, as well as the classical light sources such as thermal and coherent light, by Gamma fitting according to only one characteristic parameter $α$ or $β$. Experimental verification of beams from four-wave mixing process is in agreement with the presented mechanism, and the in fluence of temperature $T$ and laser detuning $Δ$ on the measured results are investigated. The proposal demonstrates the potential of classifying and identifying light with different nature, and the most importantly, provides a convenient and simple method to generate light sources meeting various application requirements according to the presented rules. Most notably, the bunching and superbunching are distinguishable in super-Poissonian statistics using our mechanism.

physics.optics