SearcharxivSearch

arXiv subjects

Liang Song

Publications and source records attributed to Liang Song.

At least 19 recordsLinked to original sources

Fefferman--Stein type inequalities via area and maximal functions for Schr\"odinger operators with applications

In this paper, we establish a Fefferman--Stein inequality in terms of area function and non-tangential maximal function associated with the Schr\"odinger operator $\mathcal{L} = -\Delta + V$ on stratified Lie groups $\mathcal G$, where $\Delta$ denotes the sub-Laplacian on $\mathcal G$ and $V$ is a nonnegative locally integrable function. As an application, we extend this inequality to the tensor product $\mathcal G_1 \times \mathcal G_2$ of two stratified Lie groups and develop atomic decompositions associated with the Schr\"odinger operator for functions in the Orlicz space $L\log^{+}L(\mathcal G_1 \times \mathcal G_2).$ Using these atomic decompositions, we further prove weak-type endpoint estimates for the area integral operator and the Riesz transforms associated with the Schr\"odinger operator on $L\log^{+}L(\mathcal G_1 \times \mathcal G_2),$ thereby extending the celebrated result of R.\,Fefferman and E.M.\,Stein \cite{FSt1982} to the setting of singular integrals with non-smooth kernels.

math.AP

Littlewood--Paley operators and semigroup maximal operators on CMO spaces associated to Sch\"odinger operators

Let $L=-\Delta+V$ be a Schr\"odinger operator on $\mathbb{R}^n$, where $\Delta$ is the Laplacian and $V$ satisfies the reverse H\"older inequality ${\rm RH}_q$ for some $q>n/2$. In this paper, we study the behavior of the Littlewood--Paley operators $s_L$ and $S_L$, as well as the semigroup maximal operator $T^*_L$, on the space ${\rm CMO}_L(\mathbb{R}^n)$ associated with the Schr\"odinger operator $L$. It is known from previous work that these operators are bounded on ${\rm BMO}_L(\mathbb{R}^n)$. Our main result shows that they are, in fact, mappings from ${\rm CMO}_L(\mathbb{R}^n)$ into itself. To prove this, we develop several equivalent characterizations of ${\rm CMO}_L(\mathbb{R}^n)$ and employ a refined decomposition that partitions the parameter interval at $r_B\rho(x_B)$, instead of the customary $r_B^2$ or $\rho(x_B)^2$. The new strategy allows us to overcome a key technical obstacle that arises when applying existing methods to the ${\rm CMO}_L$ setting.

math.AP

Embodied Multimedia: A Tutorial

Traditional multimedia technology has been built around optimizing content delivery for human observers, from perceptually driven compression standards to human-centric quality metrics. With the rapid rise of embodied intelligence, autonomous agents must perceive, reason, and act within the physical world in real time, exposing fundamental mismatches between conventional multimedia infrastructure and the demands of embodied tasks. In this regard, this tutorial paper formally introduces Embodied Multimedia as a cross-disciplinary research paradigm that treats multimodal data as the perceptual and communicative substrate spanning the full perception-decision-action loop. To be specific, we present a four-layer unified architecture comprising Data, Communication, Cognitive, and Evaluation layers, and provide a structured review of key enabling technologies within each layer. Furthermore, we identify five frontier application directions where Embodied Multimedia is positioned to serve a foundational role: multimedia communication, physical intelligence, embodied anomaly perception, the metaverse and interactive multimedia, and AI-driven art creation. Open technical challenges and future research directions are discussed to guide the community in this emerging field.

cs.MM

Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion

Embodied perception systems face severe challenges of dynamic environment distribution drift when they continuously interact in open physical spaces. However, the existing domain incremental awareness methods often rely on the domain id obtained in advance during the testing phase, which limits their practicability in unknown interaction scenarios. At the same time, the model often overfits to the context-specific perceptual noise, which leads to insufficient generalization ability and catastrophic forgetting. To address these limitations, we propose a domain-id and exemplar-free incremental learning framework for embodied multimedia systems, which aims to achieve robust continuous environment adaptation. This method designs a disentangled representation mechanism to remove non-essential environmental style interference, and guide the model to focus on extracting semantic intrinsic features shared across scenes, thereby eliminating perceptual uncertainty and improving generalization. We further use the weight fusion strategy to dynamically integrate the old and new environment knowledge in the parameter space, so as to ensure that the model adapts to the new distribution without storing historical data and maximally retains the discrimination ability of the old environment. Extensive experiments on multiple standard benchmark datasets show that the proposed method significantly reduces catastrophic forgetting in a completely exemplar-free and domain-id free setting, and its accuracy is better than the existing state-of-the-art methods.

cs.CV

Collaborative Adaptive Curriculum for Progressive Knowledge Distillation

Recent advances in collaborative knowledge distillation have demonstrated cutting-edge performance for resource-constrained distributed multimedia learning scenarios. However, achieving such competitiveness requires addressing a fundamental mismatch: high-dimensional teacher knowledge complexity versus heterogeneous client learning capacities, which currently prohibits deployment in edge-based visual analytics systems. Drawing inspiration from curriculum learning principles, we introduce Federated Adaptive Progressive Distillation (FAPD), a consensus-driven framework that orchestrates adaptive knowledge transfer. FAPD hierarchically decomposes teacher features via PCA-based structuring, extracting principal components ordered by variance contribution to establish a natural visual knowledge hierarchy. Clients progressively receive knowledge of increasing complexity through dimension-adaptive projection matrices. Meanwhile, the server monitors network-wide learning stability by tracking global accuracy fluctuations across a temporal consensus window, advancing curriculum dimensionality only when collective consensus emerges. Consequently, FAPD provably adapts knowledge transfer pace while achieving superior convergence over fixed-complexity approaches. Extensive experiments on three datasets validate FAPD's effectiveness: it attains 3.64% accuracy improvement over FedAvg on CIFAR-10, demonstrates 2x faster convergence, and maintains robust performance under extreme data heterogeneity ({\alpha}=0.1), outperforming baselines by over 4.5%.

cs.LG

SplatBright: Generalizable Low-Light Scene Reconstruction from Sparse Views via Physically-Guided Gaussian Enhancement

Low-light 3D reconstruction from sparse views remains challenging due to exposure imbalance and degraded color fidelity. While existing methods struggle with view inconsistency and require per-scene training, we propose SplatBright, which is, to our knowledge, the first generalizable 3D Gaussian framework for joint low-light enhancement and reconstruction from sparse sRGB inputs. Our key idea is to integrate physically guided illumination modeling with geometry-appearance decoupling for consistent low-light reconstruction. Specifically, we adopt a dual-branch predictor that provides stable geometric initialization of 3D Gaussian parameters. On the appearance side, illumination consistency leverages frequency priors to enable controllable and cross-view coherent lighting, while an appearance refinement module further separates illumination, material, and view-dependent cues to recover fine texture. To tackle the lack of large-scale geometrically consistent paired data, we synthesize dark views via a physics-based camera model for training. Extensive experiments on public and self-collected datasets demonstrate that SplatBright achieves superior novel view synthesis, cross-view consistency, and better generalization to unseen low-light scenes compared with both 2D and 3D methods.

cs.CV

Microclimatic variation in tropical canopies: A glimpse into the processes of community assembly in epiphytic bryophyte communities

Epiphytic communities offer an original framework to disentangle the contributions of environmental filters, biotic interactions and dispersal limitations to community structure at fine spatial scales. We determine here whether variations in light, microclimatic conditions and host tree size affect the variation in species composition and phylogenetic structure of epiphytic bryophyte communities, and hence, assess the contribution of environmental filtering, phylogenetic constraints and competition to community assembly.A canopy crane giving access to 1.1 ha of tropical rainforest in Yunnan (China) was employed to record hourly light and microclimatic conditions from 54 dataloggers and epiphytic bryophyte communities from 408 plots. Generalized Dissimilarity Modelling was implemented to analyse the relationship between taxonomic and phylogenetic turnover among epiphytic communities, host-tree characteristics and microclimatic variation.Within-tree vertical turnover of bryophyte communities was significantly about 30% higher than horizontal turnover among-trees. Thus, the sharp vertical variations in microclimatic conditions from tree base to canopy are more important than differences in age, reflecting the likelihood of colonization, area, and habitat conditions between young and old trees, in shaping the composition of epiphytic bryophyte communities.

q-bio.PE

The Hardy spaces $\mathcal{H}^{p}_{FIO}(\mathbb{R}^{n})$ for Fourier integral operators for $p<1$

We introduce the Hardy spaces $\mathcal{H}^{p}_{FIO}(\mathbb{R}^{n})$ for Fourier integral operators for $0<p<1$, thereby extending earlier constructions for $1\leq p\leq \infty$. We then establish various properties of these spaces, including their behavior under complex interpolation and duality, and their invariance under Fourier integral operators. We also obtain Sobolev embeddings, equivalent characterizations, and a molecular decomposition. These spaces are used in the companion article arXiv:2502.02511 to determine the sharp $\mathcal{H}^{1}(\mathbb{R}^{n})$ and $\mathrm{bmo}(\mathbb{R}^{n})$ regularity of wave equations with rough coefficients.

math.AP

EvoPSF: Online Evolution of Autonomous Driving Models via Planning-State Feedback

Recent years have witnessed remarkable progress in autonomous driving, with systems evolving from modular pipelines to end-to-end architectures. However, most existing methods are trained offline and lack mechanisms to adapt to new environments during deployment. As a result, their generalization ability diminishes when faced with unseen variations in real-world driving scenarios. In this paper, we break away from the conventional "train once, deploy forever" paradigm and propose EvoPSF, a novel online Evolution framework for autonomous driving based on Planning-State Feedback. We argue that planning failures are primarily caused by inaccurate object-level motion predictions, and such failures are often reflected in the form of increased planner uncertainty. To address this, we treat planner uncertainty as a trigger for online evolution, using it as a diagnostic signal to initiate targeted model updates. Rather than performing blind updates, we leverage the planner's agent-agent attention to identify the specific objects that the ego vehicle attends to most, which are primarily responsible for the planning failures. For these critical objects, we compute a targeted self-supervised loss by comparing their predicted waypoints from the prediction module with their actual future positions, selected from the perception module's outputs with high confidence scores. This loss is then backpropagated to adapt the model online. As a result, our method improves the model's robustness to environmental changes, leads to more precise motion predictions, and therefore enables more accurate and stable planning behaviors. Experiments on both cross-region and corrupted variants of the nuScenes dataset demonstrate that EvoPSF consistently improves planning performance under challenging conditions.

cs.RO

MMOC: Self-Supervised EEG Emotion Recognition Framework with Multi-Model Online Collaboration

Electroencephalography (EEG) emotion recognition plays a crucial role in human-computer interaction, particularly in healthcare and neuroscience. While supervised learning has been widely used, its reliance on manual annotations introduces high costs and potential bias. Self-supervised learning (SSL) offers a promising alternative by generating labels through pretext tasks. However, high inter-subject variability in EEG signals leads to significant data drift, limiting self-supervised models' generalization across unseen subjects. Traditional domain adaptation (DA) methods require access to target-domain data during training. Although domain generalization (DG) avoids this constraint, it often falls short in handling complex data drift due to limited coverage of possible target distributions. To tackle these challenges, we propose MMOC, a self-supervised framework with multi-model online collaboration (MMOC), to achieve online adaptation to unseen data. MMOC trains multiple base models using diverse strategies rooted in reconstruction and contrastive learning, enabling each model to develop distinct generalization capabilities. During inference, MMOC dynamically activates the most suitable model for each test sample via a loss-based routing mechanism that evaluates both contrastive and reconstruction losses. This dual consideration allows for a comprehensive measurement of data drift at both structural and semantic levels. Experimental results on the SEED and Dreamer datasets show that MMOC achieves state-of-the-art performance: 85.39% on SEED, and 68.77% and 69.37% on Dreamer arousal and valence dimensions, respectively. MMOC effectively mitigates inter-subject data drift, offering a practical solution for real-world EEG emotion recognition.

eess.SP

Evaluating the Sensitivity of LLMs to Prior Context

As large language models (LLMs) are increasingly deployed in multi-turn dialogue and other sustained interactive scenarios, it is essential to understand how extended context affects their performance. Popular benchmarks, focusing primarily on single-turn question answering (QA) tasks, fail to capture the effects of multi-turn exchanges. To address this gap, we introduce a novel set of benchmarks that systematically vary the volume and nature of prior context. We evaluate multiple conventional LLMs, including GPT, Claude, and Gemini, across these benchmarks to measure their sensitivity to contextual variations. Our findings reveal that LLM performance on multiple-choice questions can degrade dramatically in multi-turn interactions, with performance drops as large as 73% for certain models. Even highly capable models such as GPT-4o exhibit up to a 32% decrease in accuracy. Notably, the relative performance of larger versus smaller models is not always predictable. Moreover, the strategic placement of the task description within the context can substantially mitigate performance drops, improving the accuracy by as much as a factor of 3.5. These findings underscore the need for robust strategies to design, evaluate, and mitigate context-related sensitivity in LLMs.

cs.CL

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM's general reasoning potential. To address this limitation, we introduce the Knowledge Orthogonal Reasoning Gymnasium (KORGym), a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.

cs.CL

FreeDriveRF: Monocular RGB Dynamic NeRF without Poses for Autonomous Driving via Point-Level Dynamic-Static Decoupling

Dynamic scene reconstruction for autonomous driving enables vehicles to perceive and interpret complex scene changes more precisely. Dynamic Neural Radiance Fields (NeRFs) have recently shown promising capability in scene modeling. However, many existing methods rely heavily on accurate poses inputs and multi-sensor data, leading to increased system complexity. To address this, we propose FreeDriveRF, which reconstructs dynamic driving scenes using only sequential RGB images without requiring poses inputs. We innovatively decouple dynamic and static parts at the early sampling level using semantic supervision, mitigating image blurring and artifacts. To overcome the challenges posed by object motion and occlusion in monocular camera, we introduce a warped ray-guided dynamic object rendering consistency loss, utilizing optical flow to better constrain the dynamic modeling process. Additionally, we incorporate estimated dynamic flow to constrain the pose optimization process, improving the stability and accuracy of unbounded scene reconstruction. Extensive experiments conducted on the KITTI and Waymo datasets demonstrate the superior performance of our method in dynamic scene modeling for autonomous driving.

cs.CV

Adaptive Weighted Parameter Fusion with CLIP for Class-Incremental Learning

Class-incremental Learning (CIL) enables the model to incrementally absorb knowledge from new classes and build a generic classifier across all previously encountered classes. When the model optimizes with new classes, the knowledge of previous classes is inevitably erased, leading to catastrophic forgetting. Addressing this challenge requires making a trade-off between retaining old knowledge and accommodating new information. However, this balancing process often requires sacrificing some information, which can lead to a partial loss in the model's ability to discriminate between classes. To tackle this issue, we design the adaptive weighted parameter fusion with Contrastive Language-Image Pre-training (CLIP), which not only takes into account the variability of the data distribution of different tasks, but also retains all the effective information of the parameter matrix to the greatest extent. In addition, we introduce a balance factor that can balance the data distribution alignment and distinguishability of adjacent tasks. Experimental results on several traditional benchmarks validate the superiority of the proposed method.

cs.CV

CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion

With the proliferation of multi-modal data in large-scale visual recognition systems, enabling models to continuously acquire knowledge from evolving data streams while preserving prior information has become increasingly critical. Class-Continual Learning (CCL) addresses this challenge by incrementally incorporating new class knowledge without revisiting historical data, making it essential for real-world big data applications. While traditional CCL methods rely solely on visual features, recent advances in Vision-Language Models (VLMs) such as CLIP demonstrate significant potential for CCL by leveraging pre-trained multi-modal knowledge. However, existing approaches face challenges in mitigating catastrophic forgetting while maintaining the cross-modal generalization capabilities of VLMs. To address these limitations, we propose CalFuse, a framework that synergizes feature Calibration with parameter Fusion to enable effective multi-modal knowledge integration in continual learning scenarios. CalFuse introduces a dynamic feature calibration mechanism that adaptively balances original CLIP visual representations with task-specific features, preserving the model's intrinsic cross-modal generalization while adapting to new classes. Concurrently, a QR decomposition-based parameter fusion strategy progressively integrates newly acquired knowledge with historical task parameters, maintaining equilibrium between learning new class representations and retaining prior knowledge across sequential tasks. Extensive experiments on benchmark datasets validate the effectiveness of our approach in large-scale multi-modal continual learning settings, demonstrating superior performance over state-of-the-art methods in both average accuracy and final task retention.

cs.CV

CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos

Video Anomaly Detection (VAD) remains a fundamental yet formidable task in the video understanding community, with promising applications in areas such as information forensics and public safety protection. Due to the rarity and diversity of anomalies, existing methods only use easily collected regular events to model the inherent normality of normal spatial-temporal patterns in an unsupervised manner. Previous studies have shown that existing unsupervised VAD models are incapable of label-independent data offsets (e.g., scene changes) in real-world scenarios and may fail to respond to light anomalies due to the overgeneralization of deep neural networks. Inspired by causality learning, we argue that there exist causal factors that can adequately generalize the prototypical patterns of regular events and present significant deviations when anomalous instances occur. In this regard, we propose Causal Representation Consistency Learning (CRCL) to implicitly mine potential scene-robust causal variable in unsupervised video normality learning. Specifically, building on the structural causal models, we propose scene-debiasing learning and causality-inspired normality learning to strip away entangled scene bias in deep representations and learn causal video normality, respectively. Extensive experiments on benchmarks validate the superiority of our method over conventional deep representation learning. Moreover, ablation studies and extension validation show that the CRCL can cope with label-independent biases in multi-scene settings and maintain stable performance with only limited training data available.

cs.CV

Weak type $(1,1)$ bounds for Riesz transforms for elliptic operators in non-divergence form

Let $L=-\sum_{i,j=1}^n a_{ij}D_iD_j$ be the elliptic operator in non-divergence form with smooth real coefficients satisfying uniformly elliptic condition. Let $W$ be the global nonnegative adjoint solution. If $W\in A_2$, we prove that the Riesz transforms $\nabla L^{-\frac{1}{2}}$ is of weak type $(1,1)$ with respect to the measure $W(x)dx$. This, together with $L^2_W$ boundedness of Riesz transforms \cite{EHH}, implies that the Riesz transforms are bounded in $L^p_W$ for $1<p<2$. Our results are applicable to the case of real coefficients having sufficiently small BMO norm.

math.CA

Baichuan-M1: Pushing the Medical Capability of Large Language Models

The current generation of large language models (LLMs) is typically designed for broad, general-purpose applications, while domain-specific LLMs, especially in vertical fields like medicine, remain relatively scarce. In particular, the development of highly efficient and practical LLMs for the medical domain is challenging due to the complexity of medical knowledge and the limited availability of high-quality data. To bridge this gap, we introduce Baichuan-M1, a series of large language models specifically optimized for medical applications. Unlike traditional approaches that simply continue pretraining on existing models or apply post-training to a general base model, Baichuan-M1 is trained from scratch with a dedicated focus on enhancing medical capabilities. Our model is trained on 20 trillion tokens and incorporates a range of effective training methods that strike a balance between general capabilities and medical expertise. As a result, Baichuan-M1 not only performs strongly across general domains such as mathematics and coding but also excels in specialized medical fields. We have open-sourced Baichuan-M1-14B, a mini version of our model, which can be accessed through the following links.

cs.CL