SearcharxivSearch

arXiv subjects

Lichen Wang

Publications and source records attributed to Lichen Wang.

At least 19 recordsLinked to original sources

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-language sentences. As a result, they provide only a limited assessment of whether a model truly understands the semantic content of SL videos. To address this limitation, \textbf{we first propose a new task, Sign Language Question Answering (SLQA)}, which evaluates SL understanding by requiring models to answer arbitrary natural language questions about SL videos. Unlike previous SLU tasks, SLQA provides a more flexible and comprehensive evaluation framework that assesses multiple reasoning capabilities beyond recognition and translation. To facilitate this task, \textbf{we further construct two SignQA benchmarks} based on PHOENIX14T and CSL-Daily by automatically generating question-answer pairs from existing gloss and sentence annotations using carefully designed templates. The resulting datasets cover five complementary question categories, including position reasoning, structural reasoning, visual search, gloss recognition, and translation understanding. \textbf{Finally, we propose a simple yet effective baseline model} equipped with a Question-Conditioned Modulated Temporal Downsampling module and an in-domain knowledge transfer strategy, enabling effective knowledge transfer from existing SLU tasks while enhancing question-aware temporal feature modeling. Extensive experiments demonstrate that our baseline consistently outperforms representative vision-language models across all question categories, establishing a strong benchmark for future research on SLQA. Datasets are available at:{https://huggingface.co/datasets/hulala/SignQA-2026}.

cs.AI

Coevolutionary dynamics of cooperation, risk, and cost in collective risk games

Addressing both natural and societal challenges requires collective cooperation. Studies on collective-risk social dilemmas have shown that individual decisions are influenced by the perceived risk of collective failure. However, existing feedback evolving game models often focus on a single feedback mechanism, such as the coupling between cooperation and risk or between cooperation and cost. In many real-world scenarios, however, the level of cooperation, the cost of cooperating, and the collective risk are dynamically interlinked. Here, we present an evolutionary game model that considers the interplay of these three variables. Our analysis shows that the worst-case scenario, characterized by full defection, maximum risk, and the highest cost of cooperation, remains a stable evolutionary attractor. Nevertheless, cooperation can emerge and persist because the system also supports stable equilibria with non-zero cooperation. The system exhibits multistability, meaning that different initial conditions lead to either sustained cooperation or a tragedy of the commons. These findings highlight that initial levels of cooperation, cost, and risk collectively determine whether a population can avert a tragic outcome.

nlin.AO

Anisotropic magnetoelastic coupling in the honeycomb magnet Na$_3$Co$_2$SbO$_6$

We present magnetization and dilatometry measurements on the honeycomb cobaltate Na$_3$Co$_2$SbO$_6$ and map out its detailed field-temperature phase diagram down to sub-Kelvin temperatures. Our data for in-plane magnetic fields show a strongly anisotropic $c^{*}$-axis lattice response, which is dominated by the variation of Co--O--Co bond angles according to \textit{ab initio} calculations. At $T = 0.4$~K, the magnetization $M(B)$ exhibits step-like features that are also highly anisotropic. In the case of $B \parallel b$, a small hysteresis observed around the second field-induced magnetic transition ($B_{c2}$) indicates its first-order character, whereas divergence of the magnetic Gr\"uneisen parameter at $B_{c2}$ is suppressed upon cooling and signals the absence of quantum critical behavior upon entering the field-polarized state. None of our thermodynamic measurements provide evidence for a field-induced quantum spin liquid state near or above $B_{c2}$.

cond-mat.str-el

Ultrasensitive strain modulation of terahertz magnons at a magnetic phase transition

Antiferromagnets typically host spin-wave (magnon) excitations in the terahertz (THz) regime, offering a promising platform for high-speed magnonic information technologies. Harnessing these excitations requires sensitive control of their spectral properties. Here we use resonant x-ray diffraction and Raman scattering to demonstrate uniaxial-strain control of the antiferromagnetic (AFM) ground state and THz magnon excitations in the layered Mott insulator Ca$_2$RuO$_4$. Although the states separated by the strain-induced phase transition differ only by the sign of the weak and partially frustrated interlayer interaction, their magnon energies differ by more than 10% (~ 0.3 THz). Our theoretical analysis explains this surprising observation by tracing the origin of both the sign reversal of the interlayer coupling and the magnon energy to the spin-orbital composition of the Ru valence electrons. The extreme strain sensitivity of the THz magnon energy near a magnetic phase transition opens up pathways towards a new generation of transition-edge magnonic devices.

cond-mat.mtrl-sci

LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation

Video Large Language Models (VLLMs) unlock world-knowledge-aware video understanding through pretraining on internet-scale data and have already shown promise on tasks such as movie analysis and video question answering. However, deploying VLLMs for downstream tasks such as video recommendation remains challenging, since real systems require multi-video inputs, lightweight backbones, low-latency sequential inference, and rapid response. In practice, (1) decode-only generation yields high latency for sequential inference, (2) typical interfaces do not support multi-video inputs, and (3) constraining outputs to language discards fine-grained visual details that matter for downstream vision tasks. We argue that these limitations stem from the absence of a representation that preserves pixel-level detail while leveraging world knowledge. We present LinkedOut, a representation that extracts VLLM world knowledge directly from video to enable fast inference, supports multi-video histories, and removes the language bottleneck. LinkedOut extracts semantically grounded, knowledge-aware tokens from raw frames using VLLMs, guided by promptable queries and optional auxiliary modalities. We introduce a cross-layer knowledge fusion MoE that selects the appropriate level of abstraction from the rich VLLM features, enabling personalized, interpretable, and low-latency recommendation. To our knowledge, LinkedOut is the first VLLM-based video recommendation method that operates on raw frames without handcrafted labels, achieving state-of-the-art results on standard benchmarks. Interpretability studies and ablations confirm the benefits of layer diversity and layer-wise fusion, pointing to a practical path that fully leverages VLLM world-knowledge priors and visual reasoning for downstream vision tasks such as recommendation.

cs.CV

Strategic competition in informal risk sharing mechanism versus collective index insurance

The frequent occurrence of natural disasters has posed significant challenges to society, necessitating the urgent development of effective risk management strategies. From the early informal community-based risk sharing mechanisms to modern formal index insurance products, risk management tools have continuously evolved. Although index insurance provides an effective risk transfer mechanism in theory, it still faces the problems of basis risk and pricing in practice. At the same time, in the presence of informal community risk sharing mechanisms, the competitiveness of index insurance deserves further investigation. Here we propose a three-strategy evolutionary game model, which simultaneously examines the competitive relationship between formal index insurance purchasing (I), informal risk sharing strategies (S), and complete non-insurance (A). Furthermore, we introduce a method for calculating insurance company profits to aid in the optimal pricing of index insurance products. We find that basis risk and risk loss ratio have significant impacts on insurance adoption rate. Under scenarios with low basis risk and high loss ratios, index insurance is more popular; meanwhile, when the loss ratio is moderate, an informal risk sharing strategy is the preferred option. Conversely, when the loss ratio is low, individuals tend to forego any insurance. Furthermore, accurately assessing the degree of risk aversion and determining the appropriate ratio of risk sharing are crucial for predicting the future market sales of index insurance.

q-fin.RM

The paradigm of tax-reward and tax-punishment strategies in the advancement of public resource management dynamics

In contemporary society, the effective utilization of public resources remains a subject of significant concern. A common issue arises from defectors seeking to obtain an excessive share of these resources for personal gain, potentially leading to resource depletion. To mitigate this tragedy and ensure sustainable development of resources, implementing mechanisms to either reward those who adhere to distribution rules or penalize those who do not, appears advantageous. We introduce two models: a tax-reward model and a tax-punishment model, to address this issue. Our analysis reveals that in the tax-reward model, the evolutionary trajectory of the system is influenced not only by the tax revenue collected but also by the natural growth rate of the resources. Conversely, the tax-punishment model exhibits distinct characteristics when compared to the tax-reward model, notably the potential for bistability. In such scenarios, the selection of initial conditions is critical, as it can determine the system's path. Furthermore, our study identifies instances where the system lacks stable points, exemplified by a limit cycle phenomenon, underscoring the complexity and dynamism inherent in managing public resources using these models.

math.DS

Spin-orbit excitons in a correlated metal: Raman scattering study of Sr2RhO4

Using Raman spectroscopy to study the correlated 4$d$-electron metal Sr$_2$RhO$_4$, we observe pronounced excitations at 220 meV and 240 meV with $A_\mathrm{1g}$ and $B_\mathrm{1g}$ symmetries, respectively. We identify them as transitions between the spin-orbit multiplets of the Rh ions, in close analogy to the spin-orbit excitons in the Mott insulators Sr$_2$IrO$_4$ and $α$-RuCl$_3$. This observation provides direct evidence for the unquenched spin-orbit coupling in Sr$_2$RhO$_4$. A quantitative analysis of the data reveals that the tetragonal crystal field $Δ$ in Sr$_2$RhO$_4$ has a sign opposite to that in insulating Sr$_2$IrO$_4$, which enhances the planar $xy$ orbital character of the effective $J=1/2$ wave function. This supports a metallic ground state, and suggests that $c$-axis compression of Sr$_2$RhO$_4$ may transform it into a quasi-two-dimensional antiferromagnetic insulator.

cond-mat.str-el

iBARLE: imBalance-Aware Room Layout Estimation

Room layout estimation predicts layouts from a single panorama. It requires datasets with large-scale and diverse room shapes to train the models. However, there are significant imbalances in real-world datasets including the dimensions of layout complexity, camera locations, and variation in scene appearance. These issues considerably influence the model training performance. In this work, we propose the imBalance-Aware Room Layout Estimation (iBARLE) framework to address these issues. iBARLE consists of (1) Appearance Variation Generation (AVG) module, which promotes visual appearance domain generalization, (2) Complex Structure Mix-up (CSMix) module, which enhances generalizability w.r.t. room structure, and (3) a gradient-based layout objective function, which allows more effective accounting for occlusions in complex layouts. All modules are jointly trained and help each other to achieve the best performance. Experiments and ablation studies based on ZInD~\cite{cruz2021zillow} dataset illustrate that iBARLE has state-of-the-art performance compared with other layout estimation baselines.

cs.CV

Resonant Inelastic X-ray Scattering from Electronic Excitations in $α$-RuCl$_3$ Nanolayers

We present Ru $L_3$-edge resonant inelastic x-ray scattering (RIXS) measurements of spin-orbit and d-d excitations in exfoliated nanolayers of the Kitaev spin-liquid candidate RuCl$_3$. Whereas the spin-orbit excitations are independent of thickness, we observe a pronounced red-shift and broadening of the d-d excitations in layers with thickness below $\sim$7 nm. Aided by model calculations, we attribute these effects to distortions of the RuCl$_6$ octahedra near the surface. Our study paves the way towards RIXS investigations of electronic excitations in various other 2D materials and heterostructures.

cond-mat.str-el

Adaptive Trajectory Prediction via Transferable GNN

Pedestrian trajectory prediction is an essential component in a wide range of AI applications such as autonomous driving and robotics. Existing methods usually assume the training and testing motions follow the same pattern while ignoring the potential distribution differences (e.g., shopping mall and street). This issue results in inevitable performance decrease. To address this issue, we propose a novel Transferable Graph Neural Network (T-GNN) framework, which jointly conducts trajectory prediction as well as domain alignment in a unified framework. Specifically, a domain-invariant GNN is proposed to explore the structural motion knowledge where the domain-specific knowledge is reduced. Moreover, an attention-based adaptive knowledge learning module is further proposed to explore fine-grained individual-level feature representations for knowledge transfer. By this way, disparities across different trajectory domains will be better alleviated. More challenging while practical trajectory prediction experiments are designed, and the experimental results verify the superior performance of our proposed model. To the best of our knowledge, our work is the pioneer which fills the gap in benchmarks and techniques for practical pedestrian trajectory prediction across different domains.

cs.CV

Bone tumor suppression in rabbits by hyperthermia below the clinical safety limit using aligned magnetic bone cement

Demonstrating highly efficient alternating current (AC) magnetic field heating of nanoparticles in physiological environments under clinically safe field parameters has remained a great challenge, hindering clinical applications of magnetic hyperthermia. In this work, we report exceptionally high loss power of magnetic bone cement under clinical safety limit of AC field parameters, incorporating DC field-aligned soft magnetic Zn0.3Fe2.7O4 nanoparticles with low concentration. Under an AC field of 4 kA/m at 430 kHz, the aligned bone cement with 0.2 wt% nanoparticles achieved a temperature increase of 30 C in 180 s. This amounts to a specific loss power value of 327 W/gmetal and an intrinsic loss power of 47 nHm^2/kg, which is enhanced by 50-fold compared to randomly oriented samples. The high-performance magnetic bone cement allows for the demonstration of effective hyperthermia suppression of tumor growth in the bone marrow cavity of New Zealand White Rabbits subjecting to rapid cooling due to blood circulation, and significant enhancement of survival rate.

physics.med-ph

Semi-supervised Domain Adaptive Structure Learning

Semi-supervised domain adaptation (SSDA) is quite a challenging problem requiring methods to overcome both 1) overfitting towards poorly annotated data and 2) distribution shift across domains. Unfortunately, a simple combination of domain adaptation (DA) and semi-supervised learning (SSL) methods often fail to address such two objects because of training data bias towards labeled samples. In this paper, we introduce an adaptive structure learning method to regularize the cooperation of SSL and DA. Inspired by the multi-views learning, our proposed framework is composed of a shared feature encoder network and two classifier networks, trained for contradictory purposes. Among them, one of the classifiers is applied to group target features to improve intra-class density, enlarging the gap of categorical clusters for robust representation learning. Meanwhile, the other classifier, serviced as a regularizer, attempts to scatter the source features to enhance the smoothness of the decision boundary. The iterations of target clustering and source expansion make the target features being well-enclosed inside the dilated boundary of the corresponding source points. For the joint address of cross-domain features alignment and partially labeled data learning, we apply the maximum mean discrepancy (MMD) distance minimization and self-training (ST) to project the contradictory structures into a shared view to make the reliable final decision. The experimental results over the standard SSDA benchmarks, including DomainNet and Office-home, demonstrate both the accuracy and robustness of our method over the state-of-the-art approaches.

cs.CV

Sign Language Recognition via Skeleton-Aware Multi-Model Ensemble

Sign language is commonly used by deaf or mute people to communicate but requires extensive effort to master. It is usually performed with the fast yet delicate movement of hand gestures, body posture, and even facial expressions. Current Sign Language Recognition (SLR) methods usually extract features via deep neural networks and suffer overfitting due to limited and noisy data. Recently, skeleton-based action recognition has attracted increasing attention due to its subject-invariant and background-invariant nature, whereas skeleton-based SLR is still under exploration due to the lack of hand annotations. Some researchers have tried to use off-line hand pose trackers to obtain hand keypoints and aid in recognizing sign language via recurrent neural networks. Nevertheless, none of them outperforms RGB-based approaches yet. To this end, we propose a novel Skeleton Aware Multi-modal Framework with a Global Ensemble Model (GEM) for isolated SLR (SAM-SLR-v2) to learn and fuse multi-modal feature representations towards a higher recognition rate. Specifically, we propose a Sign Language Graph Convolution Network (SL-GCN) to model the embedded dynamics of skeleton keypoints and a Separable Spatial-Temporal Convolution Network (SSTCN) to exploit skeleton features. The skeleton-based predictions are fused with other RGB and depth based modalities by the proposed late-fusion GEM to provide global information and make a faithful SLR prediction. Experiments on three isolated SLR datasets demonstrate that our proposed SAM-SLR-v2 framework is exceedingly effective and achieves state-of-the-art performance with significant margins. Our code will be available at https://github.com/jackyjsy/SAM-SLR-v2

cs.CV

In-plane Isotropy of the Low Energy Phonon Anomalies in YBa$_{2}$Cu$_{3}$O$_{6+x}$

We study the temperature dependence of the low energy phonons in the $(H, 0, L)$ reciprocal plane of the highly ordered ortho-II YBa$_2$Cu$_3$O$_{6.55}$ cuprate high temperature superconductor by means of high-resolution inelastic x-ray scattering. Anomalies associated with the emergence of long-range charge density wave (CDW) fluctuations are observed, and are qualitatively similar to those previously observed in the $(0, K, L)$ plane. This confirms the unconventional nature of this bi-dimensional CDW, which is not soft-phonon driven. With the support of first principles calculations, the symmetry of the anomalous phonon is identified and is found to match that of the charge modulation. This suggests in turn that these anomalies originate from a direct coupling between the phonons and the collective CDW excitations.

cond-mat.supr-con

Skeleton Aware Multi-modal Sign Language Recognition

Sign language is commonly used by deaf or speech impaired people to communicate but requires significant effort to master. Sign Language Recognition (SLR) aims to bridge the gap between sign language users and others by recognizing signs from given videos. It is an essential yet challenging task since sign language is performed with the fast and complex movement of hand gestures, body posture, and even facial expressions. Recently, skeleton-based action recognition attracts increasing attention due to the independence between the subject and background variation. However, skeleton-based SLR is still under exploration due to the lack of annotations on hand keypoints. Some efforts have been made to use hand detectors with pose estimators to extract hand key points and learn to recognize sign language via Neural Networks, but none of them outperforms RGB-based methods. To this end, we propose a novel Skeleton Aware Multi-modal SLR framework (SAM-SLR) to take advantage of multi-modal information towards a higher recognition rate. Specifically, we propose a Sign Language Graph Convolution Network (SL-GCN) to model the embedded dynamics and a novel Separable Spatial-Temporal Convolution Network (SSTCN) to exploit skeleton features. RGB and depth modalities are also incorporated and assembled into our framework to provide global information that is complementary to the skeleton-based methods SL-GCN and SSTCN. As a result, SAM-SLR achieves the highest performance in both RGB (98.42\%) and RGB-D (98.53\%) tracks in 2021 Looking at People Large Scale Signer Independent Isolated SLR Challenge. Our code is available at https://github.com/jackyjsy/CVPR21Chal-SLR

cs.CV

Contradictory Structure Learning for Semi-supervised Domain Adaptation

Current adversarial adaptation methods attempt to align the cross-domain features, whereas two challenges remain unsolved: 1) the conditional distribution mismatch and 2) the bias of the decision boundary towards the source domain. To solve these challenges, we propose a novel framework for semi-supervised domain adaptation by unifying the learning of opposite structures (UODA). UODA consists of a generator and two classifiers (i.e., the source-scattering classifier and the target-clustering classifier), which are trained for contradictory purposes. The target-clustering classifier attempts to cluster the target features to improve intra-class density and enlarge inter-class divergence. Meanwhile, the source-scattering classifier is designed to scatter the source features to enhance the decision boundary's smoothness. Through the alternation of source-feature expansion and target-feature clustering procedures, the target features are well-enclosed within the dilated boundary of the corresponding source features. This strategy can make the cross-domain features to be precisely aligned against the source bias simultaneously. Moreover, to overcome the model collapse through training, we progressively update the measurement of feature's distance and their representation via an adversarial training paradigm. Extensive experiments on the benchmarks of DomainNet and Office-home datasets demonstrate the superiority of our approach over the state-of-the-art methods.

cs.CV

I3DOL: Incremental 3D Object Learning without Catastrophic Forgetting

3D object classification has attracted appealing attentions in academic researches and industrial applications. However, most existing methods need to access the training data of past 3D object classes when facing the common real-world scenario: new classes of 3D objects arrive in a sequence. Moreover, the performance of advanced approaches degrades dramatically for past learned classes (i.e., catastrophic forgetting), due to the irregular and redundant geometric structures of 3D point cloud data. To address these challenges, we propose a new Incremental 3D Object Learning (i.e., I3DOL) model, which is the first exploration to learn new classes of 3D object continually. Specifically, an adaptive-geometric centroid module is designed to construct discriminative local geometric structures, which can better characterize the irregular point cloud representation for 3D object. Afterwards, to prevent the catastrophic forgetting brought by redundant geometric information, a geometric-aware attention mechanism is developed to quantify the contributions of local geometric structures, and explore unique 3D geometric characteristics with high contributions for classes incremental learning. Meanwhile, a score fairness compensation strategy is proposed to further alleviate the catastrophic forgetting caused by unbalanced data between past and new classes of 3D object, by compensating biased prediction for new classes in the validation phase. Experiments on 3D representative datasets validate the superiority of our I3DOL framework.

cs.CV