Searcharxiv⌕ Search

arXiv subjects

Yongjie Wang

Publications and source records attributed to Yongjie Wang.

At least 19 recordsLinked to original sources

EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation

MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-side modality information, prompting a surge of diverse methods. Despite these advances, existing methods still face two critical limitations. First, they use raw modality features to construct item-item links for enriching the behavior graph, while giving limited attention to balancing collaborative and modality-aware semantics or mitigating modality noise in the process. Second, they use a uniform alignment weight across all entities and also maintain a fixed alignment strength throughout training, limiting the effectiveness of modality-behavior alignment. To address these challenges, we propose EGRA. First, instead of relying on raw modality features, it alleviates sparsity by incorporating into the behavior graph an item-item graph built from representations generated by a pretrained MMR model. This enables the graph to capture both collaborative patterns and modality aware similarities with enhanced robustness against modality noise. Moreover, it introduces a novel bi-level dynamic alignment weighting mechanism to improve modality-behavior representation alignment, which dynamically assigns alignment strength across entities according to their alignment degree, while gradually increasing the overall alignment intensity throughout training. Extensive experiments on five datasets show that EGRA significantly outperforms recent methods, confirming its effectiveness.

cs.IR↗

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such agents may retrieve public benchmark metadata, question context, or even ground-truth answers via web search. This gives rise to Search-Time Contamination (STC), where external retrieval bypasses intended reasoning and inflates measured performance. We systematically study STC in deep research agent evaluation. We define three contamination types with increasing severity, namely Benchmark Metadata Leakage, Question-Context Leakage, and Explicit Answer Leakage, and develop detection algorithms to identify them and quantify their impact on agent performance. Evaluating modern deep research agents on six public benchmarks, we find that STC is widespread and can inflate performance by up to 4%. Our findings show that existing evaluations may overestimate true reasoning ability. We therefore advocate contamination-aware practices, including isolated sandboxes, transparent search trajectories, and controlled benchmark access.

cs.CR↗

Rapid Atmospheric Vapor Deposition of H:In2O3 Transparent Conducting Oxide Thin Films

Transparent conducting oxides (TCOs) are essential for the optoelectronics industry, but there is a critical gap in cost-effective methods to rapidly deposit low sheet resistance, high transmittance films without damaging delicate materials, including emerging soft semiconductors like metal-halide perovskites. In this work, atmospheric pressure chemical vapor deposition (AP-CVD) is used to synthesise H:In2O3 films with 7.20+/-0.01 Ohm/sq sheet resistance (0.50+/-0.06 mOhm.cm resistivity) and transmittance up to 89% in the near-infrared (NIR), surpassing commercial sputter-deposited indium tin oxide. The growth rate is 40x higher than atomic layer deposition (ALD), and the AP-CVD films are fully processed under atmospheric conditions at only 140 C. Comparison of secondary ion mass spectrometry and time-of-flight elastic recoil detection analysis with changes in carrier concentration indicate that H dopants are introduced from the water oxidant. There is an increase in mobility form 40+/-10 cm2/Vs to 160+/-30 cm2/Vs when changing from O2 to H2O as the oxidant, which is attributed to H dopants passivating oxygen vacancies that act as carrier scattering centers. This work establishes AP-CVD as a promising method for manufacturing high figure-of-merit TCOs in a rapid, scalable and cost-effective manner, using mild growth conditions compatible with thermally-sensitive materials.

cond-mat.mtrl-sci↗

CausalGaze: Unveiling Hallucinations via Counterfactual Graph Intervention in Large Language Models

Despite the groundbreaking advancements made by large language models (LLMs), hallucination remains a critical bottleneck for their deployment in high-stakes domains. Existing classification-based methods mainly rely on static and passive signals from internal states, which often captures the noise and spurious correlations, while overlooking the underlying causal mechanisms. To address this limitation, we shift the paradigm from passive observation to active intervention by introducing CausalGaze, a novel hallucination detection framework based on structural causal models (SCMs). CausalGaze models LLMs' internal states as dynamic causal graphs and employs counterfactual interventions to disentangle causal reasoning paths from incidental noise, thereby enhancing model interpretability. Extensive experiments across four datasets and three widely used LLMs demonstrate the effectiveness of CausalGaze, especially achieving 3.3% improvement in AUROC on the TruthfulQA dataset compared to state-of-the-art baselines.

cs.LG↗

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks. However, it remains unclear whether the superior performance of VLMs stems from genuine vision-grounded reasoning or relies predominantly on the reasoning capabilities of their textual backbones. To systematically measure this, we introduce CrossMath, a novel multimodal reasoning benchmark designed for controlled cross-modal comparisons. Specifically, we construct each problem in text-only, image-only, and image+text formats guaranteeing identical task-relevant information, verified by human annotators. This rigorous alignment effectively isolates modality-specific reasoning differences while eliminating confounding factors such as information mismatch. Extensive evaluation of state-of-the-art VLMs reveals a consistent phenomenon: a substantial performance gap between textual and visual reasoning. Notably, VLMs excel with text-only inputs, whereas incorporating visual data (image+text) frequently degrades performance compared to the text-only baseline. These findings indicate that current VLMs conduct reasoning primarily in the textual space, with limited genuine reliance on visual evidence. To mitigate this limitation, we curate a CrossMath training set for VLM fine-tuning. Empirical evaluations demonstrate that fine-tuning on this training set significantly boosts reasoning performance across all individual and joint modalities, while yielding robust gains on two general visual reasoning tasks. Source code is available at https://github.com/xuyige/CrossMath.

cs.CV↗

From Entity-Centric to Goal-Oriented Graphs: Enhancing LLM Knowledge Retrieval in Minecraft

Large Language Models (LLMs) demonstrate impressive general capabilities but often struggle with step-by-step procedural reasoning, a critical challenge in complex interactive environments. While retrieval-augmented methods like GraphRAG attempt to bridge this gap, their fragmented entity-relation graphs hinder the construction of coherent, multi-step plans. In this paper, we propose a novel framework based on Goal-Oriented Graphs (GoGs), where each node represents a goal and edges encode logical dependencies between them. This structure enables the explicit retrieval of causal reasoning paths by identifying a high-level goal and recursively retrieving its prerequisites, forming a coherent chain to guide the LLM. Through extensive experiments on the Minecraft testbed, a domain that demands robust multi-step planning and provides rich procedural knowledge, we demonstrate that GoG substantially improves procedural reasoning and significantly outperforms GraphRAG and other state-of-the-art baselines.

cs.AI↗

Band-Like Transport and Cation Off-Centring in Ag/Bi-Based Solar Absorbers

Ag(I)-Bi(III)-based semiconductors have gained substantial attention as nontoxic, stable alternatives to lead-halide perovskites for optoelectronics, but are widely limited by carrier localization, which severely restricts diffusion lengths. The most efficient Ag/Bi solar absorber is AgBiS2, but diffusion lengths in nanocrystal films are <50 nm. Carrier localization in this rock-salt (Fm-3m) system is believed to arise from cation disorder, and so we herein investigate the layered cation-ordered analogue. Through beyond-DFT simulations combined with neutron and X-ray powder diffraction, we reveal that off-centring of Ag+ and Bi3+ cations is energetically-favoured in this cation-ordered phase. Despite local distortions in the AgS6 and BiS6 octahedra, band-like transport takes place, which, surprisingly, also occurs in the cation-disordered rock-salt phase when these materials are made as bulk powders. The cubic-phase powders have the same degree of cation disorder as the nanocrystals that have carrier localization, which suggests that extrinsic factors play a determining role. We ascribe the intrinsic band-like transport of both phases of AgBiS2 to its close packing, ensuring high electronic dimensionality. These insights offer pathways for designing solar absorbers avoiding carrier localization limitations, and call for future efforts to enhance the efficiency of AgBiS2 photovoltaics to focus on large-grained thin films, or improved nanocrystal surface passivation.

cond-mat.mtrl-sci↗

A Drinfeld Presentation of the Queer Super-Yangian

We introduce a Drinfeld presentation for the super-Yangian $\mathrm{Y}(\mathfrak{q}_n)$ associated with the queer Lie superalgebra $\mathfrak{q}_n$. The Drinfeld generators of $\mathrm{Y}(\mathfrak{q}_n)$ are obtained through a block Gauss decomposition of the generator matrix in its RTT presentation, and the Drinfeld relations are explicitly computed by utilizing a block version of its RTT relations. As a byproduct, we obtain a new expression for the central series of $\mathrm{Y}(\mathfrak{q}_n)$ in terms of Gauss generators.

math.QA↗

When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs

Large Language Models (LLMs) have enabled a wide range of applications through their powerful capabilities in language understanding and generation. However, as LLMs are trained on static corpora, they face difficulties in addressing rapidly evolving information or domain-specific queries. Retrieval-Augmented Generation (RAG) was developed to overcome this limitation by integrating LLMs with external retrieval mechanisms, allowing them to access up-to-date and contextually relevant knowledge. However, as LLMs themselves continue to advance in scale and capability, the relative advantages of traditional RAG frameworks have become less pronounced and necessary. Here, we present a comprehensive review of RAG, beginning with its overarching objectives and core components. We then analyze the key challenges within RAG, highlighting critical weakness that may limit its effectiveness. Finally, we showcase applications where LLMs alone perform inadequately, but where RAG, when combined with LLMs, can substantially enhance their effectiveness. We hope this work will encourage researchers to reconsider the role of RAG and inspire the development of next-generation RAG systems.

cs.CL↗

On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking

The explosive growth of multimodal data has driven the rapid development of multimodal entity linking (MEL) models. However, existing studies have not systematically investigated the impact of visual adversarial attacks on MEL models. We conduct the first comprehensive evaluation of the robustness of mainstream MEL models under different adversarial attack scenarios, covering two core tasks: Image-to-Text (I2T) and Image+Text-to-Text (IT2T). Experimental results show that current MEL models generally lack sufficient robustness against visual perturbations. Interestingly, contextual semantic information in input can partially mitigate the impact of adversarial perturbations. Based on this insight, we propose an LLM and Retrieval-Augmented Entity Linking (LLM-RetLink), which significantly improves the model's anti-interference ability through a two-stage process: first, extracting initial entity descriptions using large vision models (LVMs), and then dynamically generating candidate descriptive sentences via web-based retrieval. Experiments on five datasets demonstrate that LLM-RetLink improves the accuracy of MEL by 0.4%-35.7%, especially showing significant advantages under adversarial conditions. This research highlights a previously unexplored facet of MEL robustness, constructs and releases the first MEL adversarial example dataset, and sets the stage for future work aimed at strengthening the resilience of multimodal systems in adversarial environments.

cs.IR↗

CM$^3$: Calibrating Multimodal Recommendation

Alignment and uniformity are fundamental principles within the domain of contrastive learning. In recommender systems, prior work has established that optimizing the Bayesian Personalized Ranking (BPR) loss contributes to the objectives of alignment and uniformity. Specifically, alignment aims to draw together the representations of interacting users and items, while uniformity mandates a uniform distribution of user and item embeddings across a unit hypersphere. This study revisits the alignment and uniformity properties within the context of multimodal recommender systems, revealing a proclivity among extant models to prioritize uniformity to the detriment of alignment. Our hypothesis challenges the conventional assumption of equitable item treatment through a uniformity loss, proposing a more nuanced approach wherein items with similar multimodal attributes converge toward proximal representations within the hyperspheric manifold. Specifically, we leverage the inherent similarity between items' multimodal data to calibrate their uniformity distribution, thereby inducing a more pronounced repulsive force between dissimilar entities within the embedding space. A theoretical analysis elucidates the relationship between this calibrated uniformity loss and the conventional uniformity function. Moreover, to enhance the fusion of multimodal features, we introduce a Spherical Bézier method designed to integrate an arbitrary number of modalities while ensuring that the resulting fused features are constrained to the same hyperspherical manifold. Empirical evaluations conducted on five real-world datasets substantiate the superiority of our approach over competing baselines. We also shown that the proposed methods can achieve up to a 5.4% increase in NDCG@20 performance via the integration of MLLM-extracted features. Source code is available at: https://github.com/enoche/CM3.

cs.IR↗

RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval

Large Language Models (LLMs) have shown promise in character imitation, enabling immersive and engaging conversations. However, they often generate content that is irrelevant or inconsistent with a character's background. We attribute these failures to: (1) the inability to accurately recall character-specific knowledge due to entity ambiguity, and (2) a lack of awareness of the character's cognitive boundaries. To address these issues, we propose RoleRAG, a retrieval-based framework that integrates efficient entity disambiguation for knowledge indexing with a boundary-aware retriever for extracting contextually appropriate information from a structured knowledge graph. Experiments on role-playing benchmarks show that RoleRAG's calibrated retrieval helps both general-purpose and role-specific LLMs better align with character knowledge and reduce hallucinated responses.

cs.AI↗

Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?

Probing techniques have shown promise in revealing how LLMs encode human-interpretable concepts, particularly when applied to curated datasets. However, the factors governing a dataset's suitability for effective probe training are not well-understood. This study hypothesizes that probe performance on such datasets reflects characteristics of both the LLM's generated responses and its internal feature space. Through quantitative analysis of probe performance and LLM response uncertainty across a series of tasks, we find a strong correlation: improved probe performance consistently corresponds to a reduction in response uncertainty, and vice versa. Subsequently, we delve deeper into this correlation through the lens of feature importance analysis. Our findings indicate that high LLM response variance is associated with a larger set of important features, which poses a greater challenge for probe models and often results in diminished performance. Moreover, leveraging the insights from response uncertainty analysis, we are able to identify concrete examples where LLM representations align with human knowledge across diverse domains, offering additional evidence of interpretable reasoning in LLMs.

cs.AI↗

An Intelligent and Privacy-Preserving Digital Twin Model for Aging-in-Place

The population of older adults is steadily increasing, with a strong preference for aging-in-place rather than moving to care facilities. Consequently, supporting this growing demographic has become a significant global challenge. However, facilitating successful aging-in-place is challenging, requiring consideration of multiple factors such as data privacy, health status monitoring, and living environments to improve health outcomes. In this paper, we propose an unobtrusive sensor system designed for installation in older adults' homes. Using data from the sensors, our system constructs a digital twin, a virtual representation of events and activities that occurred in the home. The system uses neural network models and decision rules to capture residents' activities and living environments. This digital twin enables continuous health monitoring by providing actionable insights into residents' well-being. Our system is designed to be low-cost and privacy-preserving, with the aim of providing green and safe monitoring for the health of older adults. We have successfully deployed our system in two homes over a time period of two months, and our findings demonstrate the feasibility and effectiveness of digital twin technology in supporting independent living for older adults. This study highlights that our system could revolutionize elder care by enabling personalized interventions, such as lifestyle adjustments, medical treatments, or modifications to the residential environment, to enhance health outcomes.

cs.CY↗

Reconfigurable chiral edge states in synthetic dimensions on an integrated photonic chip

Chiral edge state is a hallmark of topological physics, which has drawn significant attention across quantum mechanics, condensed matter and optical systems. Recently, synthetic dimensions have emerged as ideal platforms for investigating chiral edge states in multiple dimensions, overcoming the limitations of real space. In this work, we demonstrate reconfigurable chiral edge states via synthetic dimensions on an integrated photonic chip. These states are realized by coupling two frequency lattices with opposite pseudospins, which are subjected to programmable artificial gauge potential and long-range coupling within a thin-film lithium niobate microring resonator. Within this system, we are able to implement versatile strategies to observe and steer the chiral edge states, including the realization and frustration of the chiral edge states in a synthetic Hall ladder, the generation of imbalanced chiral edge currents, and the regulation of chiral behaviors as chirality, single-pseudospin enhancement, and complete suppression. This work provides a reconfigurable integrated photonic platform for simulating and steering chiral edge states in synthetic space, paying the way for the realization of high-dimensional and programmable topological photonic systems on chip.

physics.optics↗

Quantum Berezinian for the Twisted Super Yangian

Motivated by an open problem proposed in Molev's book \cite[Section 2.16, Example 16]{Mo07}, we investigate the quantum Berezinian $\mathfrak{B}^{tw}(u)$ associated with the twisted super Yangian, which is a coideal sub-superalgebra of the super Yangian of the general linear Lie superalgebra. We provide an explicit formulation of $\mathfrak{B}^{tw}(u)$, and we also construct the center of the twisted super Yangian. This construction enables us to define the special twisted super Yangian, which is isomorphic to the quotient of the twisted super Yangian by its center. Moreover, we demonstrate the quantum Sylvester theorem for both the generator matrix and the quantum Berezinian.

math.QA↗

Typical representations of Takiff superalgebras

We investigate representations of the $\ell$-th Takiff superalgebras $\widetilde{\mathfrak g}_\ell := \widetilde{\mathfrak g}\otimes \mathbb C[θ]/(θ^{\ell+1})$, for $\ell>0$, associated with a basic classical and a periplectic Lie superalgebras $\widetilde{\mathfrak g}$. We introduce the odd reflections and formulate a general notion of typical representations of the Takiff superalgebras $\widetilde{\mathfrak g}_\ell$. As a consequence, we provide a complete description of the characters of the finite-dimensional modules over type I Takiff superalgebras. For the Lie superalgebras $\widetilde{\mathfrak g}= \mathfrak{gl}(m|n)$ and $\mathfrak{osp}(2|2n)$, we prove that the Kac induction functor of $\widetilde{\mathfrak g}_\ell$ leads to an equivalence from an arbitrary typical Jordan block of the category $\mathcal O$ for $\widetilde{\mathfrak g}_\ell$ to a Jordan block of the category $\mathcal O$ for the even subalgebra of $\widetilde{\mathfrak g}_\ell$. We also obtain a classification of non-singular simple Whittaker modules over the Takiff superalgebras.

math.RT↗

The Schur-Weyl duality and Invariants for classical Lie superalgebras

In this article, we provide a comprehensive characterization of invariants of classical Lie superalgebras from the super-analog of the Schur-Weyl duality in a unified way. We establish $\mathfrak{g}$-invariants of the tensor algebra $T(\mathfrak{g})$, the supersymmetric algebra $S(\mathfrak{g})$, and the universal enveloping algebra $\mathrm{U}(\mathfrak{g})$ of a classical Lie superalgebra $\mathfrak{g}$ corresponding to every element in centralizer algebras and their relationship under supersymmetrization. As a byproduct, we prove that the restriction on $T(\mathfrak{g})^{\mathfrak{g}}$ of the projection from $T(\mathfrak{g})$ to $\mathrm{U}(\mathfrak{g})$ is surjective, which enables us to determine the generators of the center $\mathcal{Z}(\mathfrak{g})$ except for $\mathfrak{g}=\mathfrak{osp}_{2m|2n}$. Additionally, we present an alternative algebraic proof of the triviality of $\mathcal{Z}(\mathfrak{p}_n)$. The key ingredient involves a technique lemma related to the symmetric group and Brauer diagrams.

math.RT↗