Searcharxiv⌕ Search

arXiv subjects

Mingzheng Li

Publications and source records attributed to Mingzheng Li.

9 recordsLinked to original sources

Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism

Owing to the unprecedented capability in semantic understanding and logical reasoning, large language models (LLMs) have shown fantastic potential in developing next-generation sequential recommender systems (RSs). However, existing LLM-based sequential RSs mostly separate index generation from sequential recommendation, leading to insufficient integration between semantic information and collaborative information. On the other hand, the neglect of user-related information hinders LLM-based sequential RSs from exploiting high-order user-item interaction patterns. In this paper, we propose the End-to-End Dual Dynamic (ED$^2$) recommender, the first LLM-based sequential RS which adopts dual dynamic index mechanism, targeting resolving the above limitations simultaneously. The dual dynamic index mechanism can not only assembly index generation and sequential recommendation into a unified LLM-backbone pipeline, but also make it practical for LLM-based sequential recommender to take advantage of user-related information. Specifically, to facilitate the LLM comprehension ability to dual dynamic index, we propose a multigrained token regulator which constructs alignment supervision based on LLMs semantic knowledge across multiple representation granularities. Moreover, the associated user collection data and a series of novel instruction tuning tasks are specially customized to capture the high-order user-item interaction patterns. Extensive experiments on three public datasets demonstrate the superiority of ED$^2$, achieving an average improvement of 19.62% in Hit-Rate and 21.11% in NDCG.

cs.IR↗

DTBIA: An Immersive Visual Analytics System for Brain-Inspired Research

The Digital Twin Brain (DTB) is an advanced artificial intelligence framework that integrates spiking neurons to simulate complex cognitive functions and collaborative behaviors. For domain experts, visualizing the DTB's simulation outcomes is essential to understanding complex cognitive activities. However, this task poses significant challenges due to DTB data's inherent characteristics, including its high-dimensionality, temporal dynamics, and spatial complexity. To address these challenges, we developed DTBIA, an Immersive Visual Analytics System for Brain-Inspired Research. In collaboration with domain experts, we identified key requirements for effectively visualizing spatiotemporal and topological patterns at multiple levels of detail. DTBIA incorporates a hierarchical workflow - ranging from brain regions to voxels and slice sections - along with immersive navigation and a 3D edge bundling algorithm to enhance clarity and provide deeper insights into both functional (BOLD) and structural (DTI) brain data. The utility and effectiveness of DTBIA are validated through two case studies involving with brain research experts. The results underscore the system's role in enhancing the comprehension of complex neural behaviors and interactions.

cs.HC↗

When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning

Multimodal Attributed Graphs (MAGs) are ubiquitous in real-world applications, encompassing extensive knowledge through multimodal attributes attached to nodes (e.g., texts and images) and topological structure representing node interactions. Despite its potential to advance diverse research fields like social networks and e-commerce, MAG representation learning (MAGRL) remains underexplored due to the lack of standardized datasets and evaluation frameworks. In this paper, we first propose MAGB, a comprehensive MAG benchmark dataset, featuring curated graphs from various domains with both textual and visual attributes. Based on MAGB dataset, we further systematically evaluate two mainstream MAGRL paradigms: $\textit{GNN-as-Predictor}$, which integrates multimodal attributes via Graph Neural Networks (GNNs), and $\textit{VLM-as-Predictor}$, which harnesses Vision Language Models (VLMs) for zero-shot reasoning. Extensive experiments on MAGB reveal following critical insights: $\textit{(i)}$ Modality significances fluctuate drastically with specific domain characteristics. $\textit{(ii)}$ Multimodal embeddings can elevate the performance ceiling of GNNs. However, intrinsic biases among modalities may impede effective training, particularly in low-data scenarios. $\textit{(iii)}$ VLMs are highly effective at generating multimodal embeddings that alleviate the imbalance between textual and visual attributes. These discoveries, which illuminate the synergy between multimodal attributes and graph topologies, contribute to reliable benchmarks, paving the way for future MAG research. The MAGB dataset and evaluation pipeline are publicly available at https://github.com/sktsherlock/MAGB.

cs.LG↗

Improving the detection sensitivity to primordial stochastic gravitational waves with reduced astrophysical foregrounds. II. Subthreshold binary neutron stars

Stochastic gravitational waves (GWs) consist of a primordial component from early Universe processes and an astrophysical component from compact binary mergers. To detect the primordial stochastic GW background (SGWB), the astrophysical foregrounds must be reduced to high precision, which is achievable for third-generation (3G) ground based GW detectors. Previous studies have shown that the foreground from individually detectable merger events can be reduced with fractional residual energy density below $10^{-3}$, and the residual foreground from subthreshold binary neutron stars (BNSs) will be the bottleneck if not well cleaned. In this work, we propose that the foreground energy density of subthreshold BNSs $Ω_{\rm sub}$ can be estimated via a population based approach from the individually detectable BNSs utilizing the isotropic orbital orientations of all BNSs, i.e., uniform distribution in $\cosι$, where $ι$ is the BNS inclination angle with respect to the line of sight. Using this approach, we find $Ω_{\rm sub}$ can be measured with percent-level uncertainty, assuming $O(10^5)$ individually detected BNSs in our simulations. This method represents a promising approach to tackling the foreground cleaning problem.

gr-qc↗

Detector induced anisotropies on the angular distribution of gravitational wave sources and opportunities of constraining horizon scale anisotropies

The cosmological principle has been verified using electromagnetic (EM) observations. However its verification with high accuracy is challenging due to various foregrounds and selection effects, and possible violation of the cosmological principle has been reported in the literature. In contrast, gravitational wave (GW) observations are free of these foregrounds and related selection biases. This may enable future GW experiments to test the cosmological principle robustly with full sky distribution of millions of standard bright/dark sirens. However, the sensitivities of GW detectors are highly anisotropic, resulting in significant instrument induced anisotropies in the observed GW catalog. We investigate these instrumental effects for 3rd generation detector networks in term of multipoles $a_{\ell m}$ of the observed GW source distribution, using Monte Carlo simulations. (1) We find that the instrument induced anisotropy primarily exists at the $m=0$ modes on large scales ($\ell \lesssim 10$), with amplitude $\langle |a_{\ell 0}|^2 \rangle \sim 10^{-3}$ for two detectors (ET-CE) and $\sim 10^{-4}$ for three detectors (ET-2CE). This anisotropy is correlated with the sky distribution of signal-to-noise ratio (SNR) and localization accuracy. Such anisotropy sets a lower limit on the detectable cosmological $a_{\ell 0}$. (2) However, we find that the instrument induced anisotropy is efficiently canceled by rotation of the Earth in $m\neq 0$ components of $a_{\ell m}$. Therefore $a_{\ell m}$ ($m\neq 0$) are clean windows to detect cosmological anisotropies. (3) We investigate the capability of 3rd generation GW experiments to measure the cosmic dipole. Through Monte Carlo simulations, we find that cosmic dipole with an amplitude of $\sim 10^{-2}$ reported in the literature can be detected/ruled out by ET-CE and ET-2CE robustly, through the measurement of $a_{11}$.

astro-ph.CO↗

Hybrid Multimodal Fusion for Humor Detection

In this paper, we present our solution to the MuSe-Humor sub-challenge of the Multimodal Emotional Challenge (MuSe) 2022. The goal of the MuSe-Humor sub-challenge is to detect humor and calculate AUC from audiovisual recordings of German football Bundesliga press conferences. It is annotated for humor displayed by the coaches. For this sub-challenge, we first build a discriminant model using the transformer module and BiLSTM module, and then propose a hybrid fusion strategy to use the prediction results of each modality to improve the performance of the model. Our experiments demonstrate the effectiveness of our proposed model and hybrid fusion strategy on multimodal fusion, and the AUC of our proposed model on the test set is 0.8972.

cs.LG↗

Localized Graph Collaborative Filtering

User-item interactions in recommendations can be naturally de-noted as a user-item bipartite graph. Given the success of graph neural networks (GNNs) in graph representation learning, GNN-based C methods have been proposed to advance recommender systems. These methods often make recommendations based on the learned user and item embeddings. However, we found that they do not perform well wit sparse user-item graphs which are quite common in real-world recommendations. Therefore, in this work, we introduce a novel perspective to build GNN-based CF methods for recommendations which leads to the proposed framework Localized Graph Collaborative Filtering (LGCF). One key advantage of LGCF is that it does not need to learn embeddings for each user and item, which is challenging in sparse scenarios. Alternatively, LGCF aims at encoding useful CF information into a localized graph and making recommendations based on such graph. Extensive experiments on various datasets validate the effectiveness of LGCF especially in sparse scenarios. Furthermore, empirical results demonstrate that LGCF provides complementary information to the embedding-based CF model which can be utilized to boost recommendation performance.

cs.IR↗

An Adaptive Graph Pre-training Framework for Localized Collaborative Filtering

Graph neural networks (GNNs) have been widely applied in the recommendation tasks and have obtained very appealing performance. However, most GNN-based recommendation methods suffer from the problem of data sparsity in practice. Meanwhile, pre-training techniques have achieved great success in mitigating data sparsity in various domains such as natural language processing (NLP) and computer vision (CV). Thus, graph pre-training has the great potential to alleviate data sparsity in GNN-based recommendations. However, pre-training GNNs for recommendations face unique challenges. For example, user-item interaction graphs in different recommendation tasks have distinct sets of users and items, and they often present different properties. Therefore, the successful mechanisms commonly used in NLP and CV to transfer knowledge from pre-training tasks to downstream tasks such as sharing learned embeddings or feature extractors are not directly applicable to existing GNN-based recommendations models. To tackle these challenges, we delicately design an adaptive graph pre-training framework for localized collaborative filtering (ADAPT). It does not require transferring user/item embeddings, and is able to capture both the common knowledge across different graphs and the uniqueness for each graph. Extensive experimental results have demonstrated the effectiveness and superiority of ADAPT.

cs.IR↗

Joint Observations of Space-based Gravitational-wave Detectors: Source Localization and Implication for Parity-violating gravity

Space-based gravitational-wave (GW) detectors, including LISA, Taiji and TianQin, are able to detect mHz GW signals produced by mergers of supermassive black hole binaries, which opens a new window for GW astronomy. In this article, we numerically estimate the potential capabilities of the future networks of multiple space-based detectors using Bayesian analysis. We modify the public package Bilby and employ the sampler PyMultiNest to analyze the simulated data of the space-based detector networks, and investigate their abilities for source localization and testing the parity symmetry of gravity. In comparison with the case of an individual detector, we find detector networks can significantly improve the source localization. While for constraining the parity symmetry of gravity, we find that detector networks and an individual detector follow the similar constraints on the parity-violating energy scale $M_{\rm PV}$. Similar analysis can be applied to other potential observations of various space-based GW detectors.

gr-qc↗