SearcharxivSearch

arXiv subjects

Guoqiang Xu

Publications and source records attributed to Guoqiang Xu.

At least 19 recordsLinked to original sources

FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder

We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicability in realistic scenarios where an object is typically observed and captured from multiple viewpoints. Although multi-view observations naturally provide complementary geometric and appearance cues, existing multimodal large models rarely explore how to effectively fuse such multi-view visual information for better cross-modal retrieval. To address this limitation, we introduce a multi-view image--3D retrieval framework named FusionBERT, which innovatively utilizes a cross-attention-based multi-view visual aggregator to adaptively integrate features from multi-view images of an object. The proposed multi-view visual encoder fuses inter-view complementary relationships and selectively emphasizes informative visual cues across multiple views to get a more robustly fused visual feature for better 3D model matching. Furthermore, FusionBERT proposes a normal-aware 3D model encoder that can further enhance the 3D geometric feature of an object model by jointly encoding point normals and 3D positions, enabling a more robust representation learning for textureless or color-degraded 3D models. Extensive image--3D retrieval experiments on both synthetic 3D models and real-world industrial mechanical objects demonstrate that FusionBERT achieves significantly higher retrieval accuracy than SOTA multimodal large models under both single-view and multi-view settings, establishing a strong baseline for multi-view multimodal retrieval.

cs.CV

Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management

The rapid increase in LLM model sizes and the growing demand for long-context inference have made memory a critical bottleneck in GPU-accelerated serving systems. Although high-bandwidth memory (HBM) on GPUs offers fast access, its limited capacity necessitates reliance on host memory (CPU DRAM) to support larger working sets such as the KVCache. However, the maximum DRAM capacity is constrained by the limited number of memory channels per CPU socket. To overcome this limitation, current systems often adopt RDMA-based disaggregated memory pools, which introduce significant challenges including high access latency, complex communication protocols, and synchronization overhead. Fortunately, the emerging CXL technology introduces new opportunities in KVCache design. In this paper, we propose Beluga, a novel memory architecture that enables GPUs and CPUs to access a shared, large-scale memory pool through CXL switches. By supporting native load/store access semantics over the CXL fabric, our design delivers near-local memory latency, while reducing programming complexity and minimizing synchronization overhead. We conduct a systematic characterization of a commercial CXL switch-based memory pool and propose a set of design guidelines. Based on Beluga, we design and implement Beluga-KVCache, a system tailored for managing the large-scale KVCache in LLM inference. Beluga-KVCache achieves an 89.6% reduction in Time-To-First-Token (TTFT) and 7.35x throughput improvement in the vLLM inference engine compared to RDMA-based solutions. To the best of our knowledge, Beluga is the first system that enables GPUs to directly access large-scale memory pools through CXL switches, marking a significant step toward low-latency, shared access to vast memory resources by GPUs.

cs.DC

One-way heat transfer in deep-subwavelength thermophotonics

Nonreciprocal thermophotonics, by breaking Lorentz reciprocity, exceeds current theoretical efficiency limits, unlocking opportunities to energy devices and thermal management. However, energy transfer in current systems is highly defect-sensitive. This sensitivity is further amplified at deep subwavelength scales by inevitable multi-source interactions, interface wrinkles, and manufacturing tolerances, making precise control of thermal photons increasingly challenging. Here, we demonstrate a topological one-way heat transport in a deep-subwavelength thermophotonic lattice. This one-way heat flow, driven by global resonances, is strongly localized at the geometric boundaries and exhibits exceptional robustness against imperfections and disorder, achieving nearly five orders of radiative enhancement. Our findings offer a blueprint for developing robust thermal systems capable of withstanding strong perturbations.

physics.optics

STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models

Though Large Language Models (LLMs) have demonstrated the powerful capabilities of few-shot learning through prompting methods, supervised training is still necessary for complex reasoning tasks. Because of their extensive parameters and memory consumption, both Parameter-Efficient Fine-Tuning (PEFT) methods and Memory-Efficient Fine-Tuning methods have been proposed for LLMs. Nevertheless, the issue of large annotated data consumption, the aim of Data-Efficient Fine-Tuning, remains unexplored. One obvious way is to combine the PEFT method with active learning. However, the experimental results show that such a combination is not trivial and yields inferior results. Through probe experiments, such observation might be explained by two main reasons: uncertainty gap and poor model calibration. Therefore, in this paper, we propose a novel approach to effectively integrate uncertainty-based active learning and LoRA. Specifically, for the uncertainty gap, we introduce a dynamic uncertainty measurement that combines the uncertainty of the base model and the uncertainty of the full model during the iteration of active learning. For poor model calibration, we incorporate the regularization method during LoRA training to keep the model from being over-confident, and the Monte-Carlo dropout mechanism is employed to enhance the uncertainty estimation. Experimental results show that the proposed approach outperforms existing baseline models on three complex reasoning tasks.

cs.CL

DINER: Debiasing Aspect-based Sentiment Analysis with Multi-variable Causal Inference

Though notable progress has been made, neural-based aspect-based sentiment analysis (ABSA) models are prone to learn spurious correlations from annotation biases, resulting in poor robustness on adversarial data transformations. Among the debiasing solutions, causal inference-based methods have attracted much research attention, which can be mainly categorized into causal intervention methods and counterfactual reasoning methods. However, most of the present debiasing methods focus on single-variable causal inference, which is not suitable for ABSA with two input variables (the target aspect and the review). In this paper, we propose a novel framework based on multi-variable causal inference for debiasing ABSA. In this framework, different types of biases are tackled based on different causal intervention methods. For the review branch, the bias is modeled as indirect confounding from context, where backdoor adjustment intervention is employed for debiasing. For the aspect branch, the bias is described as a direct correlation with labels, where counterfactual reasoning is adopted for debiasing. Extensive experiments demonstrate the effectiveness of the proposed method compared to various baselines on the two widely used real-world aspect robustness test set datasets.

cs.CL

Higher-Order Topological In-Bulk Corner State in Pure Diffusion Systems

Compared with conventional topological insulator that carries topological state at its boundaries, the higher-order topological insulator exhibits lower-dimensional gapless boundary states at its corners and hinges. Leveraging the form similarity between Schrodinger equation and diffusion equation, researches on higher-order topological insulators have been extended from condensed matter physics to thermal diffusion. Unfortunately, all the corner states of thermal higher-order topological insulator reside within the band gap. Another kind of corner state, which is embedded in the bulk states, has not been realized in pure diffusion systems so far. Here, we construct higher-dimensional Su-Schrieffer-Heeger models based on sphere-rod structure to elucidate these corner states, which we term ``in-bulk corner states". Due to the anti-Hermitian properties of diffusive Hamiltonian, we investigate the thermal behaviour of these corner states through theoretical calculation, simulation, and experiment. Furthermore, we study the different thermal behaviours of in-bulk corner state and in-gap corner state. Our results would open a different gate for diffusive topological states and provide a distinct application for efficient heat dissipation.

physics.app-ph

Aiming at the Target: Filter Collaborative Information for Cross-Domain Recommendation

Cross-domain recommender (CDR) systems aim to enhance the performance of the target domain by utilizing data from other related domains. However, irrelevant information from the source domain may instead degrade target domain performance, which is known as the negative transfer problem. There have been some attempts to address this problem, mostly by designing adaptive representations for overlapped users. Whereas, representation adaptions solely rely on the expressive capacity of the CDR model, lacking explicit constraint to filter the irrelevant source-domain collaborative information for the target domain. In this paper, we propose a novel Collaborative information regularized User Transformation (CUT) framework to tackle the negative transfer problem by directly filtering users' collaborative information. In CUT, user similarity in the target domain is adopted as a constraint for user transformation learning to filter the user collaborative information from the source domain. CUT first learns user similarity relationships from the target domain. Then, source-target information transfer is guided by the user similarity, where we design a user transformation layer to learn target-domain user representations and a contrastive loss to supervise the user collaborative information transferred. The results show significant performance improvement of CUT compared with SOTA single and cross-domain methods. Further analysis of the target-domain results illustrates that CUT can effectively alleviate the negative transfer problem.

cs.IR

Fine-grainedly Synthesize Streaming Data Based On Large Language Models With Graph Structure Understanding For Data Sparsity

Due to the sparsity of user data, sentiment analysis on user reviews in e-commerce platforms often suffers from poor performance, especially when faced with extremely sparse user data or long-tail labels. Recently, the emergence of LLMs has introduced new solutions to such problems by leveraging graph structures to generate supplementary user profiles. However, previous approaches have not fully utilized the graph understanding capabilities of LLMs and have struggled to adapt to complex streaming data environments. In this work, we propose a fine-grained streaming data synthesis framework that categorizes sparse users into three categories: Mid-tail, Long-tail, and Extreme. Specifically, we design LLMs to comprehensively understand three key graph elements in streaming data, including Local-global Graph Understanding, Second-Order Relationship Extraction, and Product Attribute Understanding, which enables the generation of high-quality synthetic data to effectively address sparsity across different categories. Experimental results on three real datasets demonstrate significant performance improvements, with synthesized data contributing to MSE reductions of 45.85%, 3.16%, and 62.21%, respectively.

cs.CL

TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision

Numerous large language model (LLM) agents have been built for different tasks like web navigation and online shopping due to LLM's wide knowledge and text-understanding ability. Among these works, many of them utilize in-context examples to achieve generalization without the need for fine-tuning, while few of them have considered the problem of how to select and effectively utilize these examples. Recently, methods based on trajectory-level retrieval with task meta-data and using trajectories as in-context examples have been proposed to improve the agent's overall performance in some sequential decision making tasks. However, these methods can be problematic due to plausible examples retrieved without task-specific state transition dynamics and long input with plenty of irrelevant context. In this paper, we propose a novel framework (TRAD) to address these issues. TRAD first conducts Thought Retrieval, achieving step-level demonstration selection via thought matching, leading to more helpful demonstrations and less irrelevant input noise. Then, TRAD introduces Aligned Decision, complementing retrieved demonstration steps with their previous or subsequent steps, which enables tolerance for imperfect thought and provides a choice for balance between more context and less noise. Extensive experiments on ALFWorld and Mind2Web benchmarks show that TRAD not only outperforms state-of-the-art models but also effectively helps in reducing noise and promoting generalization. Furthermore, TRAD has been deployed in real-world scenarios of a global business insurance company and improves the success rate of robotic process automation.

cs.AI

Deep learning-assisted active metamaterials with heat-enhanced thermal transport

Heat management is crucial for state-of-the-art applications such as passive radiative cooling, thermally adjustable wearables, and camouflage systems. Their adaptive versions, to cater to varied requirements, lean on the potential of adaptive metamaterials. Existing efforts, however, feature with highly anisotropic parameters, narrow working-temperature ranges, and the need for manual intervention, which remain long-term and tricky obstacles for the most advanced self-adaptive metamaterials. To surmount these barriers, we introduce heat-enhanced thermal diffusion metamaterials powered by deep learning. Such active metamaterials can automatically sense ambient temperatures and swiftly, as well as continuously, adjust their thermal functions with a high degree of tunability. They maintain robust thermal performance even when external thermal fields change direction, and both simulations and experiments demonstrate exceptional results. Furthermore, we design two metadevices with on-demand adaptability, performing distinctive features with isotropic materials, wide working temperatures, and spontaneous response. This work offers a framework for the design of intelligent thermal diffusion metamaterials and can be expanded to other diffusion fields, adapting to increasingly complex and dynamic environments.

physics.app-ph

Quadrupole Topological Insulator in Non-Hermitian Thermal Diffusion

Quantized bulk quadrupole moment has unveiled a nontrivial boundary state, exhibiting lower-dimensional topological edge states and simultaneously hosting the in-gap corner modes of zero dimension. All state-of-the-art strategies for topological thermal metamaterials have so far failed to observe such higher-order hierarchical features, since the absence of quantized bulk quadrupole moments in thermal diffusion fundamentally forbids the possible expansions of band topology, unlike its photonic counterpart. Here, we report a recipe of creating quantized bulk quadrupole moments in diffusion, and observe the quadrupole topological phases in non-Hermitian thermal systems. The experiments demonstrate that both the real- and imaginary-valued bands showcase the hierarchical hallmarks of bulk, gapped edge, and in-gap corner states, in stark contrast to the higher-order states only observed on real-valued bands in classic wave fields. Our findings open up unique possibilities for diffusive manipulations and establish an unexplored playground for multipolar topological physics.

cond-mat.mes-hall

Modeling User Repeat Consumption Behavior for Online Novel Recommendation

Given a user's historical interaction sequence, online novel recommendation suggests the next novel the user may be interested in. Online novel recommendation is important but underexplored. In this paper, we concentrate on recommending online novels to new users of an online novel reading platform, whose first visits to the platform occurred in the last seven days. We have two observations about online novel recommendation for new users. First, repeat novel consumption of new users is a common phenomenon. Second, interactions between users and novels are informative. To accurately predict whether a user will reconsume a novel, it is crucial to characterize each interaction at a fine-grained level. Based on these two observations, we propose a neural network for online novel recommendation, called NovelNet. NovelNet can recommend the next novel from both the user's consumed novels and new novels simultaneously. Specifically, an interaction encoder is used to obtain accurate interaction representation considering fine-grained attributes of interaction, and a pointer network with a pointwise loss is incorporated into NovelNet to recommend previously-consumed novels. Moreover, an online novel recommendation dataset is built from a well-known online novel reading platform and is released for public use as a benchmark. Experimental results on the dataset demonstrate the effectiveness of NovelNet.

cs.IR

Blackhole-Inspired Thermal Trapping with Graded Heat-Conduction Metadevices

Black holes are one of the most intriguing predictions of general relativity. So far, metadevices have enabled analogous black holes to trap light or sound in laboratory spacetime. However, trapping heat in a conductive ambient is still challenging because diffusive behaviors are directionless. Inspired by black holes, we construct graded heat-conduction metadevices to achieve thermal trapping, resorting to the imitated advection produced by graded thermal conductivities rather than the trivial solution of using insulation materials to confine thermal diffusion. We experimentally demonstrate thermal trapping for guiding hot spots to diffuse towards the center. Graded heat-conduction metadevices have advantages in energy-efficient thermal regulation because the imitated advection has a similar temperature field effect to the realistic advection that is usually driven by external energy sources. These results also provide insights into correlating transformation thermotics with other disciplines such as cosmology for emerging heat control schemes.

physics.app-ph

A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition

Analyzing human affect is vital for human-computer interaction systems. Most methods are developed in restricted scenarios which are not practical for in-the-wild settings. The Affective Behavior Analysis in-the-wild (ABAW) 2021 Contest provides a benchmark for this in-the-wild problem. In this paper, we introduce a multi-modal and multi-task learning method by using both visual and audio information. We use both AU and expression annotations to train the model and apply a sequence model to further extract associations between video frames. We achieve an AU score of 0.712 and an expression score of 0.477 on the validation set. These results demonstrate the effectiveness of our approach in improving model performance.

cs.CV

Diffusive Topological Transport in Spatiotemporal Thermal Lattices

Topological insulating phases are usually found in periodic lattices stemming from collective resonant effects, and it may thus be expected that similar features may be prohibited in thermal diffusion, given its purely dissipative and largely incoherent nature. We report the diffusion-based topological states supported by spatiotemporally-modulated advections stacked over a fluidic surface, thereby imitating a periodic propagating potential in effective thermal lattices. We observe edge and bulk states within purely nontrivial and trivial lattices, respectively. At interfaces between these two types of lattices, the diffusive system exhibits interface states, manifesting inhomogeneous thermal properties on the fluidic surface. Our findings establish a framework for topological diffusion and thermal edge/bulk states, and it may empower a distinct mechanism for flexible manipulation of robust heat and mass transfer.

physics.app-ph

Curriculum-Meta Learning for Order-Robust Continual Relation Extraction

Continual relation extraction is an important task that focuses on extracting new facts incrementally from unstructured text. Given the sequential arrival order of the relations, this task is prone to two serious challenges, namely catastrophic forgetting and order-sensitivity. We propose a novel curriculum-meta learning method to tackle the above two challenges in continual relation extraction. We combine meta learning and curriculum learning to quickly adapt model parameters to a new task and to reduce interference of previously seen tasks on the current task. We design a novel relation representation learning method through the distribution of domain and range types of relations. Such representations are utilized to quantify the difficulty of tasks for the construction of curricula. Moreover, we also present novel difficulty-based metrics to quantitatively measure the extent of order-sensitivity of a given model, suggesting new ways to evaluate model robustness. Our comprehensive experiments on three benchmark datasets show that our proposed method outperforms the state-of-the-art techniques. The code is available at the anonymous GitHub repository: https://github.com/wutong8023/AAAI_CML.

cs.CL

Guided waves in pre-stressed hyperelastic plates and tubes: Application to the ultrasound elastography of thin-walled soft materials

In vivo measurement of the mechanical properties of thin-walled soft tissues (e.g., mitral valve, artery and bladder) and in situ mechanical characterization of thin-walled artificial soft biomaterials in service are of great challenge and difficult to address via commonly used testing methods. Here we investigate the properties of guided waves generated by focused acoustic radiation force in immersed pre-stressed plates and tubes, and show that they can address this challenge. To this end, we carry out both (i) a theoretical analysis based on incremental wave motion in finite deformation theory and (ii) finite element simulations. Our analysis leads to a novel method based on the ultrasound elastography to image the elastic properties of pre-stressed thin-walled soft tissues and artificial soft materials in a non-destructive and non-invasive manner. To validate the theoretical and numerical solutions and demonstrate the usefulness of the corresponding method in practical measurements, we perform (iii) experiments on polyvinyl alcohol cryogel phantoms immersed in water, using the Verasonics V1 System equipped with a L10-5 transducer. Finally, potential clinical applications of the method have been discussed.

cond-mat.soft

Learning to Augment Expressions for Few-shot Fine-grained Facial Expression Recognition

Affective computing and cognitive theory are widely used in modern human-computer interaction scenarios. Human faces, as the most prominent and easily accessible features, have attracted great attention from researchers. Since humans have rich emotions and developed musculature, there exist a lot of fine-grained expressions in real-world applications. However, it is extremely time-consuming to collect and annotate a large number of facial images, of which may even require psychologists to correctly categorize them. To the best of our knowledge, the existing expression datasets are only limited to several basic facial expressions, which are not sufficient to support our ambitions in developing successful human-computer interaction systems. To this end, a novel Fine-grained Facial Expression Database - F2ED is contributed in this paper, and it includes more than 200k images with 54 facial expressions from 119 persons. Considering the phenomenon of uneven data distribution and lack of samples is common in real-world scenarios, we further evaluate several tasks of few-shot expression learning by virtue of our F2ED, which are to recognize the facial expressions given only few training instances. These tasks mimic human performance to learn robust and general representation from few examples. To address such few-shot tasks, we propose a unified task-driven framework - Compositional Generative Adversarial Network (Comp-GAN) learning to synthesize facial images and thus augmenting the instances of few-shot expression classes. Extensive experiments are conducted on F2ED and existing facial expression datasets, i.e., JAFFE and FER2013, to validate the efficacy of our F2ED in pre-training facial expression recognition network and the effectiveness of our proposed approach Comp-GAN to improve the performance of few-shot recognition tasks.

cs.CV