SearcharxivSearch

arXiv subjects

Liang Yin

Publications and source records attributed to Liang Yin.

16 recordsLinked to original sources

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.

cs.AI

Rotatable Antenna-Enabled Near-Field Integrated Sensing and Communication

In this paper, we propose leveraging rotatable antennas (RAs) to enhance near-field communication and sensing by exploiting a new orientation-domain spatial degree-of-freedom (DoF) provided by element-wise antenna rotation. Specifically, we investigate an RA-enabled near-field integrated sensing and communication (ISAC) system with sub-connected hybrid beamforming, where each transmit RA can independently adjust its boresight direction under a practical rotation constraint. A spherical-wave channel model incorporating orientation-dependent antenna gains is established to characterize multi-user communication and target sensing in the presence of clutters. Based on this model, a weighted communication-sensing utility maximization problem is formulated by jointly optimizing the receive beamformer, digital beamformer, analog beamformer, and RA boresight directions. To solve the resulting non-convex problem, an alternating optimization algorithm is developed by combining fractional programming, Riemannian optimization, and a spherical-cap Frank--Wolfe-based boresight update. To further understand the impact of RA rotation on near-field sensing, we derive a closed-form root Cramer--Rao bound (RCRB) expression. Simulation results demonstrate the convergence and effectiveness of the proposed algorithm. It is shown that the RA-enabled hybrid design can match or even outperform the fully-digital FPA benchmark in some regimes, indicating that the orientation-domain DoF introduced by element-wise rotation can compensate for limited RF chains. The RCRB and beampattern results further show that RA rotation improves off-broadside sensing accuracy, enhances range-domain focusing, and suppresses same-angle clutters in the near field.

eess.SP

Sensing-Aided Secure Multicast in Two-Level Rotatable Antenna-Enabled ISAC Systems: Modeling and Optimization

In physical layer security, the channel state information (CSI) of passive eavesdroppers is usually difficult to obtain, which has motivated sensing-aided secure communication (SASC). However, in secure multicast scenarios, conventional fixed-position antennas (FPAs) provide limited spatial flexibility for simultaneously serving multiple legitimate users and suppressing leakage toward possible eavesdropper directions. Motivated by this, a novel two-level rotatable antenna (RA)-enabled sensing-aided secure multicast scheme is proposed in this paper. In the proposed architecture, array-level and element-wise rotations are jointly exploited with analog beamforming for user enhancement and leakage suppression. To characterize imperfect eavesdropper sensing, the maximum likelihood estimator and the corresponding Cram\'er-Rao bound (CRB) are derived to quantify the angular estimation accuracy. Based on the derived CRB, a probabilistic angular uncertainty region is constructed. A CRB-aware max-min secrecy-rate problem is then formulated by evaluating the eavesdropper leakage over sampled high-probability directions within this region. The non-convex problem is handled through a tractable lower-bound reformulation based on Jensen's inequality and smooth approximation, followed by an alternating optimization algorithm combining manifold optimization and projected-gradient updates. Simulation results show the effectiveness and robustness of the proposed scheme compared with various benchmarks. Beam patterns further reveal that array-level and element-wise rotations play complementary roles in maintaining strong gains toward legitimate users and forming a low-gain region over the eavesdropper angular uncertainty interval.

cs.IT

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

Online Video Large Language Models (VideoLLMs) play a critical role in supporting responsive, real-time interaction. Existing methods focus on streaming perception, lacking a synchronized logical reasoning stream. However, directly applying test-time scaling methods incurs unacceptable response latency. To address this trade-off, we propose Video Streaming Thinking (VST), a novel paradigm for streaming video understanding. It supports a thinking while watching mechanism, which activates reasoning over incoming video clips during streaming. This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning latency over video playback. Furthermore, we introduce a comprehensive post-training pipeline that integrates VST-SFT, which structurally adapts the offline VideoLLM to causal streaming reasoning, and VST-RL, which provides end-to-end improvement through self-exploration in a multi-turn video interaction environment. Additionally, we devise an automated training-data synthesis pipeline that uses video knowledge graphs to generate high-quality streaming QA pairs, with an entity-relation grounded streaming Chain-of-Thought to enforce multi-evidence reasoning and sustained attention to the video stream. Extensive evaluations show that VST-7B performs strongly on online benchmarks, e.g. 79.5% on StreamingBench and 59.3% on OVO-Bench. Meanwhile, VST remains competitive on offline long-form or reasoning benchmarks. Compared with Video-R1, VST responds 15.7 times faster and achieves +5.4% improvement on VideoHolmes, demonstrating higher efficiency and strong generalization across diverse video understanding tasks. Code, data, and models will be released at https://github.com/1ranGuan/VST.

cs.CV

Rotatable Array-Aided Hybrid Beamforming for Integrated Sensing and Communication

Six-dimensional movable antenna (6DMA) technology has been proposed to enhance the performance of Integrated Sensing and Communication (ISAC) systems. However, within 6DMA-related research, studies on the ISAC system based on rotatable array (RA) remains relatively limited. Given the significant advantages of hybrid beamforming technology in balancing system performance and hardware complexity, this paper focuses on a channel model that accounts for the efficiency of the antenna radiation pattern and studies the sub-connected hybrid beamforming design for multi-user RA-aided ISAC systems. Aiming at the non-convex nature with coupled variables in this problem, this paper transforms the complex fractional objective function using the Fractional Programming (FP) method, and then proposes an algorithm based on the Alternating Optimization (AO) framework, which achieves optimization by alternately solving five subproblems. For the analog beamforming optimization subproblem, we adopt Singular Value Decomposition (SVD) method to transform the objective function, thereby deriving the closed-form update expression for the analog beamforming matrix. For the antenna rotation optimization subproblem, we derive the closed-form derivative expression of the array rotation angle and propose a two-stage Gradient Ascent (GA) based method to optimize the antenna rotation angle. Extensive simulation results demonstrate the effectiveness of the proposed RA-aided hybrid beamforming design method. It not only significantly improves the overall system performance while reducing hardware costs, but also achieves performance comparable to that of the fully-digital beamforming design with fixed-position antennas (FPA) under specific parameter configurations.

cs.ET

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

While large multi-modal models (LMMs) demonstrate promising capabilities in segmentation and comprehension, they still struggle with two limitations: inaccurate segmentation and hallucinated comprehension. These challenges stem primarily from constraints in weak visual comprehension and a lack of fine-grained perception. To alleviate these limitations, we propose LIRA, a framework that capitalizes on the complementary relationship between visual comprehension and segmentation via two key components: (1) Semantic-Enhanced Feature Extractor (SEFE) improves object attribute inference by fusing semantic and pixel-level features, leading to more accurate segmentation; (2) Interleaved Local Visual Coupling (ILVC) autoregressively generates local descriptions after extracting local features based on segmentation masks, offering fine-grained supervision to mitigate hallucinations. Furthermore, we find that the precision of object segmentation is positively correlated with the latent related semantics of the token. To quantify this relationship and the model's potential semantic inferring ability, we introduce the Attributes Evaluation (AttrEval) dataset. Our experiments show that LIRA achieves state-of-the-art performance in both segmentation and comprehension tasks. Code will be available at https://github.com/echo840/LIRA.

cs.CV

MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling

Scene text retrieval has made significant progress with the assistance of accurate text localization. However, existing approaches typically require costly bounding box annotations for training. Besides, they mostly adopt a customized retrieval strategy but struggle to unify various types of queries to meet diverse retrieval needs. To address these issues, we introduce Muti-query Scene Text retrieval with Attention Recycling (MSTAR), a box-free approach for scene text retrieval. It incorporates progressive vision embedding to dynamically capture the multi-grained representation of texts and harmonizes free-style text queries with style-aware instructions. Additionally, a multi-instance matching module is integrated to enhance vision-language alignment. Furthermore, we build the Multi-Query Text Retrieval (MQTR) dataset, the first benchmark designed to evaluate the multi-query scene text retrieval capability of models, comprising four query types and 16k images. Extensive experiments demonstrate the superiority of our method across seven public datasets and the MQTR dataset. Notably, MSTAR marginally surpasses the previous state-of-the-art model by 6.4% in MAP on Total-Text while eliminating box annotation costs. Moreover, on the MQTR benchmark, MSTAR significantly outperforms the previous models by an average of 8.5%. The code and datasets are available at https://github.com/yingift/MSTAR.

cs.CV

VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning

Recent strides in multimodal large language models (MLLMs) have significantly advanced their performance in many reasoning tasks. However, Abstract Visual Reasoning (AVR) remains a critical challenge, primarily due to limitations in perceiving abstract graphics. To tackle this issue, we investigate the bottlenecks in current MLLMs and synthesize training data to improve their abstract visual perception. First, we propose VisuRiddles, a benchmark for AVR, featuring tasks meticulously constructed to assess models' reasoning capacities across five core dimensions and two high-level reasoning categories. Second, we introduce the Perceptual Riddle Synthesizer (PRS), an automated framework for generating riddles with fine-grained perceptual descriptions. PRS not only generates valuable training data for abstract graphics but also provides fine-grained perceptual description, crucially allowing for supervision over intermediate reasoning stages and thereby improving both training efficacy and model interpretability. Our extensive experimental results on VisuRiddles empirically validate that fine-grained visual perception is the principal bottleneck and our synthesis framework markedly enhances the performance of contemporary MLLMs on these challenging tasks. Our code and dataset will be released at https://github.com/yh-hust/VisuRiddles

cs.CV

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Multimodal document understanding is a challenging task to process and comprehend large amounts of textual and visual information. Recent advances in Large Language Models (LLMs) have significantly improved the performance of this task. However, existing methods typically focus on either plain text or a limited number of document images, struggling to handle long PDF documents with interleaved text and images, especially for academic papers. In this paper, we introduce PDF-WuKong, a multimodal large language model (MLLM) that is designed to enhance multimodal question-answering (QA) for long PDF documents. PDF-WuKong incorporates a sparse sampler that operates on both text and image representations, significantly improving the efficiency and capability of the MLLM. The sparse sampler selects the paragraphs or diagrams most pertinent to user queries. To effectively train and evaluate our model, we construct PaperPDF, a dataset consisting of a broad collection of English and Chinese academic papers. Multiple strategies are proposed to build high-quality 1.1 million QA pairs along with their corresponding evidence sources. Experimental results demonstrate the superiority and high efficiency of our approach over other models on the task of long multimodal document understanding, surpassing proprietary products by an average of 8.6% on F1. Our code and dataset will be released at https://github.com/yh-hust/PDF-Wukong.

cs.CV

Contrastive Diffuser: Planning Towards High Return States via Contrastive Learning

The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (CDiffuser) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, CDiffuser groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Experiments on 14 commonly used D4RL benchmarks demonstrate the effectiveness of our proposed method. Our code is publicly available at \url{https://anonymous.4open.science/r/CDiffuser}.

cs.LG

GMOCAT: A Graph-Enhanced Multi-Objective Method for Computerized Adaptive Testing

Computerized Adaptive Testing(CAT) refers to an online system that adaptively selects the best-suited question for students with various abilities based on their historical response records. Most CAT methods only focus on the quality objective of predicting the student ability accurately, but neglect concept diversity or question exposure control, which are important considerations in ensuring the performance and validity of CAT. Besides, the students' response records contain valuable relational information between questions and knowledge concepts. The previous methods ignore this relational information, resulting in the selection of sub-optimal test questions. To address these challenges, we propose a Graph-Enhanced Multi-Objective method for CAT (GMOCAT). Firstly, three objectives, namely quality, diversity and novelty, are introduced into the Scalarized Multi-Objective Reinforcement Learning framework of CAT, which respectively correspond to improving the prediction accuracy, increasing the concept diversity and reducing the question exposure. We use an Actor-Critic Recommender to select questions and optimize three objectives simultaneously by the scalarization function. Secondly, we utilize the graph neural network to learn relation-aware embeddings of questions and concepts. These embeddings are able to aggregate neighborhood information in the relation graphs between questions and concepts. We conduct experiments on three real-world educational datasets, and show that GMOCAT not only outperforms the state-of-the-art methods in the ability prediction, but also achieve superior performance in improving the concept diversity and alleviating the question exposure. Our code is available at https://github.com/justarter/GMOCAT.

cs.IR

PolarDB-IMCI: A Cloud-Native HTAP Database System at Alibaba

Cloud-native databases have become the de-facto choice for mission-critical applications on the cloud due to the need for high availability, resource elasticity, and cost efficiency. Meanwhile, driven by the increasing connectivity between data generation and analysis, users prefer a single database to efficiently process both OLTP and OLAP workloads, which enhances data freshness and reduces the complexity of data synchronization and the overall business cost. In this paper, we summarize five crucial design goals for a cloud-native HTAP database based on our experience and customers' feedback, i.e., transparency, competitive OLAP performance, minimal perturbation on OLTP workloads, high data freshness, and excellent resource elasticity. As our solution to realize these goals, we present PolarDB-IMCI, a cloud-native HTAP database system designed and deployed at Alibaba Cloud. Our evaluation results show that PolarDB-IMCI is able to handle HTAP efficiently on both experimental and production workloads; notably, it speeds up analytical queries up to $\times149$ on TPC-H (100 $GB$). PolarDB-IMCI introduces low visibility delay and little performance perturbation on OLTP workloads (< 5%), and resource elasticity can be achieved by scaling out in tens of seconds.

cs.DB

Parallel finite volume simulation of the spherical shell dynamo with pseudo-vacuum magnetic boundary conditions

In this paper, we study the parallel simulation of the magnetohydrodynamic (MHD) dynamo in a rapidly rotating spherical shell with pseudo-vacuum magnetic boundary conditions. A second-order finite volume scheme based on a collocated quasi-uniform cubed-sphere grid is applied to the spatial discretization of the MHD dynamo equations. To ensure the solenoidal condition of the magnetic field, we adopt a widely-used approach whereby a pseudo-pressure is introduced into the induction equation. The temporal integration is split by a second-order approximate factorization approach, resulting in two linear algebraic systems both solved by a preconditioned Krylov subspace iterative method. A multi-level restricted additive Schwarz preconditioner based on domain decomposition and multigrid method is then designed to improve the efficiency and scalability. Accurate numerical solutions of two benchmark cases are obtained with our code, comparable to the existing local method results. Several large-scale tests performed on the Sunway TaihuLight supercomputer show good strong and weak scalabilities and a noticeable improvement from the multi-level preconditioner with up to 10368 processor cores.

physics.comp-ph

Heterogeneous information network model for equipment-standard system

Entity information network is used to describe structural relationships between entities. Taking advantage of its extension and heterogeneity, entity information network is more and more widely applied to relationship modeling. Recent years, lots of researches about entity information network modeling have been proposed, while seldom of them concentrate on equipment-standard system with properties of multi-layer, multi-dimension and multi-scale. In order to efficiently deal with some complex issues in equipment-standard system such as standard revising, standard controlling, and production designing, a heterogeneous information network model for equipment-standard system is proposed in this paper. Three types of entities and six types of relationships are considered in the proposed model. Correspondingly, several different similarity-measuring methods are used in the modeling process. The experiments show that the heterogeneous information network model established in this paper can reflect relationships between entities accurately. Meanwhile, the modeling process has a good performance on time consumption.

cs.IR

Bose glass and Mott glass of quasiparticles in a doped quantum magnet

The low-temperature states of bosonic fluids exhibit fundamental quantum effects at the macroscopic scale: the best-known examples are Bose-Einstein condensation (BEC) and superfluidity, which have been tested experimentally in a variety of different systems. When bosons are interacting, disorder can destroy condensation leading to a so-called Bose glass. This phase has been very elusive to experiments due to the absence of any broken symmetry and of a finite energy gap in the spectrum. Here we report the observation of a Bose glass of field-induced magnetic quasiparticles in a doped quantum magnet (Br-doped dichloro-tetrakis-thiourea-Nickel, DTN). The physics of DTN in a magnetic field is equivalent to that of a lattice gas of bosons in the grand-canonical ensemble; Br-doping introduces disorder in the hoppings and interaction strengths, leading to localization of the bosons into a Bose glass down to zero field, where it acquires the nature of an incompressible Mott glass. The transition from the Bose glass (corresponding to a gapless spin liquid) to the BEC (corresponding to a magnetically ordered phase) is marked by a novel, universal exponent governing the scaling on the critical temperature with the applied field, in excellent agreement with theoretical predictions. Our study represents the first, quantitative account of the universal features of disordered bosons in the grand-canonical ensemble.

cond-mat.str-el

Dissipation in the superconducting state of kappa-(BEDT-TTF)2Cu(NCS)2

We have studied the interlayer resistivity of the prototypical quasi-two-dimensional organic superconductor $κ$-(BEDT-TTF)$_2$Cu(NCS)$_2$ as a function of temperature, current and magnetic field, within the superconducting state. We find a region of non-zero resistivity whose properties are strongly dependent on magnetic field and current density. There is a crossover to non-Ohmic conduction below a temperature that coincides with the 2D vortex solid -- vortex liquid transition. We interpret the behaviour in terms of a model of current- and thermally-driven phase slips caused by the diffusive motion of the pancake vortices which are weakly-coupled in adjacent layers, giving rise to a finite interlayer resistance.

cond-mat.supr-con