SearcharxivSearch

arXiv subjects

Xiaoliang Wang

Publications and source records attributed to Xiaoliang Wang.

At least 19 recordsLinked to original sources

TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models

Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chains of thought but suffer from high inference latency due to autoregressive reasoning. Recent work explores using Small Reasoning Models (SRMs) to accelerate LRM inference. In this paper, we systematically characterize the capability boundaries of SRMs and identify three common types of reasoning risks: (1) path divergence, where SRMs lack the strategic ability to construct an initial plan, causing reasoning to deviate from the most probable path; (2) cognitive overload, where SRMs fail to solve particularly difficult steps; and (3) recovery inability, where SRMs lack robust self-reflection and error correction mechanisms. To address these challenges, we propose TrigReason, a trigger-based collaborative reasoning framework that replaces continuous polling with selective intervention. TrigReason delegates most reasoning to the SRM and activates LRM intervention only when necessary-during initial strategic planning (strategic priming trigger), upon detecting extraordinary overconfidence (cognitive offload trigger), or when reasoning falls into unproductive loops (intervention request trigger). The evaluation results on AIME24, AIME25, and GPQA-D indicate that TrigReason matches the accuracy of full LRMs and SpecReason, while offloading 1.70x - 4.79x more reasoning steps to SRMs. Under edge-cloud conditions, TrigReason reduces latency by 43.9\% and API cost by 73.3\%. Our code is available at \href{https://github.com/QQQ-yi/TrigReason}{https://github.com/QQQ-yi/TrigReason}

cs.AI

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory benchmarks primarily target declarative memory, specifically semantic and episodic types, where all information is explicitly presented in dialogues. In contrast, real-world actions are also governed by non-declarative memory, including habitual and procedural types, and need to be inferred from diverse digital traces. To bridge this gap, we introduce Lifebench, which features densely connected, long-horizon event simulation. It pushes AI agents beyond simple recall, requiring the integration of declarative and non-declarative memory reasoning across diverse and temporally extended contexts. Building such a benchmark presents two key challenges: ensuring data quality and scalability. We maintain data quality by employing real-world priors, including anonymized social surveys, map APIs, and holiday-integrated calendars, thus enforcing fidelity, diversity and behavioral rationality within the dataset. Towards scalability, we draw inspiration from cognitive science and structure events according to their partonomic hierarchy; enabling efficient parallel generation while maintaining global coherence. Performance results show that top-tier, state-of-the-art memory systems reach just 55.2\% accuracy, highlighting the inherent difficulty of long-horizon retrieval and multi-source integration within our proposed benchmark. The dataset and data synthesis code are available at https://github.com/1754955896/LifeBench.

cs.AI

SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens equally, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs of different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model's attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments.

cs.LG

LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation

KV Cache is commonly used to accelerate LLM inference with long contexts, yet its high memory demand drives the need for cache compression. Existing compression methods, however, are largely heuristic and lack dynamic budget allocation. To address this limitation, we introduce a unified framework for cache compression by minimizing information loss in Transformer residual streams. Building on it, we analyze the layer attention output loss and derive a new metric to compare cache entries across heads, enabling layer-wise compression with dynamic head budgets. Additionally, by contrasting cross-layer information, we also achieve dynamic layer budgets. LAVa is the first unified strategy for cache eviction and dynamic budget allocation that, unlike prior methods, does not rely on training or the combination of multiple strategies. Experiments with benchmarks (LongBench, Needle-In-A-Haystack, Ruler, and InfiniteBench) demonstrate its superiority. Moreover, our experiments reveal a new insight: dynamic layer budgets are crucial for generation tasks (e.g., code completion), while dynamic head budgets play a key role in extraction tasks (e.g., extractive QA). As a fully dynamic compression method, LAVa consistently maintains top performance across task types. Our code is available at https://github.com/MGDDestiny/Lava.

cs.LG

Consultant Decoding: Yet Another Synergistic Mechanism

The synergistic mechanism based on Speculative Decoding (SD) has garnered considerable attention as a simple yet effective approach for accelerating the inference of large language models (LLMs). Nonetheless, the high rejection rates require repeated LLMs calls to validate draft tokens, undermining the overall efficiency gain of SD. In this work, we revisit existing verification mechanisms and propose a novel synergetic mechanism Consultant Decoding (CD). Unlike SD, which relies on a metric derived from importance sampling for verification, CD verifies candidate drafts using token-level likelihoods computed solely by the LLM. CD achieves up to a 2.5-fold increase in inference speed compared to the target model, while maintaining comparable generation quality (around 100% of the target model's performance). Interestingly, this is achieved by combining models whose parameter sizes differ by two orders of magnitude. In addition, CD reduces the call frequency of the large target model to below 10%, particularly in more demanding tasks. CD's performance was even found to surpass that of the large target model, which theoretically represents the upper bound for speculative decoding.

cs.CL

daDPO: Distribution-Aware DPO for Distilling Conversational Abilities

Large language models (LLMs) have demonstrated exceptional performance across various applications, but their conversational abilities decline sharply as model size decreases, presenting a barrier to their deployment in resource-constrained environments. Knowledge distillation with Direct Preference Optimization (dDPO) has emerged as a promising approach to enhancing the conversational abilities of smaller models using a larger teacher model. However, current methods primarily focus on 'black-box' KD, which only uses the teacher's responses, overlooking the output distribution offered by the teacher. This paper addresses this gap by introducing daDPO (Distribution-Aware DPO), a unified method for preference optimization and distribution-based distillation. We provide rigorous theoretical analysis and empirical validation, showing that daDPO outperforms existing methods in restoring performance for pruned models and enhancing smaller LLM models. Notably, in in-domain evaluation, our method enables a 20% pruned Vicuna1.5-7B to achieve near-teacher performance (-7.3% preference rate compared to that of dDPO's -31%), and allows Qwen2.5-1.5B to occasionally outperform its 7B teacher model (14.0% win rate).

cs.LG

Semiparametric Conditional Factor Models in Asset Pricing

We introduce a simple and tractable methodology for estimating semiparametric conditional latent factor models. Our approach disentangles the roles of characteristics in capturing factor betas of asset returns from ``alpha.'' We construct factors by extracting principal components from Fama-MacBeth managed portfolios. Applying this methodology to the cross-section of U.S. individual stock returns, we find compelling evidence of substantial nonzero pricing errors, even though our factors demonstrate superior performance in standard asset pricing tests. Unexplained ``arbitrage'' portfolios earn high Sharpe ratios, which decline over time. Combining factors with these orthogonal portfolios produces out-of-sample Sharpe ratios exceeding 4.

econ.EM

SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding

This paper introduces SynthDoc, a novel synthetic document generation pipeline designed to enhance Visual Document Understanding (VDU) by generating high-quality, diverse datasets that include text, images, tables, and charts. Addressing the challenges of data acquisition and the limitations of existing datasets, SynthDoc leverages publicly available corpora and advanced rendering tools to create a comprehensive and versatile dataset. Our experiments, conducted using the Donut model, demonstrate that models trained with SynthDoc's data achieve superior performance in pre-training read tasks and maintain robustness in downstream tasks, despite language inconsistencies. The release of a benchmark dataset comprising 5,000 image-text pairs not only showcases the pipeline's capabilities but also provides a valuable resource for the VDU community to advance research and development in document image recognition. This work significantly contributes to the field by offering a scalable solution to data scarcity and by validating the efficacy of end-to-end models in parsing complex, real-world documents.

cs.CV

HardTaint: Production-Run Dynamic Taint Analysis via Selective Hardware Tracing

Dynamic taint analysis (DTA), as a fundamental analysis technique, is widely used in security, privacy, and diagnosis, etc. As DTA demands to collect and analyze massive taint data online, it suffers extremely high runtime overhead. Over the past decades, numerous attempts have been made to lower the overhead of DTA. Unfortunately, the reductions they achieved are marginal, causing DTA only applicable to the debugging/testing scenarios. In this paper, we propose and implement HardTaint, a system that can realize production-run dynamic taint tracking. HardTaint adopts a hybrid and systematic design which combines static analysis, selective hardware tracing and parallel graph processing techniques. The comprehensive evaluations demonstrate that HardTaint introduces only around 9% runtime overhead which is an order of magnitude lower than the state-of-the-arts, while without sacrificing any taint detection capability.

cs.CR

Secure Inter-domain Routing and Forwarding via Verifiable Forwarding Commitments

The Internet inter-domain routing system is vulnerable. On the control plane, the de facto Border Gateway Protocol (BGP) does not have built-in mechanisms to authenticate routing announcements, so an adversary can announce virtually arbitrary paths to hijack network traffic; on the data plane, it is difficult to ensure that actual forwarding path complies with the control plane decisions. The community has proposed significant research to secure the routing system. Yet, existing secure BGP protocols (e.g., BGPsec) are not incrementally deployable, and existing path authorization protocols are not compatible with the current Internet routing infrastructure. In this paper, we propose FC-BGP, the first secure Internet inter-domain routing system that can simultaneously authenticate BGP announcements and validate data plane forwarding in an efficient and incrementally-deployable manner. FC-BGP is built upon a novel primitive, name Forwarding Commitment, to certify an AS's routing intent on its directly connected hops. We analyze the security benefits of FC-BGP in the Internet at different deployment rates. Further, we implement a prototype of FC-BGP and extensively evaluate it over a large-scale overlay network with 100 virtual machines deployed globally. The results demonstrate that FC-BGP saves roughly 55% of the overhead required to validate BGP announcements compared with BGPsec, and meanwhile FC-BGP introduces a small overhead for building a globally-consistent view on the desirable forwarding paths.

cs.NI

From RDMA to RDCA: Toward High-Speed Last Mile of Data Center Networks Using Remote Direct Cache Access

In this paper, we conduct systematic measurement studies to show that the high memory bandwidth consumption of modern distributed applications can lead to a significant drop of network throughput and a large increase of tail latency in high-speed RDMA networks.We identify its root cause as the high contention of memory bandwidth between application processes and network processes. This contention leads to frequent packet drops at the NIC of receiving hosts, which triggers the congestion control mechanism of the network and eventually results in network performance degradation. To tackle this problem, we make a key observation that given the distributed storage service, the vast majority of data it receives from the network will be eventually written to high-speed storage media (e.g., SSD) by CPU. As such, we propose to bypass host memory when processing received data to completely circumvent this performance bottleneck. In particular, we design Lamda, a novel receiver cache processing system that consumes a small amount of CPU cache to process received data from the network at line rate. We implement a prototype of Lamda and evaluate its performance extensively in a Clos-based testbed. Results show that for distributed storage applications, Lamda improves network throughput by 4.7% with zero memory bandwidth consumption on storage nodes, and improves network throughput by up 17% and 45% for large block size and small size under the memory bandwidth pressure, respectively. Lamda can also be applied to latency-sensitive HPC applications, which reduces their communication latency by 35.1%.

cs.NI

Chemical principles of instability and self-organization in reacting and diffusive systems

How patterns and structures undergo symmetry breaking and self-organize within biological systems from initially homogeneous states is a key issue for biological development. The activator-inhibitor (AI) mechanism, derived from reaction-diffusion (RD) models, has been widely believed to be the elementary mechanism for biological pattern formation. This mechanism generally requires activators to be self-enhanced and diffuse more slowly than inhibitors. Here, we identify the instability sources of biological systems and derive the self-organization conditions through solving eigenvalues (dispersion relation) of the generalized RD model for two chemicals. We show that both the single AI mechanisms with long-range inhibition and activation are enough to self-organize into fully-expressed domains without the involvement of the inhibitor-inhibitor (II) mechanism, through singly enhancing the difference in self-proliferation rates of activators and inhibitors or weakening the coupling degree between them. When cross diffusion involves, both the self-enhancement and the difference in diffusion coefficients of chemicals are no longer necessary for self-organization, and the patterning mechanism can be extended to semi-inhibitor and II mechanisms. However, we show that the single activator-activator (AA) mechanism is generally unable to self-organize, even if biological domain growth is additionally involved. Moreover, adding an II system after an AI one can produce discrete and bi-stable patterns. We also observe that a higher dimensional space can solely alter the patterning principles derived from a lower dimensional space, which may be due to the instability driven by the higher degree of spatial freedom. Such results provide new insights into biological pattern formation.

physics.bio-ph

ALT: Boosting Deep Learning Performance by Breaking the Wall between Graph and Operator Level Optimizations

Deep learning models rely on highly optimized tensor libraries for efficient inference on heterogeneous hardware. Current deep compilers typically predetermine layouts of tensors and then optimize loops of operators. However, such unidirectional and one-off workflow strictly separates graph-level optimization and operator-level optimization into different system layers, missing opportunities for unified tuning. This paper proposes ALT, a compiler that performs joint graph- and operator-level optimizations for deep models. ALT provides a generic transformation module to manipulate layouts and loops with easy-to-use primitive functions. ALT further integrates an auto-tuning module that jointly optimizes graph-level data layouts and operator-level loops while guaranteeing efficiency. Experimental results show that ALT significantly outperforms state-of-the-art compilers (e.g., Ansor) in terms of both single operator performance (e.g., 1.5x speedup on average) and end-to-end inference performance (e.g., 1.4x speedup on average).

cs.LG

The evolution of cooperation: an evolutionary advantage of individuals impedes the evolution of the population

Range expansion is a universal process in biological systems, and therefore plays a part in biological evolution. Using a quantitative individual-based method based on the stochastic process, we identify that enhancing the inherent self-proliferation advantage of cooperators relative to defectors is a more effective channel to promote the evolution of cooperation in range expansion than weakening the benefit acquisition of defectors from cooperators. With this self-proliferation advantage, cooperators can rapidly colonize virgin space and establish spatial segregation more readily, which acts like a protective shield to further promote the evolution of cooperation in return. We also show that lower cell density and migration rate have a positive effect on the competition of cooperators with defectors. Biological evolution is based on competition between individuals and should therefore favor selfish behaviors. However, we observe a counterintuitive phenomenon that the evolution of a population is impeded by the fitness-enhancing chemotactic movement of individuals. This highlights a conflict between the interests of the individual and the population. The short-sighted selfish behavior of individuals may not be that favored in the competition between populations. Such information provides important implications for the handling of cooperation.

q-bio.PE

Finite volume simulation of arc: pinching arc plasma by high-frequency alternating longitudinal magnetic field

Arc plasmas have promising applications in many fields. To explore their property is of interest. This paper presents detailed pressure-based finite volume simulation of argon arc. In the modeling, the whole cathode region is coupled to electromagnetic calculations to promise the free change of current density at cathode surface. In numerical solutions, the upwind difference scheme is chosen to promise the transport property of convective terms, and the SIMPLE (Semi-Implicit Method for Pressure Linked Equations) algorithm is used to solve thermal pressure. By simulations of the free-burning argon arc, the model shows good agreement with experiment. We observe an interesting phenomenon that argon arc concentrates intensively in the high-frequency alternating longitudinal magnetic field. Different from existing constricting mechanisms, here arc achieves to be pinched through a continuous transition between shrinking and expansion. The underlying mechanism is that via collaborating with arc's motion inertia, the applied high-frequency alternating magnetic field is able to effectively play a "plasma trap" role, which leads the arc plasma to be imprisoned into a narrower space. This may provide a new approach to constrict arc.

physics.plasm-ph

Self-organization principles of cell cycles and gene expressions in the development of cell populations

A big challenge in current biology is to understand the exact self-organization mechanism underlying complex multi-physics coupling developmental processes. With multiscale computations of from subcellular gene expressions to cell population dynamics that is based on first principles, we show that cell cycles can self-organize into periodic stripes in the development of E. coli populations from one single cell, relying on the moving graded nutrient concentration profile, which provides directing positional information for cells to keep their cycle phases in place. Resultantly, the statistical cell cycle distribution within the population is observed to collapse to a universal function and shows a scale invariance. Depending on the radial distribution mode of genetic oscillations in cell populations, a transition between gene patterns is achieved. When an inhibitor-inhibitor gene network is subsequently activated by a gene-oscillatory network, cell populations with zebra stripes can be established, with the positioning precision of cell-fate-specific domains influenced by cells' speed of free motions. Such information may provide important implications for understanding relevant dynamic processes of multicellular systems, such as biological development.

q-bio.CB

Spatial alignment, group strategy and non-kin selection enable the evolution of cooperation

This article considers a mechanism to explain the emergence and evolution of social cooperation. Selfish individuals tend to benefit themselves, which makes it hard for the maintenance of cooperation between unrelated individuals. We propose and validate that a smart group strategy can effectively facilitate the evolution of cooperation, provided cooperators spatially align whilst cooperating at a new level of alliance. The general evolutionary model presented here shows that a non-kin selection effect is a possible cause for cooperation between unrelated individuals and highlights that non-kin selection may be a hallmark of biological evolution.

q-bio.PE

PSF-LO: Parameterized Semantic Features Based Lidar Odometry

Lidar odometry (LO) is a key technology in numerous reliable and accurate localization and mapping systems of autonomous driving. The state-of-the-art LO methods generally leverage geometric information to perform point cloud registration. Furthermore, obtaining point cloud semantic information which can describe the environment more abundantly will help for the registration. We present a novel semantic lidar odometry method based on self-designed parameterized semantic features (PSFs) to achieve low-drift ego-motion estimation for autonomous vehicle in realtime. We first use a convolutional neural network-based algorithm to obtain point-wise semantics from the input laser point cloud, and then use semantic labels to separate the road, building, traffic sign and pole-like point cloud and fit them separately to obtain corresponding PSFs. A fast PSF-based matching enable us to refine geometric features (GeFs) registration, reducing the impact of blurred submap surface on the accuracy of GeFs matching. Besides, we design an efficient method to accurately recognize and remove the dynamic objects while retaining static ones in the semantic point cloud, which are beneficial to further improve the accuracy of LO. We evaluated our method, namely PSF-LO, on the public dataset KITTI Odometry Benchmark and ranked #1 among semantic lidar methods with an average translation error of 0.82% in the test dataset at the time of writing.

cs.CV