SearcharxivSearch

arXiv subjects

Tao Zeng

Publications and source records attributed to Tao Zeng.

At least 19 recordsLinked to original sources

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

cs.AI

Entropy-Aware Structural Alignment for Zero-Shot Handwritten Chinese Character Recognition

Zero-shot Handwritten Chinese Character Recognition (HCCR) aims to recognize unseen characters by leveraging radical-based semantic compositions. However, existing approaches often treat characters as flat radical sequences, neglecting the hierarchical topology and the uneven information density of different components. To address these limitations, we propose an Entropy-Aware Structural Alignment Network that bridges the visual-semantic gap through information-theoretic modeling. First, we introduce an Information Entropy Prior to dynamically modulate positional embeddings via multiplicative interaction, acting as a saliency detector that prioritizes discriminative roots over ubiquitous components. Second, we construct a Dual-View Radical Tree to extract multi-granularity structural features, which are integrated via an adaptive Sigmoid-based gating network to encode both global layout and local spatial roles. Finally, a Top-K Semantic Feature Fusion mechanism is devised to augment the decoding process by utilizing the centroid of semantic neighbors, effectively rectifying visual ambiguities through feature-level consensus. Extensive experiments demonstrate that our method establishes new state-of-the-art performance, achieving an accuracy of 55.04\% on the ICDAR 2013 dataset ($m=1500$), significantly outperforming existing CLIP-based baselines in the challenging zero-shot setting. Furthermore, the framework exhibits exceptional data efficiency, demonstrating rapid adaptability with minimal support samples, achieving 92.41\% accuracy with only one support sample per class.

cs.CV

ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering

This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. The growing demand for sophisticated video analysis capabilities highlights the critical need for comprehensive, culturally-aware evaluation frameworks. ChineseVideoBench addresses this gap by providing a robust dataset and tailored evaluation metrics, enabling rigorous assessment of state-of-the-art MLLMs on complex Chinese video content. Specifically, ChineseVideoBench comprises 8 main classes and 12 sub-classes, encompassing tasks that demand both deep video understanding and nuanced Chinese linguistic and cultural awareness. Our empirical evaluations reveal that ChineseVideoBench presents a significant challenge to current MLLMs. Among the models assessed, Gemini 2.5 Pro achieves the highest performance with an overall score of 77.9%, while InternVL-38B emerges as the most competitive open-source model.

cs.CV

On the Feasibility of Exact Unitary Transformations for Many-body Hamiltonians

Exact unitary transformations play a central role in the analysis and simulation of many-body quantum systems, yet the conditions under which they can be carried out exactly and efficiently remain incompletely understood. We show that exact transformations arise whenever the adjoint action of a unitary's generator defines a linear map within a finite-dimensional operator space. In this regime, there exists a finite-degree polynomial that annihilates the adjoint map, rendering the Baker-Campbell-Hausdorff (BCH) expansion finite. We identify the role of Lie algebras and their modules in producing finite BCH expansions in all known cases. This perspective brings together previously disparate examples of exact transformations under a single unifying principle and clarifies how algebraic relations between generators and transformed operators determine the polynomial degree of the transformation. We illustrate this framework for previously known cases of efficient unitary transformations including unitary coupled-cluster and Pauli product generators. Using this framework, we propose a new class of fermionic generators that can be used for efficient transformations. The result establishes sufficient algebraic conditions for when exact unitary transformations are possible and provides new strategies for reducing their computational cost in quantum simulation and constructing feasible unitary transformations.

quant-ph

DeRainMamba: A Frequency-Aware State Space Model with Detail Enhancement for Image Deraining

Image deraining is crucial for improving visual quality and supporting reliable downstream vision tasks. Although Mamba-based models provide efficient sequence modeling, their limited ability to capture fine-grained details and lack of frequency-domain awareness restrict further improvements. To address these issues, we propose DeRainMamba, which integrates a Frequency-Aware State-Space Module (FASSM) and Multi-Directional Perception Convolution (MDPConv). FASSM leverages Fourier transform to distinguish rain streaks from high-frequency image details, balancing rain removal and detail preservation. MDPConv further restores local structures by capturing anisotropic gradient features and efficiently fusing multiple convolution branches. Extensive experiments on four public benchmarks demonstrate that DeRainMamba consistently outperforms state-of-the-art methods in PSNR and SSIM, while requiring fewer parameters and lower computational costs. These results validate the effectiveness of combining frequency-domain modeling and spatial detail enhancement within a state-space framework for single image deraining.

cs.CV

Quantum Seniority-based Subspace Expansion: Linear Combinations of Short-Circuit Unitary Transformations for the Electronic Structure Problem

Quantum SENiority-based Subspace Expansion (Q-SENSE) is a hybrid quantum-classical algorithm that interpolates between the Variational Quantum Eigensolver (VQE) and Configuration Interaction (CI) methods. It constructs Hamiltonian matrix elements on a quantum device and solves the resulting eigenvalue problem classically. Unlike other expansion-based methods -- such as Quantum Subspace Expansion (QSE), Quantum Krylov Algorithms, and the Non-Orthogonal Quantum Eigensolver -- Q-SENSE introduces seniority operators as artificial symmetries to construct orthogonal basis states. This seniority-symmetry-based approach reduces one of the primary limitations of VQE on near-term quantum hardware -- circuit depth -- at the cost of measuring additional matrix elements. The artificial symmetries also reduce the number of Hamiltonian terms that must be measured, as only a small fraction of the terms couple basis states in different seniority subspaces. With all these merits, Q-SENSE offers a scalable and resource-efficient route to quantum advantage on near-term quantum devices and in the early fault-tolerant regime.

quant-ph

SUMMA: A Multimodal Large Language Model for Advertisement Summarization

Understanding multimodal video ads is crucial for improving query-ad matching and relevance ranking on short video platforms, enhancing advertising effectiveness and user experience. However, the effective utilization of multimodal information with high commercial value still largely constrained by reliance on highly compressed video embeddings-has long been inadequate. To address this, we propose SUMMA (the abbreviation of Summarizing MultiModal Ads), a multimodal model that automatically processes video ads into summaries highlighting the content of highest commercial value, thus improving their comprehension and ranking in Douyin search-advertising systems. SUMMA is developed via a two-stage training strategy-multimodal supervised fine-tuning followed by reinforcement learning with a mixed reward mechanism-on domain-specific data containing video frames and ASR/OCR transcripts, generating commercially valuable and explainable summaries. We integrate SUMMA-generated summaries into our production pipeline, directly enhancing the candidate retrieval and relevance ranking stages in real search-advertising systems. Both offline and online experiments show substantial improvements over baselines, with online results indicating a statistically significant 1.5% increase in advertising revenue. Our work establishes a novel paradigm for condensing multimodal information into representative texts, effectively aligning visual ad content with user query intent in retrieval and recommendation scenarios.

cs.IR

Comparing Misspecified Models with Big Data: A Variational Bayesian Perspective

Optimal data detection in massive multiple-input multiple-output (MIMO) systems often requires prohibitively high computational complexity. A variety of detection algorithms have been proposed in the literature, offering different trade-offs between complexity and detection performance. In recent years, Variational Bayes (VB) has emerged as a widely used method for addressing statistical inference in the context of massive data. This study focuses on misspecified models and examines the risk functions associated with predictive distributions derived from variational posterior distributions. These risk functions, defined as the expectation of the Kullback-Leibler (KL) divergence between the true data-generating density and the variational predictive distributions, provide a framework for assessing predictive performance. We propose two novel information criteria for predictive model comparison based on these risk functions. Under certain regularity conditions, we demonstrate that the proposed information criteria are asymptotically unbiased estimators of their respective risk functions. Through comprehensive numerical simulations and empirical applications in economics and finance, we demonstrate the effectiveness of these information criteria in comparing misspecified models in the context of massive data.

econ.EM

Quantum Algorithm for Vibronic Dynamics: Case Study on Singlet Fission Solar Cell Design

Vibronic interactions between nuclear motion and electronic states are critical for the accurate modeling of photochemistry. However, accurate simulations of fully quantum non-adiabatic dynamics are often prohibitively expensive for classical methods beyond small systems. In this work, we present a quantum algorithm based on product formulas for simulating time evolution under a general vibronic Hamiltonian in real space, capable of handling an arbitrary number of electronic states and vibrational modes. We develop the first trotterization scheme for vibronic Hamiltonians beyond two electronic states and introduce an array of optimization techniques for the exponentiation of each fragment in the product formula, resulting in a remarkably low cost of implementation. To demonstrate practical relevance, we outline a proof-of-principle integration of our algorithm into a materials discovery pipeline for designing more efficient singlet fission-based organic solar cells. We estimate that $100$ fs of propagation using a second-order Trotter product formula for a $6$-state, $21$-mode model of exciton transport at an anthracene dimer requires $154$ qubits and $2.76 \times 10^6$ Toffoli gates. While a $4$-state, $246$-mode model describing charge transfer at an anthracene-fullerene interface requires $1053$ qubits and $2.66 \times 10^7$ Toffoli gates.

quant-ph

Universal low-temperature fluctuation of unconventional superconductors revealed: 'Smoking gun' leaves proper bosonic superfluidity the last theory standing

Low-temperature thermal fluctuations offer an essential window in characterizing the true nature of a quantum state of matter, a quintessential example being Fermi liquid theory. Here, we examine the leading thermal fluctuation of the superfluid density across numerous families ranging from relatively conventional to highly unconventional superconductors (MgB$_2$, bismuthates, doped buckyballs, heavy fermions, UTe$_2$, doped SrTiO$_3$, Chevrel clusters, intermetallics, organic superconductors, transition metal dichalcogenides, ruthenates, iron-pnictides, cuprates, and kagome metals). Amazingly, in all of them an unprecedented universal $T^3$ depletion materializes in the low-temperature superfluid density, even in the believed-to-be-conventional MgB$_2$. This reveals a new quantum superfluid state of matter and requires a necessary change of paradigm in describing modern superconductors. We demonstrate that such unorthodox yet generic behavior can be described by a strictly Galilean consistent theory of bosonic superfluidity hosting a long-lived 'true condensate'.

cond-mat.supr-con

ChatASU: Evoking LLM's Reflexion to Truly Understand Aspect Sentiment in Dialogues

Aspect Sentiment Understanding (ASU) in interactive scenarios (e.g., Question-Answering and Dialogue) has attracted ever-more interest in recent years and achieved important progresses. However, existing studies on interactive ASU largely ignore the coreference issue for opinion targets (i.e., aspects), while this phenomenon is ubiquitous in interactive scenarios especially dialogues, limiting the ASU performance. Recently, large language models (LLMs) shows the powerful ability to integrate various NLP tasks with the chat paradigm. In this way, this paper proposes a new Chat-based Aspect Sentiment Understanding (ChatASU) task, aiming to explore LLMs' ability in understanding aspect sentiments in dialogue scenarios. Particularly, this ChatASU task introduces a sub-task, i.e., Aspect Chain Reasoning (ACR) task, to address the aspect coreference issue. On this basis, we propose a Trusted Self-reflexion Approach (TSA) with ChatGLM as backbone to ChatASU. Specifically, this TSA treats the ACR task as an auxiliary task to boost the performance of the primary ASU task, and further integrates trusted learning into reflexion mechanisms to alleviate the LLMs-intrinsic factual hallucination problem in TSA. Furthermore, a high-quality ChatASU dataset is annotated to evaluate TSA, and extensive experiments show that our proposed TSA can significantly outperform several state-of-the-art baselines, justifying the effectiveness of TSA to ChatASU and the importance of considering the coreference and hallucination issues in ChatASU.

cs.CL

Effective electrical manipulation of topological antiferromagnet by orbital Hall effect

Electrical control of the non-trivial topology in Weyl antiferromagnet is of great interests to develop next-generation spintronic devices. Recent works suggest that spin Hall effect can switch the topological antiferromagnetic order. However, the switching efficiency remains relatively low. Here, we demonstrate effective manipulation of antiferromagnetic order in Weyl semimetal Mn3Sn by orbital Hall effect originated from metal Mn or oxide CuOx. While Mn3Sn is proven to be able to convert orbit current to spin current by itself, we find that inserting a heavy metal layer like Pt with proper thickness can effectively reduce the critical switching current density by one order of magnitude. In addition, we show that the memristor-like switching behavior of Mn3Sn can mimic the potentiation and depression processes of a synapse with high linearity, which is beneficial for constructing artificial neural network with high accuracy. Our work paves an alternative way to manipulate topological antiferromagnetic order and may inspire more high-performance antiferromagnetic functional devices.

cond-mat.mtrl-sci

Transport in the emergent Bose liquid: Bad metal, strange metal, and weak insulator, all in one system

Non-saturating high-temperature resistivity ("bad metal"), T-linear low-temperature resistivity ("strange metal"), and a crossover to activation-free growth of the resistivity in the low-temperature limit ("weak insulator") are among the most exotic behaviors widely observed in many strongly correlated materials for decades that defy the standard Fermi liquid description of solids. Here we investigate these puzzling behaviors by computing temperature-dependent optical conductivity of an emergent Bose liquid and find that it reproduces all the unexplained features of the experiments, including a featureless continuum and a well-known mid-infrared peak. Amazingly and with physically intuitive mechanisms, the corresponding doping- and temperature-dependent resistivity displays the bad metal and strange metal simultaneously and sometimes weak insulating behaviors as well. The unification of all these non-Fermi liquid behaviors in a single model suggests that a new quantum state of matter, namely the emergent Bose liquid, will guide the development of the next generation of solid state physics.

cond-mat.str-el

Fedlearn-Algo: A flexible open-source privacy-preserving machine learning platform

In this paper, we present Fedlearn-Algo, an open-source privacy preserving machine learning platform. We use this platform to demonstrate our research and development results on privacy preserving machine learning algorithms. As the first batch of novel FL algorithm examples, we release vertical federated kernel binary classification model and vertical federated random forest model. They have been tested to be more efficient than existing vertical federated learning models in our practice. Besides the novel FL algorithm examples, we also release a machine communication module. The uniform data transfer interface supports transferring widely used data formats between machines. We will maintain this platform by adding more functional modules and algorithm examples. The code is available at https://github.com/fedlearnAI/fedlearn-algo.

cs.LG

Transformer-based Conditional Variational Autoencoder for Controllable Story Generation

We investigate large-scale latent variable models (LVMs) for neural story generation -- an under-explored application for open-domain long text -- with objectives in two threads: generation effectiveness and controllability. LVMs, especially the variational autoencoder (VAE), have achieved both effective and controllable generation through exploiting flexible distributional latent representations. Recently, Transformers and its variants have achieved remarkable effectiveness without explicit latent representation learning, thus lack satisfying controllability in generation. In this paper, we advocate to revive latent variable modeling, essentially the power of representation learning, in the era of Transformers to enhance controllability without hurting state-of-the-art generation effectiveness. Specifically, we integrate latent representation vectors with a Transformer-based pre-trained architecture to build conditional variational autoencoder (CVAE). Model components such as encoder, decoder and the variational posterior are all built on top of pre-trained language models -- GPT2 specifically in this paper. Experiments demonstrate state-of-the-art conditional generation ability of our model, as well as its excellent representation learning capability and controllability.

cs.CL

BioNavi-NP: Biosynthesis Navigator for Natural Products

Nature, a synthetic master, creates more than 300,000 natural products (NPs) which are the major constituents of FDA-proved drugs owing to the vast chemical space of NPs. To date, there are fewer than 30,000 validated NPs compounds involved in about 33,000 known enzyme catalytic reactions, and even fewer biosynthetic pathways are known with complete cascade-connected enzyme catalysis. Therefore, it is valuable to make computer-aided bio-retrosynthesis predictions. Here, we develop BioNavi-NP, a navigable and user-friendly toolkit, which is capable of predicting the biosynthetic pathways for NPs and NP-like compounds through a novel (AND-OR Tree)-based planning algorithm, an enhanced molecular Transformer neural network, and a training set that combines general organic transformations and biosynthetic steps. Extensive evaluations reveal that BioNavi-NP generalizes well to identifying the reported biosynthetic pathways for 90% of test compounds and recovering the verified building blocks for 73%, significantly outperforming conventional rule-based approaches. Moreover, BioNavi-NP also shows an outstanding capacity of biologically plausible pathways enumeration. In this sense, BioNavi-NP is a leading-edge toolkit to redesign complex biosynthetic pathways of natural products with applications to total or semi-synthesis and pathway elucidation or reconstruction.

q-bio.QM

Outline to Story: Fine-grained Controllable Story Generation from Cascaded Events

Large-scale pretrained language models have shown thrilling generation capabilities, especially when they generate consistent long text in thousands of words with ease. However, users of these models can only control the prefix of sentences or certain global aspects of generated text. It is challenging to simultaneously achieve fine-grained controllability and preserve the state-of-the-art unconditional text generation capability. In this paper, we first propose a new task named "Outline to Story" (O2S) as a test bed for fine-grained controllable generation of long text, which generates a multi-paragraph story from cascaded events, i.e. a sequence of outline events that guide subsequent paragraph generation. We then create dedicate datasets for future benchmarks, built by state-of-the-art keyword extraction techniques. Finally, we propose an extremely simple yet strong baseline method for the O2S task, which fine tunes pre-trained language models on augmented sequences of outline-story pairs with simple language modeling objective. Our method does not introduce any new parameters or perform any architecture modification, except several special tokens as delimiters to build augmented sequences. Extensive experiments on various datasets demonstrate state-of-the-art conditional story generation performance with our model, achieving better fine-grained controllability and user flexibility. Our paper is among the first ones by our knowledge to propose a model and to create datasets for the task of "outline to story". Our work also instantiates research interest of fine-grained controllable generation of open-domain long text, where controlling inputs are represented by short text.

cs.CL

Ontology-based systematic classification and analysis of coronaviruses, hosts, and host-coronavirus interactions towards deep understanding of COVID-19

Given the existing COVID-19 pandemic worldwide, it is critical to systematically study the interactions between hosts and coronaviruses including SARS-Cov, MERS-Cov, and SARS-CoV-2 (cause of COVID-19). We first created four host-pathogen interaction (HPI)-Outcome postulates, and generated a HPI-Outcome model as the basis for understanding host-coronavirus interactions (HCI) and their relations with the disease outcomes. We hypothesized that ontology can be used as an integrative platform to classify and analyze HCI and disease outcomes. Accordingly, we annotated and categorized different coronaviruses, hosts, and phenotypes using ontologies and identified their relations. Various COVID-19 phenotypes are hypothesized to be caused by the backend HCI mechanisms. To further identify the causal HCI-outcome relations, we collected 35 experimentally-verified HCI protein-protein interactions (PPIs), and applied literature mining to identify additional host PPIs in response to coronavirus infections. The results were formulated in a logical ontology representation for integrative HCI-outcome understanding. Using known PPIs as baits, we also developed and applied a domain-inferred prediction method to predict new PPIs and identified their pathological targets on multiple organs. Overall, our proposed ontology-based integrative framework combined with computational predictions can be used to support fundamental understanding of the intricate interactions between human patients and coronaviruses (including SARS-CoV-2) and their association with various disease outcomes.

q-bio.OT