SearcharxivSearch

arXiv subjects

Lin Zhang

Publications and source records attributed to Lin Zhang.

At least 19 recordsLinked to original sources

Bargmann Invariants Are Insufficient for Complete Local-Unitary Orbit Discrimination

Bargmann invariants constructed from a bipartite density operator and its two lifted marginals are polynomial invariants of local-unitary conjugation. We determine the precise information encoded in these invariants. Whenever one subsystem is a qubit, the ordinary marginal-word family determines the full spectrum of the partial transpose and hence decides whether the state has the positive-partial-transpose property. In $2\otimes2$ and $2\otimes3$ systems, this yields complete separability criteria. For the two-qubit system, a finite subfamily additionally separates local-unitary orbits, and a finite extension generates the polynomial invariant ring. These three tasks already diverge for qubit-qutrit states: we exhibit full-rank, locally maximally mixed states that agree on all ordinary marginal-word invariants yet have different operator-Schmidt ranks, together with a quartic correlation invariant that separates them. When both local dimensions are at least three, the analogous collapse on the locally maximally mixed sector produces isospectral pairs consisting of one separable state and one entangled state with negative partial transpose. The ordinary Bargmann algebra therefore coincides with the full local unitary invariant ring if and only if both subsystems are qubits. The missing data are geometric: they encode the placement of global eigenspaces relative to the tensor-product decomposition

quant-ph

A Mean-Field Approach to the Dielectric Response of Bulk Superconductors for Light Dark Matter Direct Detection

The dielectric function is central to describing many-body screening effects in dark matter (DM) direct detection with condensed matter targets. Current superconducting detector analyses employ the free-electron Lindhard dielectric function to model in-medium effects, an approximation whose validity in the superconducting state remains untested. We derive the electronic dielectric function for bulk superconductors within the Bardeen-Cooper-Schrieffer (BCS) framework, incorporating the full Bogoliubov quasiparticle coherence factors in the random-phase approximation. A systematic comparison with the Lindhard function for aluminum and tungsten silicide (WSi) reveals good agreement for energy depositions $\omega\gtrsim5\Delta$, establishing the Lindhard function as a robust approximation for superconducting DM detectors operating in this regime.

hep-ph

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .

cs.CL

Marginal spectral distributions on regular bipartite unitary orbits

Fix the spectrum of a bipartite density matrix and randomize its eigenbasis according to Haar measure. We study the probability distributions induced on the spectra of the two marginal states. For arbitrary subsystem dimensions $m$ and $n$, the joint characteristic function of the reduced density matrices is expressed as a Harish-Chandra-Itzykson-Zuber integral whose external eigenvalues are the pairwise sums $x_i + y_j$. Repeated external eigenvalues are handled by confluent determinant limits. In the two-qubit case, we derive an explicit alternating-spline formula for the joint density of the two marginal Bloch radii. Its support is the Bravyi-Klyachko compatibility region. We also obtain a compact truncated-power formula for the Bloch-radius density of either individual qubit marginal. In the qubit-qutrit case, we derive a truncated-power formula for the qubit Bloch-radius density and a bivariate spline formula for the joint density of the largest and smallest eigenvalues of the qutrit marginal. The latter two variables determine the full qutrit spectrum because the trace is fixed. The derivations combine confluent HCIZ integrals, distributional Fourier inversion, orbital measures, and the SU(2) and SU(3) derivative principles. The resulting densities are piece-wise polynomial on chambers determined by subset sums of the fixed global eigenvalues, in agreement with the Duistermaat-Heckman description of projected coadjoint-orbit measures.

quant-ph

Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.

cs.SD

From Multi-Modal Paths to Executable Trajectories: A Trajectory Planning Framework for 4WIS Robots

Four-wheel independent steering (4WIS) mobile robots support multiple motion modes, offering high maneuverability in narrow and complex environments. However, existing planning methods often fail to fully exploit these capabilities, leading to suboptimal trajectory quality. To address this limitation, this paper proposes a multi-modal global trajectory planning framework that couples mode-augmented front-end search with mode-consistent segment-wise trajectory optimization. In the front-end stage, Hybrid A* is extended to a four-dimensional state space incorporating motion modes, while mode-switching-aware cost and heuristic functions embed mode decisions into the global search process. Multi-modal Reeds-Shepp curves and an intelligent terminal connection strategy are further designed to improve search efficiency. In the back-end stage, a segment-wise trajectory optimization framework based on an improved iterative safe corridor scheme is developed to convert discrete multi-modal paths into smooth, kinematically feasible trajectories with stationary mode transitions. Experimental results show that the proposed method achieves the best overall performance in safety, arrival time, terminal accuracy and computation time. Real-world experiments on a physical 4WIS robot further validate the practical effectiveness and executability of the generated trajectories, providing a flexible and high-performance solution for multi-modal mobile robot trajectory planning.

cs.RO

Bound-state-mediated remote charging of a quantum battery

Remote charging of a quantum battery (QB) is hindered by radiative leakage of the excitation into the photonic environment that acts as a mediator for energy transfer. We consider a charger-battery model consisting of two two-level systems (TLSs) that are locally coupled to two sites of a one-dimensional coupled cavity array. When their transition frequency lies outside the propagation band, the system forms atom-photon bound states with localized photonic components, and the overlap of these components lifts the degeneracy of the even- and odd-parity bound states, yielding an energy splitting that drives coherent energy transfer from the charger to the QB. In this way, the band gap suppresses resonant emission and the localized bound states mediate remote charging. From the parity-resolved spectrum, we relate the charging time to the energy splitting and the charged ergotropy to the fraction of TLS population on the bound states. Bound states closer to the band edge extend the interaction range of the TLSs but contain a large photonic fraction and are consequently more susceptible to photon loss.

quant-ph

Dielectric Response for Light Dark Matter Direct Detection Beyond the Longitudinal Approximation

The dielectric formalism for light dark matter--electron scattering in semiconductors has, to date, employed only the longitudinal dielectric function $\epsilon_L$, with the transverse response $\epsilon_T$ universally neglected on qualitative grounds. A complete derivation and quantitative evaluation of $\epsilon_T$ in this context has been lacking. We provide this derivation within the random phase approximation for a homogeneous electron gas. For silicon, we find that the transverse energy loss function is $4--6$ orders of magnitude below the longitudinal one in the bulk plasmon regime, providing the first rigorous justification for the conventional longitudinal approximation. The sizable transverse corrections appear for deposited energies $\omega\lesssim1\,\mathrm{eV}$, which lies below the energy required to reliably produce one electron-hole pair in silicon detectors. Our results provide a quantitative error assessment for existing longitudinal calculations and identify kinematic regimes where transverse corrections ought to be included for relativistic dark-matter interpretations.

hep-ph

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density. Despite growing interest, existing evaluations remain fragmented across individual datasets and narrow tasks, leaving a critical gap in unified assessment of UAV understanding and reasoning capabilities. To fill this gap, we construct UAVQA-Bench, a benchmark of 1,500 human-annotated QA pairs drawn from 13 public UAV datasets, covering 6 capability dimensions and 16 tasks in both multiple-choice and visual grounding formats. Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes: domain-toolset mismatch, unchecked error propagation, and static reasoning. Motivated by these findings, we propose UAV-MAS, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine (DSPE) that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty. UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0\%, while the 8B variant improves 8.7\% over its base model.

cs.CV

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmarks primarily focus on aerial-view scenarios. However, whether current VLMs can perform well on understanding and reasoning tasks in aerial-ground collaborative scenarios which are practical in real-world applications like rescue and infrastructure inspection remains underexplored. To address this gap, we introduce AeroGround, a comprehensive benchmark for evaluating VLMs in aerial-ground collaborative reasoning. AeroGround is built upon a simulated aerial-ground dataset containing approximately 29,000 multimodal observation groups from diverse open environments, and provides 2,250 high-quality question-answering instances covering cross-view correspondence, spatial understanding, and reasoning. Experiments on 16 pretrained VLMs, together with two domain-adapted variants, reveal a substantial gap between current models and human performance: the best model achieves an average accuracy of 54.4%, whereas humans reach 93.3%. By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.

cs.CV

Extreme Value Alpha and Crash Risk: Separating Structural Tails from Lottery Tails with LLM-Extracted Disclosure Networks

A heavy upper tail in a stock's returns is ambiguous: it can be a lottery tail, transient jump risk that investors overpay for (the MAX discount), or a structural tail, the statistical shadow of an economic reconfiguration that precedes extreme winners. Returns alone cannot separate them, so tail heat alone is not an alpha signal. Our discriminator is the firm's disclosure-measured network: a directed, span-grounded graph from 10-K filings via an auditable LLM pipeline, whose rewiring decomposes into edge birth, death, and drift. The central sign pattern: tail heat with network death is the crash side; tail heat with an intact or forming network is where structural tails and historical winners live. Pilot evidence from 24 technology firms (2014-2025) supports the crash side: upper-tail heat interacted with death mass predicts negative forward abnormal returns (monthly t = -2.9; firm-vintage t = -3.9; wild-cluster p = 0.04; robust to two-way clustering and controls), and the same configuration preceded NVIDIA's 2018 and 2022 drawdowns. This death-side signal is immediately useful as a risk-monitoring danger flag. The alpha side is directionally positive but not significant in the pilot and awaits the confirmatory test. A pre-registered replication on 50 random S&P 500 firms failed, defining the boundary: outside coherent ecosystems the disclosure graph nearly vanishes (83% of firm-vintages have zero death mass), so the discriminator exists only where firms densely document counterparties. The confirmatory design is fully pre-specified, with a gatekept primary pair, archived power simulations, positive-only winner labels, selection-corrected benchmarks, and a frozen ecosystem-coherent universe with a density gate; it activates only if the gate passes. If confirmed, tail heat becomes a conditional signal separating crash risk from structural winners.

stat.AP

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval

Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference. We derive the expected attention score distortion caused by rank r key compression and show that it yields a covariance weighted low rank objective. Under a margin condition, controlling this distortion also improves top k recall. The optimal rank r solution has a closed form asymmetric factorization obtained from the SVD of the covariance weighted query key operator. This motivates SAKI, a training free KV cache index that directly preserves attention scores rather than key reconstruction quality. Across LLaMA 3.1 8B, Qwen 2.5 7B, Mistral 7B v0.1, and Llama 3.2 3B, SAKI outperforms key PCA at every tested rank. At rank 32, it removes 13 to 30 percent of PCA's remaining top 64 recall error, including improvements from 0.748 to 0.799 on LLaMA 3.1 8B and from 0.786 to 0.850 on Qwen 2.5 7B. It improves 68 to 89 percent of attention heads per model, with the largest gains in deeper layers. Predicted score MSE reductions closely match empirical measurements, with a Pearson correlation of 0.997, while ablation studies confirm that the gains arise from optimizing the attention score objective rather than covariance weighting alone. Analysis of the scoring operator further explains why weight only, invariant subspace, and key reconstruction methods can be suboptimal. SAKI uses random-matrix theory to separate genuine covariance signal from autocorrelated sampling noise, matching PCA with only 512 calibration tokens and adding value exactly where PCA sees no reliable signal.

cs.LG

Zellige: Moldable Sequence Placement for Mixed Image-Video DiT Training

High-quality video generation requires training Diffusion Transformers (DiTs) jointly on image and video data, posing a mixed-length sequence training problem across GPUs. Existing systems rely on data parallelism (DP), context parallelism (CP), or their combination; we model these designs as disjoint-group placement and prove that they face a fundamental tradeoff between inter-group load imbalance and intra-group communication redundancy. We present Zellige, a moldable sequence placement system that jointly selects each sequence's parallelism configuration and participating ranks. Zellige consists of three components: a hardware profiler that estimates the execution time and memory consumption of candidate placements, a two-stage planner that balances compute-heavy anchor sequences and packs lighter filler sequences into the remaining capacity, and a coalesced attention engine that efficiently executes whole sequences alongside distributed-attention shards. Across 21 plans, the hardware profile predicts step makespan and peak allocated memory with mean absolute percentage errors of $3.4%$ and $1.5%$, respectively. The two-stage planner solves each batch in 33--119 ms, significantly faster than a joint-placement reference that optimizes all sequences together, while their modeled makespans differ by at most $0.32%$. In end-to-end evaluations, Zellige outperforms KnapFormer by $1.12$--$1.48\times$ on 16 A800 GPUs and $1.27$--$1.54\times$ on 32 A6000 GPUs.

cs.DC

ISAC-Enabled On-Demand UAV Charging for Wireless Rechargeable Sensor Networks

Unmanned aerial vehicles (UAVs) equipped with wireless power transfer (WPT) extend the lifetime of wireless rechargeable sensor networks (WRSNs) by delivering energy on demand. This article presents an integrated sensing and communication (ISAC)-enabled on-demand UAV charging framework coordinated by a central base station. A prioritized charging queue captures node urgency and service cost through residual energy, traffic load, estimated UAV travel time, and flight-direction alignment. This bidirectional coupling ensures that scheduling decisions shape the UAV trajectory, while updated mobility estimates from ISAC dynamically reorder the queue. ISAC-assisted estimation of UAV distance, speed, and position updates travel-time predictions under mobility uncertainty. A time-allocated partial charging policy distributes limited hover time across queued nodes according to criticality. Simulations show gains in energy usage efficiency, travel distance, and charging delay compared with representative baselines. We discuss deployment considerations, including computational overhead, scalability, and parameter selection, to aid practitioners evaluating the framework for IoT scenarios.

cs.NI

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

cs.NI

Revisiting the invariant ring of two-qubit mixed states

Local unitary equivalence serves as the cornerstone for classifying entanglement in bipartite quantum systems. Mathematically, it reduces to the study of polynomial invariants of the density matrix under the action of local unitary groups. The collection of all such polynomial invariants forms a ring, known as the invariant ring. However, identifying the complete generators of the invariant ring is the central issue. In 2007, for the two-qubit system, King et al fully characterized the structure of the invariant ring and determined its Cohen--Macaulay decomposition. In this paper, we revisit their work, with a focus on the computation of the Molien series and the construction of invariants. On one hand, we rigorously derive the Molien series via explicit contour integration over the maximal torus, filling in all previously omitted computational steps. On the other hand, we systematically construct all invariants using a graphical method, and then reduce the candidate set by applying various identities and algebraic relations, obtaining a generating set consisting of 21 invariants. This paper aims to make this important result more widely accessible to researchers in quantum information and invariant theory through the above discussions.

quant-ph

Beyond Janus Atomic Ordering: High-Throughput First-Principles Search for Hidden MoSO Monolayer Structures

Despite the growing interest in two-dimensional (2D) MoSO systems, existing studies have exclusively focused on conventional Janus structures. In this work, we perform high-throughput first-principles calculations to explore novel stable 2D MoSO monolayers. Combined with random sampling strategy, graph theory and group theory, we successfully screen out three novel non-Janus 2D MoSO monolayers from 1325 candidate structures, namely Reversed 2H-MoSO, Hybrid 2H-MoSO, and Hybrid 1T'-MoSO. Compared with Janus MoSO monolayers, the non-Janus MoSO counterparts possess lower binding energies, varying from -4.38 to -4.51 eV/atom. A systematic combination of dynamic, thermodynamic, and mechanical stability analyses corroborates their excellent structural robustness. Ab initio molecular dynamics (AIMD) simulations confirm their superior thermal resistance, with the structures remaining stable at temperatures beyond 2000 K. Interestingly, unlike the semiconducting Janus MoSO, the Hybrid 1T'-MoSO monolayer exhibits distinct metallic characteristics. Furthermore, we found that strain and curvature can enable controlled phase transitions of MoSO among semiconducting, semimetallic, and metallic phases. More importantly, the Hybrid 1T'-MoSO exhibits favorable HER activity with a Gibbs free energy of -0.002 eV, rendering it a promising candidate for hydrogen evolution catalysis. This work not only expands the family of 2D MoSO materials but also provides a reliable strategy for discovering stable functional 2D materials via high-throughput computation.

cond-mat.mtrl-sci

LLM Latent Edge Measurement: Point-in-Time Economic Graphs for Quantitative Investing from Corporate Disclosures

Standard industry classification systems such as GICS assign each firm to a single sector, but the economic relationships through which shocks propagate, such as supplier agreements, customer concentration, intellectual property licensing, cloud service dependencies, and power purchase contracts frequently cross sector boundaries and are often disclosed only in unstructured text. We formulate the construction of a firm-level adjacency matrix as a measurement problem and propose an LLM based pipeline that extracts a weighted, directed, point in time corporate network from public disclosures. Applied to the most recent 10 K and 10 K filings of 42 Nasdaq 100 constituents, the proposed pipeline produces a network containing 149 directed edges. An adversarial audit confirms 88% of sampled edges with weights of at least 0.1, increasing to 100% when economically plausible but weakly documented relationships are included. Refuted edges are concentrated entirely in the lowest-weight portion of the network. The resulting network is consistent with GICS where sector classifications are informative, exhibiting a 1.9-fold increase in within-sector connectivity, while also recovering economically meaningful cross-sector relationships that standard classifications cannot represent. Examples include nuclear power-purchase agreements connecting utilities with hyper scale technology firms and GPU-cloud dependencies within the emerging AI infrastructure ecosystem. Ablation studies further demonstrate that multi-agent fusion, inverse-document-frequency filtering, and relative thresholding each make measurable contributions to network quality.

stat.AP