Searcharxiv⌕ Search

arXiv subjects

Lin Zhang

Publications and source records attributed to Lin Zhang.

At least 127 records · Page 7Linked to original sources

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding

Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hardware-aware model-system co-design optimized for minimizing decoding costs. Step-3 innovates in two key dimensions: (1) A novel Multi-Matrix Factorization Attention (MFA) mechanism that significantly reduces both KV cache size and computation while maintaining high attention expressiveness, and (2) Attention-FFN Disaggregation (AFD), a distributed inference system that decouples attention and Feed-Forward Network (FFN) layers into specialized subsystems. This co-design achieves unprecedented cost efficiency: Step-3 significantly reduces theoretical decoding costs compared with models like DeepSeek-V3 and Qwen3 MoE 235B, with the gains widening at longer context. Step-3 achieves low cost while activating 38B parameters per token (more than DeepSeek-V3 and Qwen3 MoE 235B), demonstrating that hardware-aligned attention arithmetic intensity, MoE sparsity, and AFD are critical to cost-effectiveness. We perform a head-to-head comparison with DeepSeek-V3 in its favorable scenarios. Our implementation on Hopper GPUs achieves a decoding throughput of up to 4,039 tokens per second per GPU under 50ms TPOT SLA (4K context, FP8, no MTP). It is higher than DeepSeek-V3's 2,324 in the same setup and sets a new Pareto frontier for LLM decoding.

cs.LG↗

Multi-state imaginarity and coherence in qubit systems

Traditionally, the characterization of quantum resources has focused on individual quantum states. Recent literature, however, has increasingly explored the characterization of resources in multi-states (ordered collections of states indexed by a varying parameter). In this work, we provide a unitary-invariant framework to pinpoint imaginarity and coherence in sets of qubit states: we prove that Bloch vectors must be coplanar to be imaginarity-free and colinear to be incoherent, yielding exact rank-based tests of coherence and imaginarity, and closed-form bounds for existing robustness quantifiers, all based on two-state overlaps only. We also show that the set of imaginarity-free multi-states is not convex, and that third-order invariants completely characterize multi-state imaginarity of single-qubits but not of higher-dimensional systems. As our main technical result, we show that every Bargmann invariant of single-qubit states is determined (up to conjugation) by two-state overlaps. Beyond qubits, we give purity and system-agnostic coherence witnesses from equality constraints on higher-order invariants and connect our results to practical protocols: characterization of partial distinguishability, spin-chirality detection, and subchannel discrimination.

quant-ph↗

Packet Header Recognition Utilizing an All-Optical Reservoir Based on Reinforcement-Learning-Optimized Double-Ring Resonator

Optical packet header recognition is an important signal processing task of optical communication networks. In this work, we propose an all-optical reservoir, consisting of integrated double-ring resonators (DRRs) as nodes, for fast and accurate optical packet header recognition. As the delay-bandwidth product (DBP) of the node is a key figure-of-merit in the reservoir, we adopt a deep reinforcement learning algorithm to maximize the DBPs for various types of DRRs, which has the advantage of full parameter space optimization and fast convergence speed. Intriguingly, the optimized DBPs of the DRRs in cascaded, parallel, and embedded configurations reach the same maximum value, which is believed to be the global maximum. Finally, 3-bit and 6-bit packet header recognition tasks are performed with the all-optical reservoir consisting of the optimized cascaded rings, which have greatly reduced chip size and the desired "flat-top" delay spectra. Using this optical computing scheme, word-error rates as low as 5*10-4 and 9*10-4 are achieved for 3-bit and 6-bit packet header recognition tasks, respectively, which are one order of magnitude better than the previously reported values.

eess.SP↗

PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing

Neural speech editing enables seamless partial edits to speech utterances, allowing modifications to selected content while preserving the rest of the audio unchanged. This useful technique, however, also poses new risks of deepfakes. To encourage research on detecting such partially edited deepfake speech, we introduce PartialEdit, a deepfake speech dataset curated using advanced neural editing techniques. We explore both detection and localization tasks on PartialEdit. Our experiments reveal that models trained on the existing PartialSpoof dataset fail to detect partially edited speech generated by neural speech editing models. As recent speech editing models almost all involve neural audio codecs, we also provide insights into the artifacts the model learned on detecting these deepfakes. Further information about the PartialEdit dataset and audio samples can be found on the project page: https://yzyouzhang.com/PartialEdit/index.html.

eess.AS↗

Efficient and Accurate Prompt Optimization: the Benefit of Memory in Exemplar-Guided Reflection

Automatic prompt engineering aims to enhance the generation quality of large language models (LLMs). Recent works utilize feedbacks generated from erroneous cases to guide the prompt optimization. During inference, they may further retrieve several semantically-related exemplars and concatenate them to the optimized prompts to improve the performance. However, those works only utilize the feedback at the current step, ignoring historical and unseleccted feedbacks which are potentially beneficial. Moreover, the selection of exemplars only considers the general semantic relationship and may not be optimal in terms of task performance and matching with the optimized prompt. In this work, we propose an Exemplar-Guided Reflection with Memory mechanism (ERM) to realize more efficient and accurate prompt optimization. Specifically, we design an exemplar-guided reflection mechanism where the feedback generation is additionally guided by the generated exemplars. We further build two kinds of memory to fully utilize the historical feedback information and support more effective exemplar retrieval. Empirical evaluations show our method surpasses previous state-of-the-arts with less optimization steps, i.e., improving F1 score by 10.1 on LIAR dataset, and reducing half of the optimization steps on ProTeGi.

cs.CL↗

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic benchmarks translate to more realistic conditions? In practice, while the training set is fixed at one point in time, test-time conditions may exhibit distribution shifts relative to the training conditions, such as changes in speaker characteristics, emotional expressiveness, language and acoustic conditions, and the emergence of novel synthesis methods. Although some existing datasets target subsets of these distribution shifts, systematic analysis remains difficult due to inconsistencies between source data and synthesis systems across datasets. This difficulty is further exacerbated by the rapid development of new text-to-speech (TTS) and vocoder systems, which continually expand the diversity of synthetic speech. To enable systematic benchmarking of model performance under distribution shifts, we introduce ShiftySpeech, a large-scale benchmark comprising over 3,000 hours of synthetic speech across 7 source domains, 6 TTS systems, 12 vocoders, and 3 languages. ShiftySpeech is specifically designed to evaluate model generalization under controlled distribution shifts while ensuring broad coverage of modern synthetic speech generation techniques. It fills a key gap in current benchmarks by supporting fine-grained, controlled analysis of generalization robustness. All tested distribution shifts significantly degrade detection performance of state-of-the-art detection approaches based on self-supervised features. Overall, our findings suggest that reliance on synthetic speech detection methods in production environments should be carefully evaluated based on anticipated distribution shifts.

eess.AS↗

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the pre-dominant conversational telephone speech (CTS) domain. We explored architectures based on ResNet with different pooling mechanisms, recently introduced ReDimNet architecture, as well as a system based on the XLS-R model, which represents the family of large pre-trained self-supervised models. In open condition, we train on VoxBlink2 dataset, containing 110 thousand speakers across multiple languages. We observed a good performance and robustness of VoxBlink-trained models, and our experiments show practical recipes for developing state-of-the-art frontends for speaker recognition.

eess.AS↗

Diversity of Superradiant Phase Transitions in the Bose-Fermi System under Tight-Binding Model in the Weak-Coupling Regime

We present a comprehensive analysis of the dynamic diversity associated with superradiant phase transitions within a one-dimensional tight-binding electronic chain that is intrinsically coupled to a single-mode optical cavity. By employing the quantized electromagnetic vector potential through the Peierls substitution, the gauge-invariant coupled Bose-Fermi system facilitates momentum-dependent superradiant transitions and effectively avoids the second-order spurious phase transitions typically observed in Dicke-like models. The quantum phase transitions in this system are characterized by stable dynamics, including the displacement and squeezing of the cavity mode and the redistribution of electronic momentum in the solid chain. Distinct from multimode cavity QED systems with atomic gases, this single-mode optical configuration unveils a range of nonlinear phenomena, including multistability and diversity of spontaneous symmetry breaking. The setup allows for precise manipulation of superradiant phases in the weak coupling regime, effectively mitigating the adverse effects of quantum fluctuation divergences. The diverse attributes of these quantum phase transitions enhance our understanding of tunable quantum solid devices and underscore their potential applications in quantum information processing and metrology.

quant-ph↗

Order within disorder: spectral key generation and distribution in random lasers

In secure communication, highly random entropy sources are essential for information security. Random lasers (RLs), which arise from multiple scattering in disordered structures, are potentially ideal entropy sources. Traditionally, RLs are viewed as disordered and unpredictable. However, in this work, we present novel evidence that orderly patterns exist beneath the seemingly disordered outputs of RLs. Utilizing deep learning techniques, a variety of advanced neural network models are used to analyze the spectral data in multiple dimensions. The results show that the time series of RLs spectra are unpredictable, but spectral wavelength component intensities can be recovered due to inter-modal correlations. This finding not only breaks through the traditional perception that RLs are unpredictable, but also reveals for the first time that RLs have the dual characteristics of both randomness and determinism. Based on this new characteristic, we further expand the application field of RLs and innovatively design a new type of key generation and distribution scheme. In this scheme, the disordered property of RLs is used for key generation to ensure high randomness, while their ordered property is used for key distribution to guarantee accuracy and reliability. The scheme provides a new strategy for secure communication.

physics.optics↗

LipidBERT: A Lipid Language Model Pre-trained on METiS de novo Lipid Library

In this study, we generate and maintain a database of 10 million virtual lipids through METiS's in-house de novo lipid generation algorithms and lipid virtual screening techniques. These virtual lipids serve as a corpus for pre-training, lipid representation learning, and downstream task knowledge transfer, culminating in state-of-the-art LNP property prediction performance. We propose LipidBERT, a BERT-like model pre-trained with the Masked Language Model (MLM) and various secondary tasks. Additionally, we compare the performance of embeddings generated by LipidBERT and PhatGPT, our GPT-like lipid generation model, on downstream tasks. The proposed bilingual LipidBERT model operates in two languages: the language of ionizable lipid pre-training, using in-house dry-lab lipid structures, and the language of LNP fine-tuning, utilizing in-house LNP wet-lab data. This dual capability positions LipidBERT as a key AI-based filter for future screening tasks, including new versions of METiS de novo lipid libraries and, more importantly, candidates for in vivo testing for orgran-targeting LNPs. To the best of our knowledge, this is the first successful demonstration of the capability of a pre-trained language model on virtual lipids and its effectiveness in downstream tasks using web-lab data. This work showcases the clever utilization of METiS's in-house de novo lipid library as well as the power of dry-wet lab integration.

cs.CL↗

GAN-based Generator of Adversarial Attack on Intelligent End-to-End Autoencoder-based Communication System

Deep neural networks have been applied in wireless communications system to intelligently adapt to dynamically changing channel conditions, while the users are still under the threat of the malicious attacks due to the broadcasting property of wireless channels. However, most attack models require the knowledge of the target details, which is difficult to be implemented in real systems. Our objective is to develop an attack model with no requirement for the target information, while enhancing the block error rate. In our design, we propose a novel Generative Adversarial Networks(GANs) based attack architecture, which exploits the property of deep learning models being vulnerable to perturbations induced by dynamically changing channel conditions. In the proposed generator, the attack network is composed of convolution layer, convolution transpose layer and linear layer. Then we present the training strategy and the details of the training algorithm. Subsequently, we propose the validation strategy to evaluate the performance of the generator. Simulations are conducted and the results show that our proposed adversarial attack generator achieve better block error rate attack performance than that of benchmark schemes over Additive White Gaussian Noise (AWGN) channel, Rayleigh channel and High-Speed Railway channel.

cs.IT↗

OSDFace: One-Step Diffusion Model for Face Restoration

Diffusion models have demonstrated impressive performance in face restoration. Yet, their multi-step inference process remains computationally intensive, limiting their applicability in real-world scenarios. Moreover, existing methods often struggle to generate face images that are harmonious, realistic, and consistent with the subject's identity. In this work, we propose OSDFace, a novel one-step diffusion model for face restoration. Specifically, we propose a visual representation embedder (VRE) to better capture prior information and understand the input face. In VRE, low-quality faces are processed by a visual tokenizer and subsequently embedded with a vector-quantized dictionary to generate visual prompts. Additionally, we incorporate a facial identity loss derived from face recognition to further ensure identity consistency. We further employ a generative adversarial network (GAN) as a guidance model to encourage distribution alignment between the restored face and the ground truth. Experimental results demonstrate that OSDFace surpasses current state-of-the-art (SOTA) methods in both visual quality and quantitative metrics, generating high-fidelity, natural face images with high identity consistency. The code and model will be released at https://github.com/jkwang28/OSDFace.

cs.CV↗

Uncertainty Estimation for Trust Attribution to Speed-of-Sound Reconstruction with Variational Networks

Speed-of-sound (SoS) is a biomechanical characteristic of tissue, and its imaging can provide a promising biomarker for diagnosis. Reconstructing SoS images from ultrasound acquisitions can be cast as a limited-angle computed-tomography problem, with Variational Networks being a promising model-based deep learning solution. Some acquired data frames may, however, get corrupted by noise due to, e.g., motion, lack of contact, and acoustic shadows, which in turn negatively affects the resulting SoS reconstructions. We propose to use the uncertainty in SoS reconstructions to attribute trust to each individual acquired frame. Given multiple acquisitions, we then use an uncertainty based automatic selection among these retrospectively, to improve diagnostic decisions. We investigate uncertainty estimation based on Monte Carlo Dropout and Bayesian Variational Inference. We assess our automatic frame selection method for differential diagnosis of breast cancer, distinguishing between benign fibroadenoma and malignant carcinoma. We evaluate 21 lesions classified as BI-RADS~4, which represents suspicious cases for probable malignancy. The most trustworthy frame among four acquisitions of each lesion was identified using uncertainty based criteria. Selecting a frame informed by uncertainty achieved an area under curve of 76% and 80% for Monte Carlo Dropout and Bayesian Variational Inference, respectively, superior to any uncertainty-uninformed baselines with the best one achieving 64%. A novel use of uncertainty estimation is proposed for selecting one of multiple data acquisitions for further processing and decision making.

cs.CV↗

Discovery of a Robust Non-Janus Hybrid MoSH Monolayer as a Two-Gap Superconductor via High-Throughput Computational Screening

The atomic-scale determination of hydrogen positions in MoSH monolayers remains experimentally challenging, and existing studies are confined to Janus-type configurations. Here, we combine high-throughput structural screening with first-principles calculations to predict a novel non-Janus Hybrid 1T$^{'}$-MoSH monolayer, which energetically surpasses all previously reported MoSH phases with a binding energy of -3.02 eV. This structure emerges as a hybrid of MoS$_2$ and MoH$_2$, featuring alternating S and H atoms on both sides of the Mo layer. Comprehensive stability analyses confirm its robustness in energy, mechanics, dynamics, and thermodynamics (stable up to 1600 K). Remarkably, anisotropic Migdal-Eliashberg theory predicts Hybrid 1T$^{'}$-MoSH as a two-gap superconductor with a critical temperature T$_c$ of 16.34 K, driven by strong electron-phonon coupling ($λ$$=$1.39). Substituting Mo with Hf, Ta, or Ti drastically suppresses T$_c$ $\sim$ (0.53-2.42 K), highlighting Mo$^{'}$s unique role in enhancing superconductivity. Our work not only expands the family of 2D transition metal chalcogenides but also proposes a promising candidate for quantum technologies, bridging theoretical design to functional material discovery.

cond-mat.mtrl-sci↗

Geometry of sets of Bargmann invariants

Certain unitary-invariants, known as Bargmann invariants or multivariate traces of quantum states, have recently gained attention due to their applications in quantum information theory. However, determining the boundaries of sets of Bargmann invariants remains a theoretical challenge. In this study, we address the problem by developing a unified, dimension-independent formulation that characterizes the sets of the 3rd and 4th Bargmann invariants.In particular, our result for the set of 4th Bargmann invariants confirms the conjecture given by Fernandes \emph{et al.} [Phys.Rev.Lett.\href{https://doi.org/10.1103/PhysRevLett.133.190201}{\textbf{133}, 190201 (2024)}]. Based on the obtained results, we conjecture that the unified, dimension-independent formulation of the boundaries for sets of 3rd-order and 4th-order Bargmann invariants may extend to the general case of the $n$th-order Bargmann invariants. These results deepen our understanding of the fundamental physical limits within quantum mechanics and pave the way for novel applications of Bargmann invariants in quantum information processing and related fields.

quant-ph↗

Stellar parameters, extinction, and distances for stars in SMSS DR2 by SPar method

The availability of large datasets containing stellar parameters, distances, and extinctions for stars in the Milky Way, particularly within the Galactic disk, is essential for advancing our understanding of the Galaxy's stellar populations, structure, kinematics, and chemical evolution. In this study, we present a catalog of stellar parameters, including effective temperature (\teff), metallicity (\feh), absolute magnitudes ($M_{G}$), distances ($d$), and reddening values (\ebr), for a sample of 141 million stars from the SkyMapper Southern Survey (SMSS). These parameters are derived using the SPar algorithm, which employs a fitting procedure to match multi-band photometric observations and estimate the stellar properties (\teff, \feh, $M_G$, $d$, and \ebr) on an individual star basis, following the methodology outlined in our previous work. This study successfully determines stellar parameters and extinction values simultaneously for stars located in high and low Galactic latitudes. The resulting stellar parameters and extinction values are in good agreement with results from other studies, demonstrating the robustness of our method. We derive a temperature dispersion of 195\,K and a metallicity dispersion of 0.31\,dex when comparing our results with spectroscopic data. The catalog produced in this work provides a valuable resource for future research on the Galactic metallicity distribution function, the structure of the Galaxy, three-dimensional extinction mapping, and other related astrophysical topics.

astro-ph.GA↗

LITE: LLM-Impelled efficient Taxonomy Evaluation

This paper presents LITE, an LLM-based evaluation method designed for efficient and flexible assessment of taxonomy quality. To address challenges in large-scale taxonomy evaluation, such as efficiency, fairness, and consistency, LITE adopts a top-down hierarchical evaluation strategy, breaking down the taxonomy into manageable substructures and ensuring result reliability through cross-validation and standardized input formats. LITE also introduces a penalty mechanism to handle extreme cases and provides both quantitative performance analysis and qualitative insights by integrating evaluation metrics closely aligned with task objectives. Experimental results show that LITE demonstrates high reliability in complex evaluation tasks, effectively identifying semantic errors, logical contradictions, and structural flaws in taxonomies, while offering directions for improvement. Code is available at https://github.com/Zhang-l-i-n/TAXONOMY_DETECT .

cs.CL↗

RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model

As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitations such as high resource costs and information loss. We propose the Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model (RECKON), which directly uses reference data to evaluate models. RECKON organizes unstructured data into manageable units and generates targeted questions for each cluster, improving evaluation accuracy and efficiency. Experimental results show that RECKON reduces resource consumption by 56.5% compared to traditional methods while achieving over 97% accuracy across various domains, including world knowledge, code, legal, and biomedical datasets. Code is available at https://github.com/MikeGu721/reckon

cs.CL↗