SearcharxivSearch

arXiv subjects

Zhenyang Xiao

Publications and source records attributed to Zhenyang Xiao.

10 recordsLinked to original sources

Self-Evolving Critique Abilities in Large Language Models

Despite their remarkable performance, Large Language Models (LLMs) face a critical challenge: providing feedback for tasks where human evaluation is difficult or where LLMs potentially outperform humans. In such scenarios, leveraging the critique ability of LLMs themselves - identifying and correcting flaws - shows considerable promise. This paper explores enhancing critique abilities of LLMs, noting that current approaches rely on human annotations or more powerful models, leaving the challenge of improving critique abilities without external supervision unresolved. We introduce SCRIT (Self-evolving CRITic), a framework that trains LLMs with self-generated data to evolve their critique abilities. To address the low quality of naively generated data, we propose a contrastive-critic approach that uses reference solutions during data synthesis to enhance the model's understanding of key concepts, and incorporates a self-validation scheme to ensure data quality. The final trained model operates without any reference solutions at inference time. Implemented with Qwen2.5-72B-Instruct, a leading LLM, SCRIT demonstrates consistent improvements across a wide range of benchmarks spanning both mathematical and scientific reasoning: achieving a 10.0\% relative gain in critique-correction accuracy and a 19.0\% relative improvement in error identification F1-score. Our analysis reveals that SCRIT's performance scales positively with data and model size and enables continuous improvement through multi-round iterations.

cs.CL

Rubber band filters: optimal padding without edge artifacts

Bandpass filtering techniques are widely used in spectroscopy. However, conventional symmetric-padding filtering methods introduce boundary artifacts that distort the signal at the edges. We present a rubber band filter: a robust method for achieving band-limited filtering without these detrimental edge artifacts. The technique applies an optimal padding scheme during the filtering process, thereby overcoming longstanding challenges in achieving artifact-free filtering. Importantly, it is iterative and requires only a few extra Fourier transforms over conventional approaches. We demonstrate its superiority and versatility by applying it to three spectroscopic examples -- time-domain spectroscopy, Fourier-transform spectroscopy, and dual-comb spectroscopy.

physics.optics

Liquid combs: broadband light with equidistance and without stability

Broadband light sources with well-defined spectral structures are vital for science and technology. However, the evenly spaced lines of frequency combs represent only a small subset of all possible structured white-light sources. We demonstrate liquid combs: optical states that preserve spectral equidistance but lack temporal stability. By engineering the gain and dispersion of semiconductor laser cavities, we produce light that possesses rapid phase fluctuations but maintains relative phase differences between modes that vary identically. We show experimentally that this phenomenon occurs in multiple laser platforms -- across multiple octaves -- through the creation of a metrological technique that determines the phase differences. We also show theoretically that this is a general phenomenon that can be described using a mean-field theory. These liquid combs are attractive for many applications due to having wider bandwidths than frequency combs, and more generally, they represent the long-sought realization of structured white-light sources that are not combs.

physics.optics

RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques

Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and suggesting improvements. However, evaluating the critique capabilities of LLMs presents a significant challenge due to the open-ended nature of the task. In this work, we introduce a new benchmark designed to assess the critique capabilities of LLMs. Unlike existing benchmarks, which typically function in an open-loop fashion, our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques. Moreover, the benchmark incorporates features such as self-critique, cross-critique, and iterative critique, which are crucial for distinguishing the abilities of advanced reasoning models from more classical ones. We implement this benchmark using eight challenging reasoning tasks. We have several interesting findings. First, despite demonstrating comparable performance in direct chain-of-thought generation, classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios. Second, in self-critique and iterative critique settings, classical LLMs may even underperform relative to their baseline capabilities. We hope that this benchmark will serve as a valuable resource to guide future advancements. The code and data are available at \url{https://github.com/tangzhy/RealCritic}.

cs.CL

Fundamental scaling limits and bandwidth shaping of frequency-modulated combs

Frequency-modulated (FM) combs based on active cavities like quantum cascade lasers have recently emerged as promising light sources in many spectral regions. Unlike passive modelocking, which uses amplitude modulation to generate amplitude modulation, FM combs use phase modulation to generate phase modulation. They can therefore be regarded as a phase-domain version of passive modelocking. However, while the ultimate scaling laws of passive modelocking have long been known -- Haus showed in 1975 that pulses have a bandwidth proportional to effective gain bandwidth -- the limits of FM combs have been much less clear. Here, we show that FM combs are governed by the same fundamental limits, producing combs whose bandwidths are linear in the effective gain bandwidth. Not only do we show theoretically that the diffusive effect of gain curvature limits comb bandwidth, we also show experimentally how this limit can be increased. By adding carefully designed resonant-loss structures that are evanescently coupled to the cavity of a terahertz laser, we reduce the curvature and increase the effective gain bandwidth of the laser, demonstrating bandwidth enhancement. Our results give a new degree of freedom for the creation of active chip-scale combs and can be applied to a wide array of cavity geometries.

physics.optics

Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models

Self-correction has achieved impressive results in enhancing the style and security of the generated output from large language models (LLMs). However, recent studies suggest that self-correction might be limited or even counterproductive in reasoning tasks due to LLMs' difficulties in identifying logical mistakes. In this paper, we aim to enhance the self-checking capabilities of LLMs by constructing training data for checking tasks. Specifically, we apply the Chain of Thought (CoT) methodology to self-checking tasks, utilizing fine-grained step-level analyses and explanations to assess the correctness of reasoning paths. We propose a specialized checking format called "Step CoT Check". Following this format, we construct a checking-correction dataset that includes detailed step-by-step analysis and checking. Then we fine-tune LLMs to enhance their error detection and correction abilities. Our experiments demonstrate that fine-tuning with the "Step CoT Check" format significantly improves the self-checking and self-correction abilities of LLMs across multiple benchmarks. This approach outperforms other formats, especially in locating the incorrect position, with greater benefits observed in larger models. For reproducibility, all the datasets and code are provided in https://github.com/bammt/Learn-to-check.

cs.CL

LoMA: Lossless Compressed Memory Attention

Large Language Models (LLMs) face limitations due to the high demand on GPU memory and computational resources when handling long contexts. While sparsify the Key-Value (KV) cache of transformer model is a typical strategy to alleviate resource usage, it unavoidably results in the loss of information. We introduce Lossless Compressed Memory Attention (LoMA), a novel approach that enables lossless compression of the KV cache, thereby reducing the memory and computational demands during autoregressive generation. LoMA incorporates a specialized training or fine-tuning precedure alongside an autoregressive generation algorithm optimized for the compressed context. Our method compresses the KV cache after every $tc$ generated tokens with a compression ratio of $c$ and a target compressed length $t$, and this process occurs within a single inference pass without dependency on auxiliary models. We engineered an efficient training scheme involving specific inputs, attention masks, and position identifiers to instill this compression capability. Experimental validation has demonstrated that LoMA significantly reducing computational consumption and memory usage through achieving lossless KV cache compression.

cs.LG

Frequency combs in optically-injected terahertz ring quantum cascade lasers

Quantum cascade lasers (QCLs) have emerged as promising candidates for generating chip-scale frequency combs in mid-infrared and terahertz wavelengths. In this work, we demonstrate frequency comb formation in ring terahertz QCLs using the injection of light from a distributed feedback (DFB) laser. The DFB design frequency is chosen to match the modes of the ring cavity (near 3.3 THz), and light from the DFB is injected into the ring QCL via a bus waveguide. By controlling the power and frequency of the optical injection, we show experimentally and theoretically that combs can be selectively formed and controlled in the ring cavity. The potential for soliton generation and efficient power extraction through the bus waveguide presents exciting opportunities for integrating this technology in compact comb designs.

physics.optics

Mao-Zedong At SemEval-2023 Task 4: Label Represention Multi-Head Attention Model With Contrastive Learning-Enhanced Nearest Neighbor Mechanism For Multi-Label Text Classification

The study of human values is essential in both practical and theoretical domains. With the development of computational linguistics, the creation of large-scale datasets has made it possible to automatically recognize human values accurately. SemEval 2023 Task 4\cite{kiesel:2023} provides a set of arguments and 20 types of human values that are implicitly expressed in each argument. In this paper, we present our team's solution. We use the Roberta\cite{liu_roberta_2019} model to obtain the word vector encoding of the document and propose a multi-head attention mechanism to establish connections between specific labels and semantic components. Furthermore, we use a contrastive learning-enhanced K-nearest neighbor mechanism\cite{su_contrastive_2022} to leverage existing instance information for prediction. Our approach achieved an F1 score of 0.533 on the test set and ranked fourth on the leaderboard.

cs.CL

Optical-pump terahertz-probe spectroscopy of the topological crystalline insulator Pb1-xSnxSe through the topological phase transition

Topological crystalline insulators -- topological insulators whose properties are guaranteed by crystalline symmetry -- can potentially provide a promising platform for terahertz optoelectronic devices, as their properties can be tuned on demand when layered in heterostructures. We perform the first optical-pump terahertz-probe spectroscopy of topological crystalline insulators, using them to study the dynamics of Pb1-xSnxSe as a function of temperature. At low temperatures, excitation of Dirac fermions leads to an increase in terahertz transmission; from this negative photoconductivity, the intrasubband relaxation rate of 6 ps is extracted. At high temperatures where only massive fermions exist, the free-carrier losses induced by the pump reduce the terahertz transmission for the duration of the 27 ps interband lifetime. Both effects are present at temperatures near the topological-to-trivial transition. Our experimental observations provide critical details for potential applications of Pb1-xSnxSe and provide a direct measurement of the topological character of Pb1-xSnxSe heterostructures.

physics.optics