Searcharxiv⌕ Search

arXiv subjects

Yang He

Publications and source records attributed to Yang He.

At least 55 records · Page 3Linked to original sources

Dataset Color Quantization: A Training-Oriented Framework for Dataset-Level Compression

Large-scale image datasets are fundamental to deep learning, but their high storage demands pose challenges for deployment in resource-constrained environments. While existing approaches reduce dataset size by discarding samples, they often ignore the significant redundancy within each image -- particularly in the color space. To address this, we propose Dataset Color Quantization (DCQ), a unified framework that compresses visual datasets by reducing color-space redundancy while preserving information crucial for model training. DCQ achieves this by enforcing consistent palette representations across similar images, selectively retaining semantically important colors guided by model perception, and maintaining structural details necessary for effective feature learning. Extensive experiments across CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-1K show that DCQ significantly improves training performance under aggressive compression, offering a scalable and robust solution for dataset-level storage reduction.

cs.CV↗

Physics-Informed Regression Modelling for Vertical Facade Surface Temperature: A Tropical Case Study on Solar-reflective Material

Urban heat islands (UHIs) pose a critical challenge in densely populated cities and tropical climates where large amounts of energy are used to meet the cooling demand. To address this, Building and Construction Authority (BCA) of Singapore provides incentives for passive cooling such as using of solar-reflective material in its Green Mark guidelines. Thus, understanding about its real-world effectiveness in tropical urban environments is required. This study evaluated the effectiveness of solar-reflective cool paint using a hybrid modelling framework combining a transient physical model and data driven model through field measurements. Several machine learning algorithms were compared including multiple-linear regression (MLR), random forest regressor (RF), AdaBoost regressor (AB), extreme gradient boosting regressor (XGB), and TabPFN regressor (TPR). The results indicated that the transient physical model overestimated facade temperatures in the lower temperature ranges. The physics-informed MLR achieved best performance with improved accuracy for pre-cool paint (R2=0.96, RMSE=0.83C) and post-cool paint (R2=0.95, RMSE=0.65C) scenarios, reducing RMSE by 26% and 44%, respectively. The hybrid model also effectively predicted hourly heat fluxes revealing substantial reductions in surface temperature and heat storage with increasing albedo. The maximum net heat flux q_net was reduced by about 30-65 W/m2 in the post-cool paint stage (albedo = 0.73) compared to the pre-cool paint stage (albedo = 0.31). As albedo increases from 0.1 to 0.9, the sensitivity analysis predicts that the maximum daytime surface temperature will decrease by about 11C and the peak heat release of the net heat flux will decrease significantly from about 161 W/m2 to 27 W/m2.

physics.app-ph↗

Chip-scale ultrafast soliton laser

Femtosecond laser, owing to their ultrafast time scales and broad frequency bandwidths, have substantially changed fundamental science over the past decades, from chemistry and bio-imaging to quantum physics. Critically, many emerging industrial-scale photonic technologies -- such as optical interconnects, AI accelerators, quantum computing, and LiDAR -- also stand to benefit from their massive frequency parallelism. However, achieving a femtosecond-scale laser on-chip, constrained by size and system power input, has remained a long-standing challenge. Here, we demonstrate the first on-chip femtosecond laser, enabled by a new mechanism -- photorefraction-assisted soliton (PAS) mode-locking. Operating from a simple, low-voltage electrical supply, the laser provides deterministic, turn-key generation of sub-90-fs solitons. Furthermore, it provides electronic reconfigurability of its pulse properties and features an exceptional optical coherence with a 53 Hz intrinsic comb linewidth. This demonstration removes a key barrier to the full integration of chip-scale photonic systems for next-generation sensing, communication, metrology, and computing.

physics.optics↗

SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs

Multimodal Large Language Models (MLLMs) typically process a large number of visual tokens, leading to considerable computational overhead, even though many of these tokens are redundant. Existing visual token pruning methods primarily focus on selecting the most salient tokens based on attention scores, resulting in the semantic incompleteness of the selected tokens. In this paper, we propose a novel visual token pruning strategy, called \textbf{S}aliency-\textbf{C}overage \textbf{O}riented token \textbf{P}runing for \textbf{E}fficient MLLMs (SCOPE), to jointly model both the saliency and coverage of the selected visual tokens to better preserve semantic completeness. Specifically, we introduce a set-coverage for a given set of selected tokens, computed based on the token relationships. We then define a token-coverage gain for each unselected token, quantifying how much additional coverage would be obtained by including it. By integrating the saliency score into the token-coverage gain, we propose our SCOPE score and iteratively select the token with the highest SCOPE score. We conduct extensive experiments on multiple vision-language understanding benchmarks using the LLaVA-1.5 and LLaVA-Next models. Experimental results demonstrate that our method consistently outperforms prior approaches. Our code is available at \href{https://github.com/kinredon/SCOPE}{https://github.com/kinredon/SCOPE}.

cs.CV↗

Chiral supersolid and dissipative time crystal in Rydberg-dressed Bose-Einstein condensates with Raman-induced spin-orbit coupling

Spin-orbit coupling (SOC) is one of the crucial factors that affect the chiral symmetry of matter by causing the spatial symmetry breaking of the system. We find that Raman-induced SOC can induce a chiral supersolid phase with a helical antiskyrmion lattice in balanced Rydberg-dressed two-component Bose-Einstein condensates (BECs) in a harmonic trap by modulating the Raman coupling strength. This is in stark contrast to the mirror symmetric supersolid phase containing skyrmion-antiskyrmion lattice pair for the case of Rashba SOC. Two ground-state phase diagrams are presented as a function of the Rydberg interaction and the Raman-induced SOC. It is shown that the interplay among Raman-induced SOC, Rydberg interactions, and nonlinear contact interactions favors rich ground-state structures, including half-quantum vortex phase, stripe supersolid phase, toroidal stripe phase with a central Anderson-Toulouse coreless vortex, checkerboard supersolid phase, mirror symmetric supersolid phase, chiral supersolid phase and standing-wave supersolid phase. In addition, the effects of rotation and in-plane quadrupole magnetic field on the ground state of the system are analyzed. In these two cases, the chiral supersolid phase is broken and the ground state tends to form a miscible phase. Furthermore, we demonstrate that when the initial state is a chiral supersolid phase the rotating harmonic trapped system sustains dissipative continuous time crystal by studying the rotational dynamic behaviors of the system.

cond-mat.quant-gas↗

Attosecond electron bunch generation by an intense high-order harmonic pulse interacting with a thin target

Laser-accelerated electron bunches and the secondary radiation sources they produce exhibit unique temporal resolution for probing ultrafast physical processes due to their ultrashort pulse duration. The inherently short temporal profile of these pulses leads to extremely high peak bunch currents, thereby enabling a wide range of practical applications. In this study, we propose an innovative method for generating such bunch by utilizing high-harmonics generated through laser-plasma interaction as the driving pulse, which subsequently interacts with a thin target to produce an attosecond electron bunch. Using this method, we successfully generated an electron bunch characterized by excellent collimation and an ultra-short duration of approximately 100 attoseconds, representing a substantial reduction in bunch duration. The total bunch charge achieved was 0.38 nC, with an emittance of $4.5 \times 10^{-3} \, \text{mm} \cdot \text{mrad}$ and a divergence angle of approximately $10^\circ$. Moreover, by systematically analyzing the effects of laser intensity and target positioning, we determined an optimized set of simulation parameters. This research establishes a robust foundation for the generation of ultrashort electron bunches and opens new prospects for their application in advanced high-energy and attosecond physics experiments.

physics.acc-ph↗

Artificial ferroelectric-like hysteresis in antiferroelectrics with non-uniform disorder

Antiferroelectrics exhibit unique double-hysteresis polarization loops, which have garnered significant attention due to their potential applications such as energy storage, electromechanical transduction, as well as synapse devices. However, numerous antiferroelectric materials have been reported to display signs of hysteresis loops resembling those of ferroelectric materials, and a comprehensive understanding remains elusive. In this work, we provide a phenomenological model that reproduces such widely observed artificial ferroelectric hysteresis with a superposition of numerous disordered antiferroelectric loops that have varying antiferroelectric-to-ferroelectric transition fields, particularly when these field ranges intersect. Experimentally, we realized such artificial ferroelectric-like hysteresis loops in the prototypical antiferroelectric PbZrO$_3$ and PbHfO$_3$ thin films, by introducing non-uniform local disorder (e.g., defects) via fine-tuning of the film growth conditions. These ferroelectric-like states are capable of persisting for several hours prior to transitioning back into the thermodynamically stable antiferroelectric ground state. Those results provide insights into the fundamental impact of disorder on the AFE properties and new possibilities of disorder-tailored functions.

cond-mat.mtrl-sci↗

Synthesizing Optimal Object Selection Predicates for Image Editing using Lattices

Image editing is a common task across a wide range of domains, from personal use to professional applications. Despite advances in computer vision, current tools still demand significant manual effort for editing tasks that require repetitive operations on images with many objects. In this paper, we present a novel approach to automating the image editing process using program synthesis. We propose a new algorithm based on lattice structures to automatically synthesize object selection predicates for image editing from positive and negative examples. By leveraging the algebraic properties of lattices, our algorithm efficiently synthesizes an optimal object selection predicate among multiple correct solutions. We have implemented our technique and evaluated it on 100 tasks over 20 images. The evaluation result demonstrates our tool is effective and efficient, which outperforms state-of-the-art synthesizers and LLM-based approaches.

cs.PL↗

Etching-free dual-lift-off for direct patterning of epitaxial oxide thin films

Although monocrystalline oxide films offer broad functional capabilities, their practical use is hampered by challenges in patterning. Traditional patterning relies on etching, which can be costly and prone to issues like film or substrate damage, under-etching, over-etching, and lateral etching. In this study, we introduce a dual-lift-off method for direct patterning of oxide films, circumventing the etching process and associated issues. Our method involves an initial lift-off of amorphous Sr$_3$Al$_2$O$_6$ or Sr$_4$Al$_2$O$_7$ ($a$SAO) through stripping the photoresist, followed by a subsequent lift-off of the functional oxide thin films by dissolving the $a$SAO layer. $a$SAO functions as a ``high-temperature photoresist", making it compatible with the high-temperature growth of monocrystalline oxides. Using this method, patterned ferromagnetic La$_{0.67}$Sr$_{0.33}$MnO$_{3}$ and ferroelectric BiFeO$_3$ were fabricated, accurately mirroring the shape of the photoresist. Our study presents a straightforward, flexible, precise, environmentally friendly, and cost-effective method for patterning high-quality oxide thin films.

cond-mat.mtrl-sci↗

Three-dimensional hyperspectral imaging with optical microcombs

Optical frequency combs have revolutionised time and frequency metrology [1, 2]. The advent of microresonator-based frequency combs ('microcombs' [3-5]) is set to lead to the miniaturisation of devices that are ideally suited to a wide range of applications, including microwave generation [6, 7], ranging [8-10], the precise calibration of astronomical spectrographs [11], neuromorphic computing [12, 13], high-bandwidth data communications[14], and quantum-optics [15, 16] platforms. Here, we introduce a new microcomb application for three-dimensional imaging. Our method can simultaneously determine the chemical identity and full three-dimensional geometry, including size, shape, depth, and spatial coordinates, of particulate matter ranging from micrometres to millimetres in size across nearly $10^5$ distinct image pixels. We demonstrate our technique using millimetre-sized plastic specimens (i.e. microplastics measuring less than 5 mm). We combine amplitude and phase analysis and achieve a throughput exceeding $1.2~10^6$ pixels per second with micrometre-scale precision. Our method leverages the defining feature of microcombs - their large line spacing - to enable precise spectral diagnostics using microcombs with a repetition frequency of 1 THz. Our results suggest scalable operation over several million pixels and nanometre-scale axial resolution. Coupled with its high-speed, label-free and multiplexed capabilities, our approach provides a promising basis for environmental sensing, particularly for the real-time detection and characterisation of microplastic pollutants in aquatic ecosystems [17].

physics.optics↗

TCFNet: Bidirectional face-bone transformation via a Transformer-based coarse-to-fine point movement network

Computer-aided surgical simulation is a critical component of orthognathic surgical planning, where accurately simulating face-bone shape transformations is significant. The traditional biomechanical simulation methods are limited by their computational time consumption levels, labor-intensive data processing strategies and low accuracy. Recently, deep learning-based simulation methods have been proposed to view this problem as a point-to-point transformation between skeletal and facial point clouds. However, these approaches cannot process large-scale points, have limited receptive fields that lead to noisy points, and employ complex preprocessing and postprocessing operations based on registration. These shortcomings limit the performance and widespread applicability of such methods. Therefore, we propose a Transformer-based coarse-to-fine point movement network (TCFNet) to learn unique, complicated correspondences at the patch and point levels for dense face-bone point cloud transformations. This end-to-end framework adopts a Transformer-based network and a local information aggregation network (LIA-Net) in the first and second stages, respectively, which reinforce each other to generate precise point movement paths. LIA-Net can effectively compensate for the neighborhood precision loss of the Transformer-based network by modeling local geometric structures (edges, orientations and relative position features). The previous global features are employed to guide the local displacement using a gated recurrent unit. Inspired by deformable medical image registration, we propose an auxiliary loss that can utilize expert knowledge for reconstructing critical organs.Compared with the existing state-of-the-art (SOTA) methods on gathered datasets, TCFNet achieves outstanding evaluation metrics and visualization results. The code is available at https://github.com/Runshi-Zhang/TCFNet.

cs.CV↗

MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents

Modern task-oriented dialogue (TOD) systems increasingly rely on large language model (LLM) agents, leveraging Retrieval-Augmented Generation (RAG) and long-context capabilities for long-term memory utilization. However, these methods are primarily based on semantic similarity, overlooking task intent and reducing task coherence in multi-session dialogues. To address this challenge, we introduce MemGuide, a two-stage framework for intent-driven memory selection. (1) Intent-Aligned Retrieval matches the current dialogue context with stored intent descriptions in the memory bank, retrieving QA-formatted memory units that share the same goal. (2) Missing-Slot Guided Filtering employs a chain-of-thought slot reasoner to enumerate unfilled slots, then uses a fine-tuned LLaMA-8B filter to re-rank the retrieved units by marginal slot-completion gain. The resulting memory units inform a proactive strategy that minimizes conversational turns by directly addressing information gaps. Based on this framework, we introduce the MS-TOD, the first multi-session TOD benchmark comprising 132 diverse personas, 956 task goals, and annotated intent-aligned memory targets, supporting efficient multi-session task completion. Evaluations on MS-TOD show that MemGuide raises the task success rate by 11% (88% -> 99%) and reduces dialogue length by 2.84 turns in multi-session settings, while maintaining parity with single-session benchmarks.

cs.CL↗

Microwave Engineering of Tunable Spin Interactions with Superconducting Qubits

Quantum simulation has emerged as a powerful framework for investigating complex many - body phenomena. A key requirement for emulating these dynamics is the realization of fully controllable quantum systems enabling various spin interactions. Yet, quantum simulators remain constrained in the types of attainable interactions. Here we demonstrate experimental realization of multiple microwave - engineered spin interactions in superconducting quantum circuits. By precisely controlling the native XY interaction and microwave drives, we achieve tunable spin Hamiltonians including: (i) XYZ spin models with continuously adjustable parameters, (ii) transverse - field Ising systems, and (iii) Dzyaloshinskii - Moriya interacting systems. Our work expands the toolbox for analogue - digital quantum simulation, enabling exploration of a wide range of exotic quantum spin models.

quant-ph↗

Many-body delocalization with a two-dimensional 70-qubit superconducting quantum simulator

Quantum many-body systems with sufficiently strong disorder can exhibit a non-equilibrium phenomenon, known as the many-body localization (MBL), which is distinct from conventional thermalization. While the MBL regime has been extensively studied in one dimension, its existence in higher dimensions remains elusive, challenged by the avalanche instability. Here, using a 70-qubit two-dimensional (2D) superconducting quantum simulator, we experimentally explore the robustness of the MBL regime in controlled finite-size 2D systems. We observe that the decay of imbalance becomes more pronounced with increasing system sizes, scaling up from 21, 42 to 70 qubits, with a relatively large disorder strength, and for the first time, provide an evidence for the many-body delocalization in 2D disordered systems. Our experimental results are consistent with the avalanche theory that predicts the instability of MBL regime beyond one spatial dimension. This work establishes a scalable platform for probing high-dimensional non-equilibrium phases of matter and their finite-size effects using superconducting quantum circuits.

quant-ph↗

Geography of Landau-Ginzburg models and threefold syzygies

We study the behavior of toric Landau-Ginzburg models under extremal contraction and minimal model program. We also establish a relation between the moduli space of toric Landau-Ginzburg models and the geography of central models. We conjecture that there is a correspondence between extremal contractions and minimal model program on Fano varieties, and degenerations of their associated toric Landau-Ginzburg models written explicitly. We prove the conjectures for smooth toric varieties, as well as general smooth Fano varieties in dimensions 2 and 3. As an application, we compute the elementary syzygies for smooth Fano threefolds.

math.AG↗

Self-Route: Automatic Mode Switching via Capability Estimation for Efficient Reasoning

While reasoning-augmented large language models (RLLMs) significantly enhance complex task performance through extended reasoning chains, they inevitably introduce substantial unnecessary token consumption, particularly for simpler problems where Short Chain-of-Thought (Short CoT) suffices. This overthinking phenomenon leads to inefficient resource usage without proportional accuracy gains. To address this issue, we propose Self-Route, a dynamic reasoning framework that automatically selects between general and reasoning modes based on model capability estimation. Our approach introduces a lightweight pre-inference stage to extract capability-aware embeddings from hidden layer representations, enabling real-time evaluation of the model's ability to solve problems. We further construct Gradient-10K, a model difficulty estimation-based dataset with dense complexity sampling, to train the router for precise capability boundary detection. Extensive experiments demonstrate that Self-Route achieves comparable accuracy to reasoning models while reducing token consumption by 30-55\% across diverse benchmarks. The proposed framework demonstrates consistent effectiveness across models with different parameter scales and reasoning paradigms, highlighting its general applicability and practical value.

cs.CL↗

Graphiti: Bridging Graph and Relational Database Queries

This paper presents an automated reasoning technique for checking equivalence between graph database queries written in Cypher and relational queries in SQL. To formalize a suitable notion of equivalence in this setting, we introduce the concept of database transformers, which transform database instances between graph and relational models. We then propose a novel verification methodology that checks equivalence modulo a given transformer by reducing the original problem to verifying equivalence between a pair of SQL queries. This reduction is achieved by embedding a subset of Cypher into SQL through syntax-directed translation, allowing us to leverage existing research on automated reasoning for SQL while obviating the need for reasoning simultaneously over two different data models. We have implemented our approach in a tool called Graphiti and used it to check equivalence between graph and relational queries. Our experiments demonstrate that Graphiti is useful both for verification and refutation and that it can uncover subtle bugs, including those found in Cypher tutorials and academic papers.

cs.PL↗