Searcharxiv⌕ Search

arXiv subjects

Hong Yang

Publications and source records attributed to Hong Yang.

At least 37 records · Page 2Linked to original sources

Unidirectional reflection lasing based on destructive interference and Bragg scattering modulation in defective atomic lattice

The novel and ingenious scheme we propose for achieving unidirectional reflection lasing (URL) involves integrating a one-dimensional (1D) defective atomic lattice with a coherent gain atomic system. Its physical essence lies in the fact that the right-side reflectivity is drastically reduced due to the destructive interference between primary and secondary reflections, whereas on the left-side primary reflection is effectively suppressed and the secondary reflection is efficiently enhanced, ultimately reaching the lasing threshold. Through numerical results and further analyses, we have elucidated how to precisely tailor the lattice parameters and coupling fields to control destructive interference point (DIP), thereby realizing URL and enabling its active modulation. Our scheme is experimentally feasible and not only effectively circumvents the stringent conditions faced in directly realizing URL, providing a new pathway, but also beneficial for integrating active photonic devices into compact quantum networks and may improve the efficiency of optical information transmission.

physics.optics↗

VALLR-Pin: Uncertainty-Factorized Visual Speech Recognition for Mandarin with Pinyin Guidance

Visual speech recognition (VSR) aims to transcribe spoken content from silent lip-motion videos and is particularly challenging in Mandarin due to severe viseme ambiguity and pervasive homophones. We propose VALLR-Pin, a two-stage Mandarin VSR framework that extends the VALLR architecture by explicitly incorporating Pinyin as an intermediate representation. In the first stage, a shared visual encoder feeds dual decoders that jointly predict Mandarin characters and their corresponding Pinyin sequences, encouraging more robust visual-linguistic representations. In the second stage, an LLM-based refinement module takes the predicted Pinyin sequence together with an N-best list of character hypotheses to resolve homophone-induced ambiguities. To further adapt the LLM to visual recognition errors, we fine-tune it on synthetic instruction data constructed from model-generated Pinyin-text pairs, enabling error-aware correction. Experiments on public Mandarin VSR benchmarks demonstrate that VALLR-Pin consistently improves transcription accuracy under multi-speaker conditions, highlighting the effectiveness of combining phonetic guidance with lightweight LLM refinement.

cs.CV↗

Unidirectional spectral singularity lasing in a defective atomic lattice

We propose an efficient scheme for achieving mode-tunable unidirectional reflection lasing (URL) by establishing a coherent gain atomic system to amplify the probe field and ingeniously designing the one-dimensional (1D) defective atomic lattice. This lattice not only replaces the resonant cavity to provide a distributed feedback mechanism but also breaks the spatial symmetry of the probe susceptibility. Correspondingly, the URL can be characterized by a non-Hermitian degenerate spectral singularity (NHDSS), where the two eigenvalues of the inverse scattering matrix are engineered to satisfy $λ_{S^{-1}}^{+}\simeq λ_{S^{-1}}^{-}\rightarrow 0$. This intriguing NHDSS depends on the probe susceptibility and the Bragg condition, both of which can be modulated by adjusting the external optical field and lattice structure, rendering the scheme experimentally feasible. Our approach achieves both nonreciprocity and lasing oscillation in a single system, significantly enhancing the efficiency of optical information transmission and facilitating the integration of active photonic devices into compact quantum networks.

physics.optics↗

Bizard: A Community-Driven Platform for Accelerating and Enhancing Biomedical Data Visualization

Biomedical research increasingly relies on heterogeneous, high-dimensional datasets, yet effective visualization remains hindered by fragmented code resources, steep programming barriers, and limited domain-specific guidance. Bizard is an open-source visualization code repository engineered to streamline data analysis in biomedical research. It aggregates a diverse array of executable visualization scripts, empowering researchers to select and tailor optimal graphical methods for their specific investigative demands. The platform features an intuitive interface equipped with sophisticated browsing and filtering capabilities, exhaustive tutorials, and interactive discussion forums that foster knowledge dissemination. Through its community-driven paradigm, Bizard promotes continual refinement and functional expansion, establishing itself as an essential resource for elevating biomedical data visualization and analytical standards. By harnessing Bizard's infrastructure, researchers can augment their visualization proficiency, propel methodological progress, and enhance interpretive rigor, ultimately accelerating precision medicine and personalized therapeutics. Bizard is freely accessible at https://openbiox.github.io/Bizard/.

q-bio.GN↗

Can We Ignore Labels In Out of Distribution Detection?

Out-of-distribution (OOD) detection methods have recently become more prominent, serving as a core element in safety-critical autonomous systems. One major purpose of OOD detection is to reject invalid inputs that could lead to unpredictable errors and compromise safety. Due to the cost of labeled data, recent works have investigated the feasibility of self-supervised learning (SSL) OOD detection, unlabeled OOD detection, and zero shot OOD detection. In this work, we identify a set of conditions for a theoretical guarantee of failure in unlabeled OOD detection algorithms from an information-theoretic perspective. These conditions are present in all OOD tasks dealing with real-world data: I) we provide theoretical proof of unlabeled OOD detection failure when there exists zero mutual information between the learning objective and the in-distribution labels, a.k.a. 'label blindness', II) we define a new OOD task - Adjacent OOD detection - that tests for label blindness and accounts for a previously ignored safety gap in all OOD detection benchmarks, and III) we perform experiments demonstrating that existing unlabeled OOD methods fail under conditions suggested by our label blindness theory and analyze the implications for future research in unlabeled OOD methods.

cs.LG↗

Embodied Intelligence in Disassembly: Multimodal Perception Cross-validation and Continual Learning in Neuro-Symbolic TAMP

With the rapid development of the new energy vehicle industry, the efficient disassembly and recycling of power batteries have become a critical challenge for the circular economy. In current unstructured disassembly scenarios, the dynamic nature of the environment severely limits the robustness of robotic perception, posing a significant barrier to autonomous disassembly in industrial applications. This paper proposes a continual learning framework based on Neuro-Symbolic task and motion planning (TAMP) to enhance the adaptability of embodied intelligence systems in dynamic environments. Our approach integrates a multimodal perception cross-validation mechanism into a bidirectional reasoning flow: the forward working flow dynamically refines and optimizes action strategies, while the backward learning flow autonomously collects effective data from historical task executions to facilitate continual system learning, enabling self-optimization. Experimental results show that the proposed framework improves the task success rate in dynamic disassembly scenarios from 81.68% to 100%, while reducing the average number of perception misjudgments from 3.389 to 1.128. This research provides a new paradigm for enhancing the robustness and adaptability of embodied intelligence in complex industrial environments.

cs.RO↗

Nonreciprocal Entanglement by Dynamically Encircling a Nexus

Nonreciprocal entanglement, characterized by inherently robust operation, is a cornerstone for quantum information processing and communications. However, it remains a great challenge to achieve nonreciprocal entanglement characterized by stability and robustness against environmental fluctuations. Here, we propose a universal nonlinear mechanism to engineer magnetic-free nonreciprocity in dissipative optomechanics by utilizing bistability, a phenomenon ubiquitous across nonlinear physical systems. By dynamically encircling the nexus of bistability, a cusp converged by the bistable surfaces, we obtain nonreciprocal displacement and then utilize it to achieve robust nonreciprocal entanglement. Owing to the unique landscape of bistability, our nonreciprocal displacement and entanglements exhibit stability and robustness through closed-loop operations. Our work presents a foundational framework for leveraging nonlinearity to achieve nonreciprocal quantum information processing. It paves new avenues for exploring nonreciprocal quantum information processing and designing backaction-immune quantum metrology with nonlinearity.

quant-ph↗

Unidirectional lasing via vacuum induced coherent in defective atomic lattice

We skillfully utilized vacuum induced coherence to amplify the probe light, and then successfully achieved both nonreciprocal reflection and lasing oscillation in a single physical system by leveraging the distributed feedback and spatial symmetry breaking effect of the one-dimensional defective atomic lattice. This innovative scheme for realizing unidirectional reflection lasing (URL) is based on both non-Hermitian degeneracy and spectral singularity (NHDSS, means $λ_{+}^{-1}\simeqλ_{-}^{-1}\rightarrow0$). Therefore, we analyze the modulation of parameters such as the lattice structure and external optical fields in this system to find NHDSS point, and further verified the conditions for its occurrence by solving the transcendental equation of susceptibility satisfying the NHDSS point, as well as analyzed its physical essence. Our mechanism is not only beneficial for the integration of photonic devices in quantum networks, but also greatly improves the efficiency of optical information transmission.

physics.optics↗

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

Existing human motion Q\&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across multiple query types without manually designed modules. Unlike existing implicit reasoning approaches that infer reasoning operations from question words, our model directly conditions on structured program functions, ensuring a more precise execution of reasoning steps. Additionally, we introduce a program-guided reading mechanism, which dynamically selects multi-level motion representations from a pretrained motion Vision Transformer (ViT), capturing both high-level semantics and fine-grained motion cues. The reasoning module iteratively refines memory representations, leveraging structured program functions to extract relevant information for different query types. Our model achieves state-of-the-art performance on Babel-QA and generalizes to a newly constructed motion Q\&A dataset based on HuMMan, demonstrating its adaptability across different motion reasoning datasets. Code and dataset are available at: https://github.com/LUNAProject22/IMoRe.

cs.CV↗

Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic

Action-driven stochastic human motion prediction aims to generate future motion sequences of a pre-defined target action based on given past observed sequences performing non-target actions. This task primarily presents two challenges. Firstly, generating smooth transition motions is hard due to the varying transition speeds of different actions. Secondly, the action characteristic is difficult to be learned because of the similarity of some actions. These issues cause the predicted results to be unreasonable and inconsistent. As a result, we propose two memory banks, the Soft-transition Action Bank (STAB) and Action Characteristic Bank (ACB), to tackle the problems above. The STAB stores the action transition information. It is equipped with the novel soft searching approach, which encourages the model to focus on multiple possible action categories of observed motions. The ACB records action characteristic, which produces more prior information for predicting certain actions. To fuse the features retrieved from the two banks better, we further propose the Adaptive Attention Adjustment (AAA) strategy. Extensive experiments on four motion prediction datasets demonstrate that our approach consistently outperforms the previous state-of-the-art. The demo and code are available at https://hyqlat.github.io/STABACB.github.io/.

cs.CV↗

HEAT:History-Enhanced Dual-phase Actor-Critic Algorithm with A Shared Transformer

For a single-gateway LoRaWAN network, this study proposed a history-enhanced two-phase actor-critic algorithm with a shared transformer algorithm (HEAT) to improve network performance. HEAT considers uplink parameters and often neglected downlink parameters, and effectively integrates offline and online reinforcement learning, using historical data and real-time interaction to improve model performance. In addition, this study developed an open source LoRaWAN network simulator LoRaWANSim. The simulator considers the demodulator lock effect and supports multi-channel, multi-demodulator and bidirectional communication. Simulation experiments show that compared with the best results of all compared algorithms, HEAT improves the packet success rate and energy efficiency by 15% and 95%, respectively.

cs.NI↗

Optimizing Multi-Gateway LoRaWAN via Cloud-Edge Collaboration and Knowledge Distillation

For large-scale multi-gateway LoRaWAN networks, this study proposes a cloud-edge collaborative resource allocation and decision-making method based on edge intelligence, HEAT-LDL (HEAT-Local Distill Lyapunov), which realizes collaborative decision-making between gateways and terminal nodes. HEAT-LDL combines the Actor-Critic architecture and the Lyapunov optimization method to achieve intelligent downlink control and gateway load balancing. When the signal quality is good, the network server uses the HEAT algorithm to schedule the terminal nodes. To improve the efficiency of autonomous decision-making of terminal nodes, HEAT-LDL performs cloud-edge knowledge distillation on the HEAT teacher model on the terminal node side. When the downlink decision instruction is lost, the terminal node uses the student model and the edge decider based on prior knowledge and local history to make collaborative autonomous decisions. Simulation experiments show that compared with the optimal results of all compared algorithms, HEAT-LDL improves the packet success rate and energy efficiency by 20.5% and 88.1%, respectively.

cs.NI↗

Energy-Efficient Flat Precoding for MIMO Systems

This paper addresses the suboptimal energy efficiency of conventional digital precoding schemes in multiple-input multiple-output (MIMO) systems. Through an analysis of the power amplifier (PA) output power distribution associated with conventional precoders, it is observed that these power distributions can be quite uneven, resulting in large PA backoff (thus low efficiency) and high power consumption. To tackle this issue, we propose a novel approach called flat precoding, which aims to control the flatness of the power distribution within a desired interval. In addition to reducing PA power consumption, flat precoding offers the advantage of requiring smaller saturation levels for PAs, which reduces the size of PAs and lowers the cost. To incorporate the concept of flat power distribution into precoding design, we introduce a new lower-bound per-antenna power constraint alongside the conventional sum power constraint and the upper-bound per-antenna power constraint. By adjusting the lower-bound and upper-bound values, we can effectively control the level of flatness in the power distribution. We then seek to find a flat precoder that satisfies these three sets of constraints while maximizing the weighted sum rate (WSR). In particular, we develop efficient algorithms to design weighted minimum mean squared error (WMMSE) and zero-forcing (ZF)-type precoders with controllable flatness features that maximize WSR. Numerical results demonstrate that complete flat precoding approaches, where the power distribution is a straight line, achieve the best trade-off between spectral efficiency and energy efficiency for existing PA technologies. We also show that the proposed ZF and WMMSE precoding methods can approach the performance of their conventional counterparts with only the sum power constraint, while significantly reducing PA size and power consumption.

cs.IT↗

LLM4GNAS: A Large Language Model Based Toolkit for Graph Neural Architecture Search

Graph Neural Architecture Search (GNAS) facilitates the automatic design of Graph Neural Networks (GNNs) tailored to specific downstream graph learning tasks. However, existing GNAS approaches often require manual adaptation to new graph search spaces, necessitating substantial code optimization and domain-specific knowledge. To address this challenge, we present LLM4GNAS, a toolkit for GNAS that leverages the generative capabilities of Large Language Models (LLMs). LLM4GNAS includes an algorithm library for graph neural architecture search algorithms based on LLMs, enabling the adaptation of GNAS methods to new search spaces through the modification of LLM prompts. This approach reduces the need for manual intervention in algorithm adaptation and code modification. The LLM4GNAS toolkit is extensible and robust, incorporating LLM-enhanced graph feature engineering, LLM-enhanced graph neural architecture search, and LLM-enhanced hyperparameter optimization. Experimental results indicate that LLM4GNAS outperforms existing GNAS methods on tasks involving both homogeneous and heterogeneous graphs.

cs.LG↗

Correlated Rydberg Electromagnetically Induced Transparencys

In the regime of Rydberg electromagnetically induced transparency, we study the correlated behaviors between the transmission spectra of a pair of probe fields passing through respective parallel one-dimensional cold Rydberg ensembles. Due to the van der Waals (vdW) interactions between Rydberg atoms, each ensemble exhibits a local optical nonlinearity, where the output EIT spectra are sensitive to both the input probe intensity and the photonic statistics. More interestingly, a nonlocal optical nonlinearity emerges between two spatially separated ensembles, as the probe transmissivity and probe correlation at the exit of one Rydberg ensemble can be manipulated by the probe field at the input of the other Rydberg ensemble. Realizing correlated Rydberg EITs holds great potential for applications in quantum control, quantum network, quantum walk and so on.

quant-ph↗

Drone Data Analytics for Measuring Traffic Metrics at Intersections in High-Density Areas

This study employed over 100 hours of high-altitude drone video data from eight intersections in Hohhot to generate a unique and extensive dataset encompassing high-density urban road intersections in China. This research has enhanced the YOLOUAV model to enable precise target recognition on unmanned aerial vehicle (UAV) datasets. An automated calibration algorithm is presented to create a functional dataset in high-density traffic flows, which saves human and material resources. This algorithm can capture up to 200 vehicles per frame while accurately tracking over 1 million road users, including cars, buses, and trucks. Moreover, the dataset has recorded over 50,000 complete lane changes. It is the largest publicly available road user trajectories in high-density urban intersections. Furthermore, this paper updates speed and acceleration algorithms based on UAV elevation and implements a UAV offset correction algorithm. A case study demonstrates the usefulness of the proposed methods, showing essential parameters to evaluate intersections and traffic conditions in traffic engineering. The model can track more than 200 vehicles of different types simultaneously in highly dense traffic on an urban intersection in Hohhot, generating heatmaps based on spatial-temporal traffic flow data and locating traffic conflicts by conducting lane change analysis and surrogate measures. With the diverse data and high accuracy of results, this study aims to advance research and development of UAVs in transportation significantly. The High-Density Intersection Dataset is available for download at https://github.com/Qpu523/High-density-Intersection-Dataset.

eess.IV↗

Perception- and Fidelity-aware Reduced-Reference Super-Resolution Image Quality Assessment

With the advent of image super-resolution (SR) algorithms, how to evaluate the quality of generated SR images has become an urgent task. Although full-reference methods perform well in SR image quality assessment (SR-IQA), their reliance on high-resolution (HR) images limits their practical applicability. Leveraging available reconstruction information as much as possible for SR-IQA, such as low-resolution (LR) images and the scale factors, is a promising way to enhance assessment performance for SR-IQA without HR for reference. In this letter, we attempt to evaluate the perceptual quality and reconstruction fidelity of SR images considering LR images and scale factors. Specifically, we propose a novel dual-branch reduced-reference SR-IQA network, \ie, Perception- and Fidelity-aware SR-IQA (PFIQA). The perception-aware branch evaluates the perceptual quality of SR images by leveraging the merits of global modeling of Vision Transformer (ViT) and local relation of ResNet, and incorporating the scale factor to enable comprehensive visual perception. Meanwhile, the fidelity-aware branch assesses the reconstruction fidelity between LR and SR images through their visual perception. The combination of the two branches substantially aligns with the human visual system, enabling a comprehensive SR image evaluation. Experimental results indicate that our PFIQA outperforms current state-of-the-art models across three widely-used SR-IQA benchmarks. Notably, PFIQA excels in assessing the quality of real-world SR images.

eess.IV↗

Antidirected hamiltonian paths in $k$-hypertournaments

A $k$-hypertournament $H$ on $n$ vertices is a pair $(V(H),A(H))$, where $V(H)$ is a set of vertices and $A(H)$ is a set of $k$-tuples of vertices, called arcs, such that for any $k$-subset $S$ of $V(H)$, $A(H)$ contains exactly one of the $k!$ $k$-tuples whose entries belong to $S$. Clearly, a 2-hypertournament is a tournament. An antidirected path in $H$ is a sequence $x_1 a_1 x_2 a_2 x_3 \ldots x_{t-1} a_{t-1} x_t$ of distinct vertices $x_1, x_2, \ldots, x_t$ and distinct arcs $a_1, a_{2},\ldots, a_{t-1}$ such that for any $i\in \{2,3,\ldots, t-1\}$, either $x_{i-1}$ precedes $x_{i}$ in $a_{i-1}$ and $x_{i+1}$ precedes $x_{i}$ in $a_{i}$, or $x_{i}$ precedes $x_{i-1}$ in $a_{i-1}$ and $x_{i}$ precedes $x_{i+1}$ in $a_{i}$. An antidirected path that includes all vertices of $H$ is known as an antidirected hamiltonian path. In this paper, we prove that except for four hypertournaments, $T_3^{c}, T_5^{c}, T_7^{c}$ and $H_{4}$, every $k$-hypertournament with $n$ vetices, where $2\leq k\leq n-1$, has an antidirected hamiltonian path, which extends Grünbaum's theorem on tournaments (except for three tournaments, $T_3^{c}, T_5^{c}$ and $T_7^{c}$, every tournament has an antidirected hamiltonian path).

math.CO↗