SearcharxivSearch

arXiv subjects

Yidong Huang

Publications and source records attributed to Yidong Huang.

At least 19 recordsLinked to original sources

Magic-free coexisting photonic and phononic moiré flat bands

Moiré flat bands enhance localization and interactions through suppressed group velocity, but existing approaches largely target a single physical field because distinct excitations generally require different, finely tuned magic configurations. Here we introduce a flat-band mechanism based on strong diffractive hybridization among moiré-folded bands. Period-mismatched modulations open distinct coupling channels whose hybridization renormalizes the band dispersion. An effective Hamiltonian shows that increasing the diffractive coupling progressively suppresses the group velocity, driving the system toward a flat-band regime without field-specific magic configurations. This coupling-induced mechanism enables band flattening across distinct physical excitations. We demonstrate this mechanism in a single-layer moiré optomechanical crystal, where photonic and phononic flat bands are simultaneously realized, and their localized modes and optomechanical interaction are experimentally observed. Beyond photonic and phononic systems, this mechanism may extend to other wave and quasiparticle platforms, providing a general route to co-localizing and coupling distinct physical fields in moiré systems.

physics.optics

Aligning Heterogeneous DFT Datasets: A Graph Neural Network Approach to Cross-Functional Formation Energies

Heterogeneous density functional theory (DFT) calculations, particularly plane-wave implementations, introduce systematic formation energy errors ranging from tens to hundreds of meV/atom, depending on the selection of exchange-correlation functionals, kinetic energy cutoffs, pseudopotentials, and dispersion corrections. As demonstrated by the MatPES dataset, identical structures can exhibit an average energy discrepancy of 107 meV/atom between PBE and r2SCAN calculations. Such method-dependent discrepancies hinder the integration of multi-source DFT data, greatly limiting the scale and quality of datasets for training robust materials AI models. Here, we resolve this fundamental data silo barrier via graph-based transfer learning. Leveraging 380,190 structurally paired PBE-r2SCAN entries from the MatPES database, we train a structure-aware graph neural network to predict cross-functional energy residuals and align inconsistent DFT energy scales. By adopting GPTFF model architecture, the model converts conventional PBE energies to r2SCAN-level accuracy with a mean absolute error of 14.3 meV/atom, compared with 18.2 meV/atom achieved by CHGNet. This versatile approach effectively upgrades massive legacy PBE datasets to high-precision r2SCAN standards. It enables reliable predictions of phase stability, battery voltage profiles, and reaction thermodynamics, while allowing the integration of multi-source DFT data to advance the development of high-performance materials foundation models.

cond-mat.mtrl-sci

The Mechanistic Emergence of Symbol Grounding in Language Models

Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Recent work has shown preliminary evidence that grounding may emerge in (vision-)language models trained at scale without using explicit grounding objectives. Yet, the specific loci of this emergence and the mechanisms that drive it remain largely unexplored. To address this problem, we introduce a controlled evaluation framework that systematically traces how symbol grounding arises within the internal computations through mechanistic and causal analysis. Our findings show that grounding concentrates in middle-layer computations and is implemented through the aggregate mechanism, where attention heads aggregate the environmental ground to support the prediction of linguistic forms. This phenomenon replicates in multimodal dialogue and across architectures (Transformers and state-space models), but not in unidirectional LSTMs. Our results provide behavioral and mechanistic evidence that symbol grounding can emerge in language models, with practical implications for predicting and potentially controlling the reliability of generation.

cs.CL

Trapping 11,000 Atoms in a Tweezer Array Generated by a Single Metasurface

The scalability of physical qubit numbers is a central challenge toward a universal fault-tolerant quantum computer. The inherent scalability of atom array quantum computers stems from the identical nature of atomic qubits, so the available qubit resource is primarily limited by the number of atoms that can be trapped and controlled. Here, we robustly trap 11,000 individual atoms in a tweezer array, thereby enabling the available qubit resource to reach the tens-of-thousands scale for the first time among all quantum computation platforms. This advance is enabled by a single metasurface, approximately 2 cm in diameter, that generates the entire tweezer array without the need for microscope objectives, thereby maximizing laser-power efficiency. The large aperture ensures a working distance of about 1.5 cm, allowing the metasurface to be placed outside the vacuum cell and avoiding the technical complications of in-vacuum operation. We further characterize the randomly loaded atom array using the statistical theory of percolation phase transitions. This work takes an important first step toward a quantum computer at the 10,000-qubit scale.

quant-ph

Stabilization-Free H(curl) and H(div)-Conforming Virtual Element Method

Standard Virtual Element Method (VEM) requires stabilization terms that significantly affect the numerical computation performance. In this work, we propose a stabilization-free VEM for general order \(\mathbf{H}(\operatorname{\mathbf{curl}})\) and \(\mathbf{H}(\operatorname{div})\)-conforming spaces by constructing novel serendipity projectors and corresponding serendipity spaces with minimum number of DoFs. Our approach handles the full De Rham complex chain in \(\mathbb{R}^3\) while preserving essential properties including boundary continuity and commutativity. Since the number of DoFs are minimized, computational overhead is greatly reduced. The optimal approximation properties are rigorously proven and validated through Maxwell eigenvalue problems with numerical experiments.

math.NA

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general video quality, extending it to human motion remains bottlenecked by a reward signal that cannot reliably score motion realism. Existing video rewards primarily rely on 2D perceptual signals, without explicitly modeling the 3D body state, contact, and dynamics underlying articulated human motion, and often assign high scores to videos with floating bodies or physically implausible movements. To address this, we propose PhyMotion, a structured, fine-grained motion reward that grounds recovered 3D human trajectories in a physics simulator and evaluates motion quality along multiple dimensions of physical feasibility. Concretely, we recover SMPL body meshes from generated videos, retarget them onto a humanoid in the MuJoCo physics simulator, and evaluate the resulting motion along three axes: kinematic plausibility, contact and balance consistency, and dynamic feasibility. Each component provides a continuous and interpretable signal tied to a specific aspect of motion quality, allowing the reward to capture which aspects of motion are physically correct or violated. Experiments show that PhyMotion achieves stronger correlation with human judgments than existing reward formulations. These gains carry over to RL-based post-training, where optimizing PhyMotion leads to larger and more consistent improvements than optimizing existing rewards, improving motion realism across both autoregressive and bidirectional video generators under both automatic metrics and blind human evaluation (+68 Elo gain). Ablations show that the three axes provide complementary supervision signals, while the reward preserves overall video generation quality with only modest training overhead.

cs.CV

Quantum Interaction Between Free Electrons and Light Involving First-order and Second-order Process

Photon-induced Near-field Electron Microscopy (PINEM) effect has revealed the quantum interaction between free electrons and optical near filed, which demonstrated plenty of novel phenomena of manipulating free electron wave packet and detecting/shaping quantum photonic states. However, free electrons generally only absorb/emit one photon at a time, while the physical mechanism and phenomena of free electron-two-photon interaction have not been studied yet. Moreover, the relationship between PINEM and Kapitza-Dirac (KD) effect and nonlinear Compton scattering is still unclear. Here we develop the full quantum theory of electron-photon interaction considering the two-photon process. It is revealed that the emission/absorption of two photons by electrons can be greatly enhanced by manipulating the electric field component of optical near field, and the quantum interference between single-photon and two-photon processes can occur in some circumstances, which affects the photon number state, electron energy states and electron-photon entanglement. Meanwhile, it is found that the KD effect (elastic electron-photon scattering) and nonlinear Compton scattering (inelastic electron-photon scattering) are also a kind of two-photon process and the distribution of electrons can be deduced analytically based on the full quantum theory. Our work uncovers the possible abundant phenomena when free electron interacting with two photons, paves the way for more in-depth studies of nonlinear processes in electron-photon quantum interactions in the future.

quant-ph

Arbitrary-order exceptional points in a nanomechanical cavity

Higher-order exceptional points (EPs) govern non-Hermitian system dynamics through their enriched and sharpened spectral topology, yet the intrinsic topological fragility hinders robust experimental realization. Here, we present a scalable architecture that implements arbitrary-order EPs via a recurrent network comprising a single nanomechanical resonator and unlimited virtual resonators. We experimentally realize mechanical EPs up to the seventh order and confirm this architecture's scalability. Moreover, we reveal that the fundamental noise component and the measured signal share the same system coupling channel and thus undergo identical root-response amplification near EPs of arbitrary order, consistent with our signal-to-noise ratio measurements. Our work establishes a general platform for exploring higher-order EP-based phenomena while clarifying the fundamental boundary of non-Hermitian sensitivity enhancement across diverse physical systems.

physics.optics

Direct Generation of an Array with 78400 Optical Tweezers Using a Single Metasurface

Scalability remains a major challenge in building practical fault-tolerant quantum computers. Currently, the largest number of qubits achieved across leading quantum platforms ranges from hundreds to thousands. In atom arrays, scalability is primarily constrained by the capacity to generate large numbers of optical tweezers, and conventional techniques using acousto-optic deflectors or spatial light modulators struggle to produce arrays much beyond $\sim 10,000$ tweezers. Moreover, these methods require additional microscope objectives to focus the light into micrometer-sized spots, which further complicates system integration and scalability. Here, we demonstrate the experimental generation of an optical tweezer array containing $280\times 280$ spots using a metasurface, nearly an order of magnitude more than most existing systems. The metasurface leverages a large number of subwavelength phase-control pixels to engineer the wavefront of the incident light, enabling both large-scale tweezer generation and direct focusing into micron-scale spots without the need for a microscope. This result shifts the scalability bottleneck for atom arrays from the tweezer generation hardware to the available laser power. Furthermore, the array shows excellent intensity uniformity exceeding $90\%$, making it suitable for homogeneous single-atom loading and paving the way for trapping arrays of more than $10,000$ atoms in the near future.

cond-mat.quant-gas

Planning with Sketch-Guided Verification for Physics-Aware Video Generation

Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-shot plans that are typically limited to simple motions, or iterative refinement which requires multiple calls to the video generator, incuring high computational cost. To overcome these limitations, we propose SketchVerify, a training-free, sketch-verification-based planning framework that improves motion planning quality with more dynamically coherent trajectories (i.e., physically plausible and instruction-consistent motions) prior to full video generation by introducing a test-time sampling and verification loop. Given a prompt and a reference image, our method predicts multiple candidate motion plans and ranks them using a vision-language verifier that jointly evaluates semantic alignment with the instruction and physical plausibility. To efficiently score candidate motion plans, we render each trajectory as a lightweight video sketch by compositing objects over a static background, which bypasses the need for expensive, repeated diffusion-based synthesis while achieving comparable performance. We iteratively refine the motion plan until a satisfactory one is identified, which is then passed to the trajectory-conditioned generator for final synthesis. Experiments on WorldModelBench and PhyWorldBench demonstrate that our method significantly improves motion quality, physical realism, and long-term consistency compared to competitive baselines while being substantially more efficient. Our ablation study further shows that scaling up the number of trajectory candidates consistently enhances overall performance.

cs.CV

On-chip Time-bin to Path Qubit Encoding Converter via Thin Film Lithium Niobate Photonics Chip

The development of quantum internet demands on-chip quantum processor nodes and interconnection between the nodes. Path-encoded photonic qubits are suitable for on-chip quantum information processors, while time-bin encoded ones are good at long-distance communication. It is necessary to develop an on-chip converter between the two encodings to satisfy the needs of the quantum internet. In this work, a quantum photonic circuit is proposed to convert time-bin-encoded photonic qubits to path-encoded ones via a thin-film lithium niobate high-speed optical switch and low-loss matched optical delay lines. The performance of the encoding converter is demonstrated by the experiment of time-bin to path encoding conversion on the fabricated sample chip. The converted path qubits have an average fidelity higher than 97%. The potential of the encoding converter on applications in quantum networks is demonstrated by the experiments of entanglement distribution and quantum key distribution. The results show that the on-chip encoding converter can serve as a foundational component in the future quantum internet, bridging the gap between quantum information transmission and on-chip processing based on photons.

quant-ph

Multi-channel electrically tunable varifocal metalens with compact multilayer polarization-dependent metasurfaces and liquid crystals

As an essential module of optical systems, varifocal lens usually consists of multiple mechanically moving lenses along the optical axis. The recent development of metasurfaces with tunable functionalities holds the promise of miniaturizing varifocal lens. However, existing varifocal metalenses are hard to combine electrical tunability with scalable number and range of focal lengths, thus limiting the practical applications. Our previous work shows that the electrically tunable channels could be increased to 2N by cascading N polarization-dependent metasurfaces with liquid crystals (LCs). Here, we demonstrated a compact eight-channel electrically tunable varifocal metalens with three single-layer polarization-multiplexed bi-focal metalens and three LC cells. The total thickness of the device is ~6 mm, while the focal lengths could be switched among eight values within the range of 3.6 to 9.6 mm. The scheme is scalable in number and range of focal lengths and readily for further miniaturization. We believe that our proposal would open new possibilities of miniaturized imaging systems, AR/VR displays, LiDAR, etc.

physics.optics

Polarization Decoupling Multi-Port Beam-Splitting Metasurface for Miniaturized Magneto-Optical Trap

In regular magneto-optical trap (MOT) systems, the delivery of six circularly polarized (CP) cooling beams requires complex and bulky optical arrangements including waveplates, mirrors, retroreflectors, etc. To address such technique challenges, we have proposed a beam delivery system for miniaturized MOT entirely based on meta-devices. The key component is a novel multi-port beam-splitting (PD-MPBS) metasurface that relies on both propagation phase and geometric phase. The fabricated samples exhibit high beam-splitting power uniformity (within 4.4%) and polarization purities (91.29%~93.15%). By leveraging such beam-splitting device as well as reflective beam-expanding meta-device, an integrated six-beam delivery system for miniaturized MOT application has been implemented. The experimental results indicate that six expanded beams have been successfully delivered with uniform power (within 9.5%), the desired CP configuration and large overlapping volume (76.2 mm^3). We believe that a miniaturized MOT with the proposed beam delivery system is very promising for portable application of cold atom technology in precision measurement, atomic clock, quantum simulation and computing, etc.

physics.optics

Demonstration of Time-reversal Symmetric Two-Dimensional Photonic Topological Anderson Insulator

Recently, the impact of disorder on topological properties has attracted significant attention in photonics, especially the intriguing disorder-induced topological phase transitions in photonic topological Anderson insulators (PTAIs). However, the reported PTAIs are based on time-reversal symmetry broken systems or quasi-three-dimensional time-reversal invariant system, both of which would limit the applications in integrated optics. Here, we realize a time-reversal symmetric two-dimensional PTAI on silicon platform within the near-IR wavelength range, taking the advantageous valley degree of freedom of photonic crystal. A low-threshold topological Anderson phase transition is observed by applying disorder to the critical topologically trivial phase. Conversely, we have also realized extremely robust topologically protected edge states based on the stable topological phase. Both two phenomena are validated through theoretical Dirac Hamiltonian analysis, numerical simulations, and experimental measurements. Our proposed structure holds promise to achieve near-zero topological phase transition thresholds, which breaks the conventional cognition that strong disorder is required to induce the phase transition. It significantly alleviates the difficulty of manipulating disorder and could be extended to other systems, such as condensed matter systems where strong disorder is hard to implement. This work is also beneficial to construct highly robust photonic integrated circuits serving for on-chip photonic and quantum optic information processing. Moreover, this work also provides an outstanding platform to investigate on-chip integrated disordered systems.

physics.optics

Chip-to-chip photonic quantum teleportation over optical fibers of 12.3km

Quantum teleportation is a crucial function in quantum networks. The implementation of photonic quantum teleportation could be highly simplified by quantum photonic circuits. To extend chip-to-chip teleportation distance, more effort is needed on both chip design and system implementation. In this work, we demonstrate a chip-to-chip photonic quantum teleportation over optical fibers under the scenario of star-topology quantum network. Time-bin encoded quantum states are used to achieve a long teleportation distance. Three photonic quantum circuits are designed and fabricated on a single chip, each serving specific functions: heralded single-photon generation at the user node, entangled photon pair generation and Bell state measurement at the relay node, and projective measurement of the teleported photons at the central node. The unbalanced Mach-Zehnder interferometers (UMZI) for time-bin encoding in these quantum photonic circuits are optimized to reduce insertion losses and suppress noise photons generated on the chip. Besides, an active feedback system is employed to suppress the impact of fiber length fluctuation between the circuits, achieving a stable quantum interference for the Bell state measurement in the relay node. As the result, a photonic quantum teleportation over optical fibers of 12.3km is achieved based on these quantum photonic circuits, showing the potential of chip integration on the development of quantum networks.

quant-ph

Spectral Convolutional Neural Network Chip for In-sensor Edge Computing of Incoherent Natural Light

Convolutional neural networks (CNNs) are representative models of artificial neural networks (ANNs). However, the considerable power consumption and limited computing speed of electrical computing platforms restrict further CNN development on edge devices. Optical neural networks are considered next-generation physical implementations of ANNs, but their capabilities are limited by on-chip integration scale and requirement for coherent light sources. This study proposes a spectral convolutional neural network (SCNN) of incoherent natural light by an optical convolutional layer (OCL) and a reconfigurable electrical backend. The OCL is implemented by integrating very large-scale, pixel-aligned spectral filters on a CMOS image sensor on a 12-inch wafer, facilitating highly parallel spectral vector-inner products of incident light. It accepts broadband incoherent natural light containing two spatial and one spectral dimension directly as input with the function of matter meta-imaging. This unique optoelectronic framework empowers in-sensor optical analog computing at extremely high energy efficiency because the OCL is driven by the energy of the information carrier, i. e. natural light. To the best of our knowledge, this is the first integrated optical computing utilizing natural light. We employ the same SCNN chip for completely different real-world complex tasks,and achieve accuracies of over 96% for pathological diagnosis and almost 100% for face anti-spoofing at video rates. The SCNN framework has an unprecedented new function of substance identification, provides a feasible optoelectronic and integrated optical CNN implementation for edge devices or cellphones, providing them with practical and powerful edge computing abilities and facilitating diverse applications, such as intelligent robotics, industrial automation, medical diagnosis, and remote sensing.

physics.optics

Map Optical Properties to Subwavelength Structures Directly via a Diffusion Model

Subwavelength photonic structures and metamaterials provide revolutionary approaches for controlling light. The inverse design methods proposed for these subwavelength structures are vital to the development of new photonic devices. However, most of the existing inverse design methods cannot realize direct mapping from optical properties to photonic structures but instead rely on forward simulation methods to perform iterative optimization. In this work, we exploit the powerful generative abilities of artificial intelligence (AI) and propose a practical inverse design method based on latent diffusion models. Our method maps directly the optical properties to structures without the requirement of forward simulation and iterative optimization. Here, the given optical properties can work as "prompts" and guide the constructed model to correctly "draw" the required photonic structures. Experiments show that our direct mapping-based inverse design method can generate subwavelength photonic structures at high fidelity while following the given optical properties. This may change the method used for optical design and greatly accelerate the research on new photonic devices.

physics.optics

SUANPAN: Scalable Photonic Linear Vector Machine

Photonic linear operation is a promising approach to handle the extensive vector multiplications in artificial intelligence techniques due to the natural bosonic parallelism and high-speed information transmission of photonics. Although it is believed that maximizing the interaction of the light beams is necessary to fully utilize the parallelism and tremendous efforts have been made in past decades, the achieved dimensionality of vector-matrix multiplication is very limited due to the difficulty of scaling up a tightly interconnected or highly coupled optical system. Additionally, there is still a lack of a universal photonic computing architecture that can be readily merged with existing computing system to meet the computing power demand of AI techniques. Here, we propose a programmable and reconfigurable photonic linear vector machine to perform only the inner product of two vectors, formed by a series of independent basic computing units, while each unit is just one pair of light-emitter and photodetector. Since there is no interaction among light beams inside, extreme scalability could be achieved by simply duplicating the independent basic computing unit while there is no requirement of large-scale analog-to-digital converter and digital-to-analog converter arrays. Our architecture is inspired by the traditional Chinese Suanpan or abacus and thus is denoted as photonic SUANPAN. As a proof of principle, SUANPAN architecture is implemented with an 8*8 vertical cavity surface emission laser array and an 8*8 MoTe2 two-dimensional material photodetector array. We believe that our proposed photonic SUANPAN is capable of serving as a fundamental linear vector machine that can be readily merged with existing electronic digital computing system and is potential to enhance the computing power for future various AI applications.

physics.optics