SearcharxivSearch

arXiv subjects

Ning Han

Publications and source records attributed to Ning Han.

17 recordsLinked to original sources

Causality Sum Rules in Conventional Scattering Matrices

Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is explicit in the conventional incoming-outgoing matrix, whereas causality sum rules are usually formulated only after transforming the response into auxiliary variables. Here we show that these rules can be written directly in the conventional scattering matrix by removing the time advance introduced by the reference domain. Using the earliest-arrival delay of each channel, we define a domain-delayed matrix that preserves real-frequency passivity while restoring the causal time origin. Under explicit analyticity, transparency, and regularity assumptions, this matrix becomes a Schur function, enabling a Cayley-Herglotz construction. The resulting projected and determinant bounds constrain coherent channel superpositions and aggregate multichannel loss. The framework recovers Rozanov's absorber limit and spherical-multipole sum rules, while extending causality bounds to measurable quantities including insertion loss, suppressed singular-value channels, and conditional lossless delay-bandwidth trade-offs. Our work directly connects fundamental causality theory with experimentally accessible scattering data. The initial theoretical route is autonomously explored by Qiushi Engine, an AI research system for open-ended scientific discovery, and subsequently verified, refined, and developed by the authors, demonstrating a hybrid AI-human discovery workflow.

physics.optics

World Models for Robotic Manipulation: A Survey

Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world models provide this capability by predicting task-relevant future evolution under robot intervention, yet the term now spans latent dynamics models, action-conditioned video generators, three- and four-dimensional scene predictors, physics-informed simulators, and predictive modules inside vision-language-action systems. This breadth has fragmented the literature and obscured the design choices that matter for manipulation. We survey world models for robotic manipulation through three questions: what future representation is predicted, how prediction is connected to action, and when prediction is used in the robot-learning pipeline. We operationally define a world model as an action-conditioned predictive system and distinguish it from perception modules, inverse models, policies, rewards, and value functions. We then organize existing work into five representation families, develop a functional taxonomy that separates integrated prediction-action models from explicit predictive planners, and characterize infrastructure roles including synthetic experience generation, candidate filtering, search-based evaluation, learned environments, and outcome verification. We further map these roles across pretraining, post-training, and inference adaptation, review 34 manipulation datasets, and synthesize evaluation protocols for predictive fidelity, task performance, and simulator reliability. This survey shows that world models are evolving from task-specific dynamics predictors into predictive infrastructure for robot learning, while exposing open challenges in contact modeling, hallucination control, action alignment, and benchmarking under closed-loop use.

cs.RO

End-to-end autonomous scientific discovery on a real optical platform

Scientific research has long been human-led, driving new knowledge and transformative technologies through the continual revision of questions, methods and claims as evidence accumulates. Although large language model (LLM)-based agents are beginning to move beyond assisting predefined research workflows, none has yet demonstrated end-to-end autonomous discovery in a real physical system that produces a nontrivial result supported by experimental evidence. Here we introduce Qiushi Discovery Engine, an LLM-based agentic system for end-to-end autonomous scientific discovery on a real optical platform. Qiushi Engine combines nonlinear research phases, Meta-Trace memory and a dual-layer architecture to maintain adaptive and stable research trajectories across long-horizon investigations involving thousands of LLM-mediated reasoning, measurement and revision actions. It autonomously reproduces a published transmission-matrix experiment on a non-original platform and converts an abstract coherence-order theory into experimental observables, providing, to our knowledge, the first observation of this class of coherence-order structure. More importantly, in an open-ended study involving 145.9 million tokens, 3,242 LLM calls, 1,242 tool calls, 163 research notes and 44 scripts, Qiushi Engine proposes and experimentally validates optical bilinear interaction, a physical mechanism structurally analogous to a core operation in Transformer attention. This AI-discovered mechanism suggests a route towards high-speed, energy-efficient optical hardware for pairwise computation. To our knowledge, this is the first demonstration of an AI agentic system autonomously identifying and experimentally validating a nontrivial, previously unreported physical mechanism, marking a milestone for research-level autonomous agents.

cs.AI

Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection

Recent weakly supervised video anomaly detection methods have achieved significant advances by employing unified frameworks for joint optimization. However, this paradigm is limited by a fundamental sensitivity-stability trade-off, as the conflicting objectives for detecting transient and sustained anomalies lead to either fragmented predictions or over-smoothed responses. To address this limitation, we propose DeSC, a novel Decoupled Sensitivity-Consistency framework that trains two specialized streams using distinct optimization strategies. The temporal sensitivity stream adopts an aggressive optimization strategy to capture high-frequency abrupt changes, whereas the semantic consistency stream applies robust constraints to maintain long-term coherence and reduce noise. Their complementary strengths are fused through a collaborative inference mechanism that reduces individual biases and produces balanced predictions. Extensive experiments demonstrate that DeSC establishes new state-of-the-art performance by achieving 89.37% AUC on UCF-Crime (+1.29%) and 87.18% AP on XD-Violence (+2.22%). Code is available at https://github.com/imzht/DeSC.

cs.CV

Realizing anomalous Floquet non-Abelian band topology in photonic scattering networks

The concept of multi-gap topology has recently been shown to give rise to uncharted phases beyond conventional single-gap classifications. These phases relate to band nodes with non-Abelian quaternion charges and momentum-space braiding processes characterized by new invariants such as paradigmatic Euler class, phenomena that intrinsically require at least two spatial dimensions. Extending such phases into the non-equilibrium regime is predicted to unlock even richer multi-gap topologies beyond static settings, yet their experimental realization has remained elusive due to the stringent requirements on dimensionality, symmetry, and dynamical control. Here, we theoretically demonstrate and, for the first time, experimentally realize two-dimensional (2D) Floquet non-Abelian band topology in photonic scattering networks. Within this platform, we uncover a sequence of topological phenomena unique to 2D multi-gap systems far from equilibrium, including anomalous multi-gap phases interconnected by band nodes, Floquet Euler transfer, gapped phases with anomalous Dirac string configurations, and Floquet-induced non-Abelian braiding of band nodes. In addition, we observe Floquet-periodic anomalous edge states across multiple gaps, providing experimental signatures of these sought-after 2D multi-gap Floquet topological phases. Our results establish photonic scattering networks as a practical and versatile route to non-Abelian Floquet systems, opening avenues for dynamical topological physics with braiding capability and robust photonic functionalities.

physics.optics

Observation of Custodial Chiral Symmetry in Memristive Topological Insulators

The concept of custodial symmetry, a residual symmetry that protects physical observables from large quantum corrections, has been a cornerstone of high-energy physics, but its experimental observation has remained unexplored. Building on recent theoretical work [Phys. Rev. Lett. 128, 097701 (2022)], we report the first experimental observation of classical analog of custodial chiral symmetry in a memristive Su-Schrieffer-Heeger (SSH) circuit. We provide direct experimental evidence for custodial symmetry through the measurement of the correction to the Lagrangian. This Lagrangian correction, which mimics a mass term in field theory, vanishes smoothly as the perturbation is reduced. We also demonstrate that topological edge states in the memristive SSH circuit remain localized at the boundary, protected by custodial chiral symmetry. This work opens new avenues for emulating field-theoretic symmetries and nonlinear dynamics in memristive platforms.

cond-mat.mes-hall

IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval

In recent years, significant developments have been made in both video retrieval and video moment retrieval tasks, which respectively retrieve complete videos or moments for a given text query. These advancements have greatly improved user satisfaction during the search process. However, previous work has failed to establish meaningful "interaction" between the retrieval system and the user, and its one-way retrieval paradigm can no longer fully meet the personalization and dynamic needs of at least 80.8\% of users. In this paper, we introduce the Interactive Video Corpus Retrieval (IVCR) task, a more realistic setting that enables multi-turn, conversational, and realistic interactions between the user and the retrieval system. To facilitate research on this challenging task, we introduce IVCR-200K, a high-quality, bilingual, multi-turn, conversational, and abstract semantic dataset that supports video retrieval and even moment retrieval. Furthermore, we propose a comprehensive framework based on multi-modal large language models (MLLMs) to help users interact in several modes with more explainable solutions. The extensive experiments demonstrate the effectiveness of our dataset and framework.

cs.CV

GrOCE:Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models

Concept erasure aims to remove harmful, inappropriate, or copyrighted content from text-to-image diffusion models while preserving non-target semantics. However, existing methods either rely on costly fine-tuning or apply coarse semantic separation, often degrading unrelated concepts and lacking adaptability to evolving concept sets. In this paper, we propose Graph-Guided Online Concept Erasure (GrOCE), a training-free framework that performs precise and context-aware online removal of target concepts. GrOCE constructs dynamic semantic graphs to identify clusters of target concepts and selectively suppress their influence within text prompts. It consists of three synergistic components: (1) dynamic semantic graph construction (Construct) incrementally builds a weighted graph over vocabulary concepts to capture semantic affinities; (2) adaptive cluster identification (Identify) extracts a target concept cluster through multi-hop traversal and diffusion-based scoring to quantify semantic influence; and (3) selective severing (Sever) removes semantic components associated with the target cluster from the text prompt while retaining non-target semantics and the global sentence structure. Extensive experiments demonstrate that GrOCE achieves state-of-the-art performance on the Concept Similarity (CS) and Fr\'echet Inception Distance (FID) metrics, offering efficient, accurate, and stable concept erasure.

cs.CV

Prescribed Performance Control of Deformable Object Manipulation in Spatial Latent Space

Manipulating three-dimensional (3D) deformable objects presents significant challenges for robotic systems due to their infinite-dimensional state space and complex deformable dynamics. This paper proposes a novel model-free approach for shape control with constraints imposed on key points. Unlike existing methods that rely on feature dimensionality reduction, the proposed controller leverages the coordinates of key points as the feature vector, which are extracted from the deformable object's point cloud using deep learning methods. This approach not only reduces the dimensionality of the feature space but also retains the spatial information of the object. By extracting key points, the manipulation of deformable objects is simplified into a visual servoing problem, where the shape dynamics are described using a deformation Jacobian matrix. To enhance control accuracy, a prescribed performance control method is developed by integrating barrier Lyapunov functions (BLF) to enforce constraints on the key points. The stability of the closed-loop system is rigorously analyzed and verified using the Lyapunov method. Experimental results further demonstrate the effectiveness and robustness of the proposed method.

cs.RO

Realization of Weyl elastic metamaterials with spin skyrmions

Topological elastic metamaterials provide a topologically robust way to manipulate the phononic energy and information beyond the conventional approaches. Among various topological elastic metamaterials, Weyl elastic metamaterials stand out, as they are unique to three dimensions and exhibit numerous intriguing phenomena and potential applications. To date, however, the realization of Weyl elastic metamaterials remains elusive, primarily due to the full-vectoral nature of elastic waves and the complicated couplings between polarizations, leading to complicated and tangled three-dimensional (3D) bandstructures that unfavorable for experimental demonstration. Here, we overcome the challenge and realize an ideal, 3D printed, all-metallic Weyl elastic metamaterial with low dissipation losses. Notably, the elastic spin of the excitations around the Weyl points exhibits skyrmion textures, a topologically stable structure in real space. Utilizing 3D laser vibrometry, we reveal the projection of the Weyl points, the Fermi arcs and the unique spin characteristics of the topological surface states. Our work extends the Weyl metamaterials to elastic waves and paves a topological way to robust manipulation of elastic waves in 3D space.

physics.app-ph

Photonic antiferromagnetic topological insulator with a single surface Dirac cone

Antiferromagnetism, characterized by magnetic moments aligned in alternating directions with a vanished ensemble average, has garnered renewed interest for its potential applications in spintronics and axion dynamics. The synergy between antiferromagnetism and topology can lead to the emergence of an exotic topological phase unique to certain magnetic order, termed antiferromagnetic topological insulators (AF TIs). A hallmark signature of AF TIs is the presence of a single surface Dirac cone--a feature typically associated with strong three-dimensional (3D) topological insulators--only on certain symmetry-preserving crystal terminations. However, the direct observation of this phenomenon poses a significant challenge. Here, we have theoretically and experimentally discovered a 3D photonic AF TI hosting a single surface Dirac cone protected by the combined symmetry of time reversal and half-lattice translation. Conceptually, our setup can be viewed as a z-directional stack of two-dimensional Chern insulators, with adjacent layers oppositely magnetized to form a 3D type-A AF configuration. By measuring both bulk and surface states, we have directly observed the symmetry-protected gapless single-Dirac-cone surface state, which shows remarkable robustness against random magnetic disorders. Our work constitutes the first realization of photonic AF TIs and photonic analogs of strong topological insulators, opening a new chapter for exploring novel topological photonic devices and phenomena that incorporate additional magnetic degrees of freedom.

physics.optics

A high-performance nitrogen-rich ZIF-8-derived Fe-Co-NC electrocatalyst for the oxygen reduction reaction

Exploring and developing low-cost, high-performance, and stable catalysts for the oxygen reduction reaction (ORR) is of great importance, though progress is ongoing. In this work, a simple evaporation-pyrolysis method is employed to synthesize 3D porous electrocatalysts, referred to as Fe-Co-NC, which are carbonization products at 900{\deg}C. These catalysts are composed of carbon nanoparticles and metallic FeCo doped nitrogen-enriched carbon nanotubes, produced through the carbonization of pristine ZIF-8, resulting in highly efficient and durable ORR electrocatalysts. The Fe-Co-NC structure exhibits a 3D open porous texture, abundant active sites, favorable nitrogen bonding, and a high specific surface area, all of which contribute to excellent ORR activity. The optimal Fe-Co-NC catalyst demonstrates exceptional ORR performance, with a high onset potential of 0.96 V and a half-wave potential of 0.86 V versus RHE, rivaling the performance of commercially available Pt/C in an alkaline electrolyte. Additionally, the Fe-Co-NC catalyst shows superior stability compared to Pt/C, which is crucial for the development of novel electrocatalysts based on non-precious metals.

cond-mat.mtrl-sci

Three-dimensional topological valley photonics

Topological valley photonics, which exploits valley degree of freedom to manipulate electromagnetic waves, offers a practical and effective pathway for various classical and quantum photonic applications across the entire spectrum. Current valley photonics, however, has been limited to two dimensions, which typically suffer from out-of-plane losses and can only manipulate the flow of light in planar geometries. Here, we have theoretically and experimentally developed a framework of three-dimensional (3D) topological valley photonics with a complete photonic bandgap and vectorial valley contrasting physics. Unlike the two-dimensional counterparts with a pair of valleys characterized by scalar valley Chern numbers, the 3D valley systems exhibit triple pairs of valleys characterized by valley Chern vectors, enabling the creation of vectorial bulk valley vortices and canalized chiral valley surface states. Notably, the valley Chern vectors and the circulating propagation direction of the valley surface states are intrinsically governed by the right-hand-thumb rule. Our findings reveal the vectorial nature of the 3D valley states and highlight their potential applications in 3D waveguiding, directional radiation, and imaging.

physics.optics

Boundary-induced topological chiral extended states in Weyl metamaterial waveguides

In topological physics, it is commonly understood that the existence of the boundary states of a topological system is inherently dictated by its bulk. A classic example is that the surface Fermi arc states of a Weyl system are determined by the chiral charges of Weyl points within the bulk. Contrasting with this established perspective, here, we theoretically and experimentally discover a family of topological chiral bulk states extending over photonic Weyl metamaterial waveguides, solely induced by the waveguide boundaries, independently of the waveguide width. Notably, these bulk states showcase discrete momenta and function as wormhole tunnels that connect Fermi-arc surface states living in different two dimensional spaces via a third dimension. Our work offers a magneticfield-free mechanism for robust chiral bulk transport of waves and highlights the boundaries as a new degree of freedom to regulate bulk Weyl quasiparticles.

physics.optics

Real higher-order Weyl photonic crystal

Higher-order Weyl semimetals are a family of recently predicted topological phases simultaneously showcasing unconventional properties derived from Weyl points, such as chiral anomaly, and multidimensional topological phenomena originating from higher-order topology. The higher-order Weyl semimetal phases, with their higher-order topology arising from quantized dipole or quadrupole bulk polarizations, have been demonstrated in phononics and circuits. Here, we experimentally discover a class of higher-order Weyl semimetal phase in a three-dimensional photonic crystal (PhC), exhibiting the concurrence of the surface and hinge Fermi arcs from the nonzero Chern number and the nontrivial generalized real Chern number, respectively, coined a real higher-order Weyl PhC. Notably, the projected two-dimensional subsystem with kz = 0 is a real Chern insulator, belonging to the Stiefel-Whitney class with real Bloch wavefunctions, which is distinguished fundamentally from the Chern class with complex Bloch wavefunctions. Our work offers an ideal photonic platform for exploring potential applications and material properties associated with the higher-order Weyl points and the Stiefel-Whitney class of topological phases.

cond-mat.mes-hall

Efficient Cross-Modal Video Retrieval with Meta-Optimized Frames

Cross-modal video retrieval aims to retrieve the semantically relevant videos given a text as a query, and is one of the fundamental tasks in Multimedia. Most of top-performing methods primarily leverage Visual Transformer (ViT) to extract video features [1, 2, 3], suffering from high computational complexity of ViT especially for encoding long videos. A common and simple solution is to uniformly sample a small number (say, 4 or 8) of frames from the video (instead of using the whole video) as input to ViT. The number of frames has a strong influence on the performance of ViT, e.g., using 8 frames performs better than using 4 frames yet needs more computational resources, resulting in a trade-off. To get free from this trade-off, this paper introduces an automatic video compression method based on a bilevel optimization program (BOP) consisting of both model-level (i.e., base-level) and frame-level (i.e., meta-level) optimizations. The model-level learns a cross-modal video retrieval model whose input is the "compressed frames" learned by frame-level optimization. In turn, the frame-level optimization is through gradient descent using the meta loss of video retrieval model computed on the whole video. We call this BOP method as well as the "compressed frames" as Meta-Optimized Frames (MOF). By incorporating MOF, the video retrieval model is able to utilize the information of whole videos (for training) while taking only a small number of input frames in actual implementation. The convergence of MOF is guaranteed by meta gradient descent algorithms. For evaluation, we conduct extensive experiments of cross-modal video retrieval on three large-scale benchmarks: MSR-VTT, MSVD, and DiDeMo. Our results show that MOF is a generic and efficient method to boost multiple baseline methods, and can achieve a new state-of-the-art performance.

cs.CV

BiC-Net: Learning Efficient Spatio-Temporal Relation for Text-Video Retrieval

The task of text-video retrieval aims to understand the correspondence between language and vision, has gained increasing attention in recent years. Previous studies either adopt off-the-shelf 2D/3D-CNN and then use average/max pooling to directly capture spatial features with aggregated temporal information as global video embeddings, or introduce graph-based models and expert knowledge to learn local spatial-temporal relations. However, the existing methods have two limitations: 1) The global video representations learn video temporal information in a simple average/max pooling manner and do not fully explore the temporal information between every two frames. 2) The graph-based local video representations are handcrafted, it depends heavily on expert knowledge and empirical feedback, which may not be able to effectively mine the higher-level fine-grained visual relations. These limitations result in their inability to distinguish videos with the same visual components but with different relations. To solve this problem, we propose a novel cross-modal retrieval framework, Bi-Branch Complementary Network (BiC-Net), which modifies transformer architecture to effectively bridge text-video modalities in a complementary manner via combining local spatial-temporal relation and global temporal information. Specifically, local video representations are encoded using multiple transformer blocks and additional residual blocks to learn spatio-temporal relation features, calling the module a Spatio-Temporal Residual transformer (SRT). Meanwhile, Global video representations are encoded using a multi-layer transformer block to learn global temporal features. Finally, we align the spatio-temporal relation and global temporal features with the text feature on two embedding spaces for cross-modal text-video retrieval.

cs.CV