SearcharxivSearch

arXiv subjects

Hong Qi

Publications and source records attributed to Hong Qi.

At least 19 recordsLinked to original sources

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

cs.CV

FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial paradigm for adapting large language models (LLMs) under constrained computational budgets. However, standard PEFT methods often struggle in multi-task fine-tuning settings, where diverse optimization objectives induce task interference and limited parameter budgets lead to representational deficiency. While recent approaches incorporate mixture-of-experts (MoE) to alleviate these issues, they predominantly operate in the spatial domain, which may introduce structural redundancy and parameter overhead. To overcome these limitations, we reformulate adaptation in the spectral domain. Our spectral analysis reveals that different tasks exhibit distinct frequency energy distributions, and that LLM layers display heterogeneous frequency sensitivities. Motivated by these insights, we propose FourierMoE, which integrates the MoE architecture with the inverse discrete Fourier transform (IDFT) for frequency-aware adaptation. Specifically, FourierMoE employs a frequency-adaptive router to dispatch tokens to experts specialized in distinct frequency bands. Each expert learns a set of conjugate-symmetric complex coefficients, preserving complete phase and amplitude information while theoretically guaranteeing lossless IDFT reconstruction into real-valued spatial weights. Extensive evaluations across 28 benchmarks, multiple model architectures, and scales demonstrate that FourierMoE consistently outperforms competitive baselines in both single-task and multi-task settings while using significantly fewer trainable parameters. These results highlight the promise of spectral-domain expert adaptation as an effective and parameter-efficient paradigm for LLM fine-tuning.

cs.LG

Multibanded Reduced Order Quadrature Techniques for Gravitational Wave Inference

Reduced-order quadrature (ROQ) is commonly used to accelerate parameter estimation in gravitational wave astronomy; however, constructing ROQ bases can be computationally costly, particularly for longer-duration signals. We propose a modified construction strategy based on PyROQ that accelerates this process by performing the basis search using multiband waveforms, without compromising the desired likelihood accuracy. We use this altered method to construct a set of ROQs in the sub-solar mass (SSM) range using the IMRPhenomXAS_NRTidalV3 waveform. Compared to PyROQ's standard ROQ method, we find a decrease in basis size of 20% to 30% and observe a decrease in basis construction time by 5 to 20 times, reducing from two weeks to a couple of days. We verify the bases built with this method by injecting simulated gravitational waves into LIGO-Virgo-KAGRA design noise and recovering the parameters, and we find that they preserve the likelihood accuracy and maintain consistent parameter estimation results.

gr-qc

A Realistic Projection for Constraining Neutron Star Equation of State with the LIGO-Virgo-KAGRA Detector Network in the A+ Era

The LIGO-Virgo-KAGRA network in the upcoming A+ era with upgrades of both Advanced LIGO and Advanced Virgo will enable more frequent and precise observations of binary neutron star (BNS) mergers, improving constraints on the neutron star equation of state (EOS). In this study, we applied reduced order quadrature techniques for full parameter estimation of 3,000 simulated gravitational wave signals from BNS mergers at A+ sensitivity following three EOS models: HQC18, SLY230A, and MPA1. We found that tidal deformability tends to be overestimated at higher mass and underestimated at lower mass. We postprocessed the parameter estimation results to present our EOS recovery accuracies, identify biases within EOS constraints and their causes, and quantify the needed corrections.

gr-qc

Quantum Bayesian Inference with Renormalization for Gravitational Waves

Advancements in gravitational-wave interferometers, particularly the next generation, are poised to profoundly impact gravitational wave astronomy and multimessenger astrophysics. A hybrid quantum algorithm is proposed to carry out quantum inference of parameters from compact binary coalescences detected in gravitational-wave interferometers. It performs quantum Bayesian Inference with Renormalization and Downsampling (qBIRD). We choose binary black hole (BBH) mergers from LIGO observatories as the first case to test the algorithm, but its application can be extended to more general instances. The quantum algorithm is able to generate corner plots of relevant parameters such as chirp mass, mass ratio, spins, etc. by inference of simulated gravitational waves with known injected parameter values with zero noise, Gaussian noise and real data, thus recovering an accuracy equivalent to that of classical Markov Chain Monte Carlo inferences. The simulations are performed with sets of 2 and 4 parameters. These results enhance the possibilities to extend our capacity to track signals from coalescences over longer durations and at lower frequencies extending the accuracy and promptness of gravitational wave parameter estimation.

quant-ph

Radiative thermal switch via metamaterials made of vanadium dioxide-coated nanoparticles

In this work, a thermal switch is proposed based on the phase-change material vanadium dioxide (VO2) within the framework of near-field radiative heat transfer (NFRHT). The radiative thermal switch consists of two metamaterials filled with core-shell nanoparticles, with the shell made of VO2. Compared to traditional VO2 slabs, the proposed switch exhibits a more than 2-times increase in the switching ratio, reaching as high as 90.29% with a 100 nm vacuum gap. The improved switching effect is attributed to the capability of the VO2 shell to couple with the core, greatly enhancing heat transfer with the insulating VO2, while blocking the motivation of the core in the metallic state of VO2. As a result, this efficiently enlarges the difference in photonic characteristics between the insulating and metallic states of the structure, thereby improving the ability to rectify the NFRHT. The proposed switch opens pathways for active control of NFRHT and holds practical significance for developing thermal photon-based logic circuits.

physics.app-ph

Splitting of temperature distributions due to dual-channel photon heat exchange in many-body systems

We investigate the radiative heat transfer and spatial distributions of stationary temperatures in periodic many-body systems composed of alternating slabs of two different materials. We show that temperature distributions exhibit an alternating spatial pattern and split into two distinct components, with each component corresponding to one of the two materials. Spatial temperature variations following the periodicity of the structure can be attributed to a dual-channel photon heat exchange through a long-range coupling of electromagnetic modes supported by bodies of the same material. We also analyze the thermal relaxation of the temperatures in the system to verify potential applications in dynamical situations. The results reveal that tunable nonmonotonic temperature variations can be also designed and utilized at a transient mode. The dual-channel mechanism to control temperature distributions proposed in the present work may pave new avenues for prospective applications in nano devices, especially for thermal photon-driven logic circuitry and thermal management.

physics.optics

Multiple magnetoplasmon polaritons of magneto-optical graphene in near-field radiative heat transfer

Graphene, as a two-dimensional magneto-optical material, supports magnetoplasmon polaritons (MPP) when exposed to an applied magnetic field. Recently, MPP of a single-layer graphene has shown an excellent capability in the modulation of near-field radiative heat transfer (NFRHT). In this study, we present a comprehensive theoretical analysis of NFRHT between two multilayered graphene structures, with a particular focus on the multiple MPP effect. We reveal the physical mechanism and evolution law of the multiple MPP, and we demonstrate that the multiple MPP allow one to mediate, enhance, and tune the NFRHT by appropriately engineering the properties of graphene, the number of graphene sheets, the intensity of magnetic fields, as well as the geometric structure of systems. We show that the multiple MPP have a quite significant distinction relative to the single MPP or multiple surface plasmon polaritons (SPPs) in terms of modulating and manipulating NFRHT.

cond-mat.mes-hall

Performance improvement of three-body radiative diode driven by graphene surface plasmon polaritons

As an analogue to electrical diode, a radiative thermal diode allows radiation to transfer more efficiently in one direction than in the opposite direction by operating in a contactless mode. In this study, we demonstrated that, within the framework of three-body photon thermal tunneling, the rectification performance of three-body radiative diode can be greatly improved by bringing graphene into the system. The system is composed of three parallel slabs, with the hot and cold terminals of the diode coated with graphene films, and the intermediate body made of vanadium dioxide (VO2). The rectification factor of the proposed radiative thermal diode reaches 300 % with a 350 nm separation distance between the hot and cold terminals of the diode. With the help of graphene, the rectification performance of the radiative thermal diode can be improved by over 11 times. By analyzing the spectral heat flux and energy transmission coefficients, it was found that the improved performance is primarily attributed to the surface plasmon polaritons (SPPs) of graphene. They excite the modes of insulating VO2 in the forward-biased scenario by forming strongly coupled modes between graphene and VO2, and thus dramatically enhance the heat flux. While, for the reverse-biased scenario, the VO2 is at its metallic state and thus graphene SPPs cannot work by three-body photon thermal tunneling. Furthermore, the improvement was also investigated for different chemical potentials of graphene, and geometric parameters of the three-body system. Our findings demonstrate the feasibility of using thermal-photon-based logical circuits, creating radiation-based communication technology, and implementing thermal management approaches at the nanoscale.

physics.optics

Dynamic Spatial-temporal Hypergraph Convolutional Network for Skeleton-based Action Recognition

Skeleton-based action recognition relies on the extraction of spatial-temporal topological information. Hypergraphs can establish prior unnatural dependencies for the skeleton. However, the existing methods only focus on the construction of spatial topology and ignore the time-point dependence. This paper proposes a dynamic spatial-temporal hypergraph convolutional network (DST-HCN) to capture spatial-temporal information for skeleton-based action recognition. DST-HCN introduces a time-point hypergraph (TPH) to learn relationships at time points. With multiple spatial static hypergraphs and dynamic TPH, our network can learn more complete spatial-temporal features. In addition, we use the high-order information fusion module (HIF) to fuse spatial-temporal information synchronously. Extensive experiments on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets show that our model achieves state-of-the-art, especially compared with hypergraph methods.

cs.CV

Type-supervised sequence labeling based on the heterogeneous star graph for named entity recognition

Named entity recognition is a fundamental task in natural language processing, identifying the span and category of entities in unstructured texts. The traditional sequence labeling methodology ignores the nested entities, i.e. entities included in other entity mentions. Many approaches attempt to address this scenario, most of which rely on complex structures or have high computation complexity. The representation learning of the heterogeneous star graph containing text nodes and type nodes is investigated in this paper. In addition, we revise the graph attention mechanism into a hybrid form to address its unreasonableness in specific topologies. The model performs the type-supervised sequence labeling after updating nodes in the graph. The annotation scheme is an extension of the single-layer sequence labeling and is able to cope with the vast majority of nested entities. Extensive experiments on public NER datasets reveal the effectiveness of our model in extracting both flat and nested entities. The method achieved state-of-the-art performance on both flat and nested datasets. The significant improvement in accuracy reflects the superiority of the multi-layer labeling strategy.

cs.CL

End-to-End Entity Detection with Proposer and Regressor

Named entity recognition is a traditional task in natural language processing. In particular, nested entity recognition receives extensive attention for the widespread existence of the nesting scenario. The latest research migrates the well-established paradigm of set prediction in object detection to cope with entity nesting. However, the manual creation of query vectors, which fail to adapt to the rich semantic information in the context, limits these approaches. An end-to-end entity detection approach with proposer and regressor is presented in this paper to tackle the issues. First, the proposer utilizes the feature pyramid network to generate high-quality entity proposals. Then, the regressor refines the proposals for generating the final prediction. The model adopts encoder-only architecture and thus obtains the advantages of the richness of query semantics, high precision of entity localization, and easiness of model training. Moreover, we introduce the novel spatially modulated attention and progressive refinement for further improvement. Extensive experiments demonstrate that our model achieves advanced performance in flat and nested NER, achieving a new state-of-the-art F1 score of 80.74 on the GENIA dataset and 72.38 on the WeiboNER dataset.

cs.CL

Skeleton-based Action Recognition via Temporal-Channel Aggregation

Skeleton-based action recognition methods are limited by the semantic extraction of spatio-temporal skeletal maps. However, current methods have difficulty in effectively combining features from both temporal and spatial graph dimensions and tend to be thick on one side and thin on the other. In this paper, we propose a Temporal-Channel Aggregation Graph Convolutional Networks (TCA-GCN) to learn spatial and temporal topologies dynamically and efficiently aggregate topological features in different temporal and channel dimensions for skeleton-based action recognition. We use the Temporal Aggregation module to learn temporal dimensional features and the Channel Aggregation module to efficiently combine spatial dynamic channel-wise topological features with temporal dynamic topological features. In addition, we extract multi-scale skeletal features on temporal modeling and fuse them with an attention mechanism. Extensive experiments show that our model results outperform state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.

cs.CV

Spirit Distillation: A Model Compression Method with Multi-domain Knowledge Transfer

Recent applications pose requirements of both cross-domain knowledge transfer and model compression to machine learning models due to insufficient training data and limited computational resources. In this paper, we propose a new knowledge distillation model, named Spirit Distillation (SD), which is a model compression method with multi-domain knowledge transfer. The compact student network mimics out a representation equivalent to the front part of the teacher network, through which the general knowledge can be transferred from the source domain (teacher) to the target domain (student). To further improve the robustness of the student, we extend SD to Enhanced Spirit Distillation (ESD) in exploiting a more comprehensive knowledge by introducing the proximity domain which is similar to the target domain for feature extraction. Results demonstrate that our method can boost mIOU and high-precision accuracy by 1.4% and 8.2% respectively with 78.2% segmentation variance, and can gain a precise compact network with only 41.8% FLOPs.

cs.CV

Spirit Distillation: Precise Real-time Semantic Segmentation of Road Scenes with Insufficient Data

Semantic segmentation of road scenes is one of the key technologies for realizing autonomous driving scene perception, and the effectiveness of deep Convolutional Neural Networks(CNNs) for this task has been demonstrated. State-of-art CNNs for semantic segmentation suffer from excessive computations as well as large-scale training data requirement. Inspired by the ideas of Fine-tuning-based Transfer Learning (FTT) and feature-based knowledge distillation, we propose a new knowledge distillation method for cross-domain knowledge transference and efficient data-insufficient network training, named Spirit Distillation(SD), which allow the student network to mimic the teacher network to extract general features, so that a compact and accurate student network can be trained for real-time semantic segmentation of road scenes. Then, in order to further alleviate the trouble of insufficient data and improve the robustness of the student, an Enhanced Spirit Distillation (ESD) method is proposed, which commits to exploit a more comprehensive general features extraction capability by considering images from both the target and the proximity domains as input. To our knowledge, this paper is a pioneering work on the application of knowledge distillation to few-shot learning. Persuasive experiments conducted on Cityscapes semantic segmentation with the prior knowledge transferred from COCO2017 and KITTI demonstrate that our methods can train a better student network (mIOU and high-precision accuracy boost by 1.4% and 8.2% respectively, with 78.2% segmentation variance) with only 41.8% FLOPs (see Fig. 1).

cs.CV

Testing the Black Hole No-hair Theorem with Galactic Center Stellar Orbits

Theoretical investigations have provided proof-of-principle calculations suggesting measurements of stellar or pulsar orbits near the Galactic Center could strongly constrain the properties of the Galactic Center black hole, local matter, and even the theory of gravity itself. In this work, we develop both a Markov chain Monte Carlo and an analytic model (Fisher matrix) to understand what properties are well-constrained and why. We conclude that existing astrometric measurements cannot constrain the spin of the Galactic Center black hole. Extrapolating to the precision and cadence of future experiments, we anticipate that the black hole spin can be measured with the known star S2. Our calculations show that we can measure the dimensionless black hole spin to a precision of $\sim$0.1 with weekly measurements of the orbit of S2 for 40 years using the GRAVITY telescope's best resolution at the Galactic Center. An analytic expression is derived for the measurement uncertainty of the black hole spin using the Fisher matrix in terms of observation strategy, star's orbital parameters, and instrument resolution. From it, we conclude that highly eccentric orbits can provide better constraints on the spin and that an orbit with a higher eccentricity is more favorable even when the orbital period is longer. We also apply it to S62, S4711, and S4714 and show whether they can constrain the black hole spin sooner than S2. If in addition, future measurements include the discovery of a new, tighter stellar orbit, then future data could conceivably enable tests of strong field gravity, by directly measuring the black hole quadrupole moment. Our simulations show that with a stellar orbit similar to that of S2 but at one-fifth the distance to the Galactic Center and GRAVITY's resolution limits, we can start to test the no-hair theorem with 20 years of weekly orbital measurements.

astro-ph.GA

Activation Map Adaptation for Effective Knowledge Distillation

Model compression becomes a recent trend due to the requirement of deploying neural networks on embedded and mobile devices. Hence, both accuracy and efficiency are of critical importance. To explore a balance between them, a knowledge distillation strategy is proposed for general visual representation learning. It utilizes our well-designed activation map adaptive module to replace some blocks of the teacher network, exploring the most appropriate supervisory features adaptively during the training process. Using the teacher's hidden layer output to prompt the student network to train so as to transfer effective semantic information.To verify the effectiveness of our strategy, this paper applied our method to cifar-10 dataset. Results demonstrate that the method can boost the accuracy of the student network by 0.6% with 6.5% loss reduction, and significantly improve its training speed.

cs.CV

PyROQ: a Python-based Reduced Order Quadrature Building Code for Fast Gravitational Wave Inference

The next generation of gravitational-wave observatories will reach low frequency limits on the orders of a few Hz, thus enabling the detection of gravitational wave signals of very long duration. The run time of standard parameter estimation techniques with these long waveforms can be months or even years, making it impractical with existing Bayesian inference pipelines. Reduced order modeling and reduced order quadrature integration rule have recently been exploited as promising techniques that can greatly reduce parameter estimation computational costs. We describe a Python-based reduced order quadrature building code, PyROQ, which builds the reduced order quadrature data needed to accelerate parameter estimation of gravitational waves. We present the first bases for the IMRPhenomXPHM waveform model of binary-black-hole coalescences, including subdominant harmonic modes and precessing spins effects. Furthermore, the code infrastructure makes it directly applicable to the gravitational wave inference for space-borne detectors such as the Laser Interferometer Space Antenna (LISA).

gr-qc