SearcharxivSearch

arXiv subjects

Jiahao Hu

Publications and source records attributed to Jiahao Hu.

At least 19 recordsLinked to original sources

Quillen equivalences for truncated multicomplexes

We establish Quillen equivalences between successive truncations of multicomplexes for the model structures that detect weak equivalences simultaneously on prescribed pages of the two associated spectral sequences. In particular, the natural comparison from $4$-multicomplexes to bicomplexes, which is relevant to the comparison between almost complex and complex geometry, is a Quillen equivalence. This answers a question of Cirici, Livernet, and Whitehouse.

math.AT

Can Text-to-Image Models Draw from the Right Frame of Reference?

Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference. For example, ``the left of'' may refer to the viewer's image coordinates or to the intrinsic orientation of an object, leading to different expected layouts. Existing T2I benchmarks reveal important layout failures, yet they rarely isolate whether models can follow a specified frame of reference when it differs from camera view. To mitigate this gap, we introduce FoR-T2I, a benchmark for evaluating this distinction with 1,200 prompt pairs built from controlled spatial layouts. In each pair, the camera-view (Cam) prompt states the target relation in camera view, while the frame-of-reference (FoR) prompt describes the same target placement through an oriented anchor object. Across 22 closed-source and open-source T2I models, mean final accuracy is 41.8\% lower on FoR prompts than on matched Cam prompts; even the best-performing model achieves only 44.3\% FoR accuracy. This suggests that current models struggle more when the same layout is described through an object's orientation rather than directly in image coordinates. We further analyze this gap by relation type and camera view, compare several training-free prompting and feedback-based mitigation strategies, and propose a VLM-gated rewriting approach that selects rewritten prompts using visual feedback, improving average FoR accuracy from 25.0\% to 29.2\% under the same generation budget.

cs.CV

Test-Time Curriculum for Open-Set AIGC Detection

AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to emerge. Although existing methods improve cross-generator generalization through better representations or training data construction, they typically follow a static train-once-and-deploy paradigm and cannot adapt after deployment. In this work, we study open-set AIGC image detection from a test-time adaptation perspective. We propose Test-Time Curriculum (TTC), a simple and model-agnostic framework that adapts a detector on unlabeled test data through curriculum-based self-training. TTC starts from highly reliable pseudo-labeled samples and progressively incorporates harder yet informative cases, while enforcing class-balanced selection to reduce biased updates under generator shift. To further improve pseudo-label quality, we introduce Cross-Scale Pseudo-Label Refinement, which aggregates complementary evidence across multiple resolutions for more reliable adaptation, and applies noisy-or fusion at inference to strengthen final predictions. In addition, we construct AIGCGuard, a new benchmark containing 3,100 representative real images and 124,000 generated images from 40 of the most advanced open-source and proprietary text-to-image models. Extensive experiments on five benchmarks show that TTC substantially improves overall detection performance under diverse unseen-generator shifts, establishing a practical and effective test-time adaptation framework for open-set generated image detection.

cs.CV

Bott-Chern and Aeppli homotopy

This paper introduces Bott-Chern and Aeppli homotopy sets for a fibrant class of bisimplicial sets and establishes their basic properties. In positive bidegrees, Bott-Chern homotopy sets carry natural monoid structures, while Aeppli homotopy sets carry natural group structures. They are related by a loop-space comparison: after a bidegree shift, the Aeppli homotopy groups of X are naturally identified with the Bott-Chern homotopy monoids of the loop space of X. In particular, the Bott-Chern homotopy monoids of loop spaces are groups. To justify our definitions, we show that the Bott-Chern homotopy monoids of a bisimplicial abelian group are naturally isomorphic to the Bott-Chern homology groups of its associated normalized Moore bicomplex. An analogous statement holds for Aeppli homotopy.

math.AT

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-turn interactions while leveraging the strong reasoning capacity of large language models. However, since signals about user needs are often weak, indirect, and can only be disambiguated through multi-turn interaction, existing emotional support methods often struggle to acquire and generalize relevant conversational knowledge efficiently. To bridge this gap, we introduce User-Aware Active Knowledge Acquisition (UKA), a gradient-free active dialogue learning framework that explicitly represents uncertainty about user needs and incorporates active learning into both knowledge acquisition and response selection.We propose a Theory-of-Mind uncertainty estimation mechanism that allows the model to prioritize responses, thereby eliciting more informative user feedback. UKA is capable of efficiently exploring user-aligned conversational knowledge during training while maintaining robustness at test time. Experiments across multiple dialogue benchmarks and model architectures demonstrate that our approach consistently outperforms strong baselines in dialogue quality and user alignment.

cs.CL

Safer Trajectory Planning with CBF-guided Diffusion Model for Unmanned Aerial Vehicles

Safe and agile trajectory planning is essential for autonomous systems, especially during complex aerobatic maneuvers. Motivated by the recent success of diffusion models in generative tasks, this paper introduces AeroTrajGen, a novel framework for diffusion-based trajectory generation that incorporates control barrier function (CBF)-guided sampling during inference, specifically designed for unmanned aerial vehicles (UAVs). The proposed CBF-guided sampling addresses two critical challenges: (1) mitigating the inherent unpredictability and potential safety violations of diffusion models, and (2) reducing reliance on extensively safety-verified training data. During the reverse diffusion process, CBF-based guidance ensures collision-free trajectories by seamlessly integrating safety constraint gradients with the diffusion model's score function. The model features an obstacle-aware diffusion transformer architecture with multi-modal conditioning, including trajectory history, obstacles, maneuver styles, and goal, enabling the generation of smooth, highly agile trajectories across 14 distinct aerobatic maneuvers. Trained on a dataset of 2,000 expert demonstrations, AeroTrajGen is rigorously evaluated in simulation under multi-obstacle environments. Simulation results demonstrate that CBF-guided sampling reduces collision rates by 94.7% compared to unguided diffusion baselines, while preserving trajectory agility and diversity. Our code is open-sourced at https://github.com/RoboticsPolyu/CBF-DMP.

cs.RO

Performance Evaluation of Straw Tubes with Muon Beams at CERN

We present results from two test beam campaigns that investigate the performance of straw tube detectors as potential candidates for an FCC-ee straw tracker. These studies were carried out at CERN using 150 GeV muon beams. Dedicated algorithms were developed to determine both single tube spatial resolution for the primary coordinate in the $r-\phi$ plane and spatial resolution for the secondary coordinate along the tube direction within a straw chamber. Detection efficiency was also evaluated as a function of the extrapolated hit position for each tube. Both datasets showed consistent results for spatial resolutions and efficiency. Our findings will help establish benchmark performance metrics and provide valuable insight for future design, optimization, and construction of straw chambers for high-precision tracking applications.

physics.ins-det

Enhanced Hydrogen Electrolyzer with Integrated Energy Storage to Provide Grid-Forming Services for Off-Grid ReP2H Application

This article proposes an energy storage-enhanced hydrogen electrolyzer (ESEHE) to provide grid-forming (GFM) services for off-grid renewable power to hydrogen (ReP2H) systems. Unlike conventional ReP2H systems that use a centralized energy storage (ES) plant, the proposed topology directly connects batteries to the DC buses of electrolysis rectifiers. A tailored virtual synchronous machine (VSM) control framework enables the electrolyzer to autonomously provide real and reactive power support. A coordinated frequency-splitting energy extraction strategy is designed to exploit both the battery and the electrolysis stack's electrical double-layer (EDL) effect on different timescales, maximizing active power support while mitigating battery and stack degradation. An adaptive equalization control strategy is further developed to balance the battery state of charge (SOC) among multiple ESEHEs operating in parallel, which optimizes energy distribution and extends battery life. Real-time simulations on StarSim validate the proposed topology and control strategies. Techno-economic analysis shows that, compared with conventional off-grid ReP2H systems based on a centralized ES plant, the ESEHE improves overall energy efficiency by 0.23% and reduces the initial total converter investment cost by roughly 6%, mainly due to the elimination of bidirectional AC/DC conversion and its associated losses in the centralized ES plant.

math.OC

Soliton-Assisted Massive Signal Broadcasting via Exceptional Points

Chip-scale all-optical signal broadcasting enables data replication from an optical signal to a large number of wavelength channels, playing a critical role in enabling massive-throughput optical communication and computing systems. The underlying process is four-wave mixing between an optical signal and a multi-wavelength pump source via optical Kerr nonlinearity. To enhance the generally weak nonlinearity, high-quality (Q) microcavities are commonly used to achieve practical efficiency. However, the ultra-narrow linewidths of high Q cavities prohibit achieving massive throughput broadcasting due to Fourier reciprocity. Here, we overcome this challenge by harnessing a parity-time symmetric coupled-cavity system that supports equally spaced exceptional points in the frequency domain. This design seamlessly integrates generation of dissipative Kerr soliton comb source and all-optical signal broadcasting into a unified nonlinear process. As a result, we realize soliton-assisted intracavity massive signal broadcasting with a channel count exceeding 100 over 200 nm wavelength range, resulting in Terabit-per-second aggregated rates. This throughput surpasses the intrinsic microcavity linewidth constraint (~200 MHz) by over three orders of magnitude. We further demonstrate the utility of this approach through an optical convolutional accelerator, highlighting its potential to enable transformative capabilities in photonic computing. Our work establishes a new paradigm for chip-scale photonic processing devices based on non-Hermitian optical design.

physics.optics

OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation Balance

Vision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method's efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal.

cs.AI

Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM

Training Long-Context Large Language Models (LLMs) is challenging, as hybrid training with long-context and short-context data often leads to workload imbalances. Existing works mainly use data packing to alleviate this issue, but fail to consider imbalanced attention computation and wasted communication overhead. This paper proposes Hierarchical Balance Packing (HBP), which designs a novel batch-construction method and training recipe to address those inefficiencies. In particular, the HBP constructs multi-level data packing groups, each optimized with a distinct packing length. It assigns training samples to their optimal groups and configures each group with the most effective settings, including sequential parallelism degree and gradient checkpointing configuration. To effectively utilize multi-level groups of data, we design a dynamic training pipeline specifically tailored to HBP, including curriculum learning, adaptive sequential parallelism, and stable loss. Our extensive experiments demonstrate that our method significantly reduces training time over multiple datasets and open-source models while maintaining strong performance. For the largest DeepSeek-V2 (236B) MoE model, our method speeds up the training by 2.4$\times$ with competitive performance. Codes will be released at https://github.com/ModelTC/HBP.

cs.LG

Integrated Planning and Control on Manifolds: Factor Graph Representation and Toolkit

Model predictive control (MPC) faces significant limitations when applied to systems evolving on nonlinear manifolds, such as robotic attitude dynamics and constrained motion planning, where traditional Euclidean formulations struggle with singularities, over-parameterization, and poor convergence. To overcome these challenges, this paper introduces FactorMPC, a factor-graph based MPC toolkit that unifies system dynamics, constraints, and objectives into a modular, user-friendly, and efficient optimization structure. Our approach natively supports manifold-valued states with Gaussian uncertainties modeled in tangent spaces. By exploiting the sparsity and probabilistic structure of factor graphs, the toolkit achieves real-time performance even for high-dimensional systems with complex constraints. The velocity-extended on-manifold control barrier function (CBF)-based obstacle avoidance factors are designed for safety-critical applications. By bridging graphical models with safety-critical MPC, our work offers a scalable and geometrically consistent framework for integrated planning and control. The simulations and experimental results on the quadrotor demonstrate superior trajectory tracking and obstacle avoidance performance compared to baseline methods. To foster research reproducibility, we have provided open-source implementation offering plug-and-play factors.

cs.RO

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle is the absence of open-source, large-scale video editing datasets based on real-world data, as constructing such datasets is both time-consuming and costly. Moreover, video data requires a significantly larger number of tokens for representation, which substantially increases the training costs for video editing models. Lastly, current video editing models offer limited interactivity, often making it difficult for users to express their editing requirements effectively in a single attempt. To address these challenges, this paper introduces a dataset VIVID-10M and a baseline model VIVID. VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, which comprises 9.7M samples that encompass a wide range of video editing tasks. VIVID is a Versatile and Interactive VIdeo local eDiting model trained on VIVID-10M, which supports entity addition, modification, and deletion. At its core, a keyframe-guided interactive video editing mechanism is proposed, enabling users to iteratively edit keyframes and propagate it to other frames, thereby reducing latency in achieving desired outcomes. Extensive experimental evaluations show that our approach achieves state-of-the-art performance in video local editing, surpassing baseline methods in both automated metrics and user studies. The VIVID-10M dataset are open-sourced at https://kwaivgi.github.io/VIVID/.

cs.CV

Super-efficient optical frequency division referenced to {\mu}Hz Schawlow-Townes-linewidth quantum-noise-limited lasers

Optical frequency division (OFD) implements the conversion of ultra-stable optical frequencies into microwave frequencies through an optical frequency comb flywheel, generating microwave oscillators with record-low phase noise and time jitter. However, conventional OFD systems face significant trade-off between division complexity and noise suppression due to severe thermal noise and technical noise in the optical frequency references. Here, we address this challenge by generating common-cavity bi-color Brillouin lasers as the optical frequency references, which operate at the fundamental quantum noise limit with Schawlow-Townes linewidth on the 10 {\mu}Hz level. Enabled by these ultra-coherent reference lasers, our OFD system uses a dramatically simplified comb divider with an unprecedented small division factor of 10, and generates 10 GHz microwave signal with exceptional phase noise of -65 dBc/Hz at 1Hz, -155 dBc/Hz at 10 kHz, and -172 dBc/Hz at 10 MHz offset. Moreover, to fully harness the spectral purity of the OFD technology, here we implement broadband frequency synthesis directly referenced to the OFD oscillator, covering 5 to 20 GHz with millisecond tuning time. Our work redefines the trade-off between noise suppression and division complexity in OFD, paving the way for compact, high-performance microwave synthesis for next-generation atomic clocks, quantum sensors, and low-noise radar systems.

physics.optics

Further Characterization of the JadePix-3 CMOS Pixel Sensor for the CEPC Vertex Detector: in Dependence of Substrate Reverse Bias

The Circular Electron-Positron Collider (CEPC), a proposed next-generation $e^+e^-$ collider to enable high-precision studies of the Higgs boson and potential new physics, imposes rigorous demands on detector technologies, particularly the vertex detector. JadePix-3 is a prototype Monolithic Active Pixel Sensor (MAPS) designed for the CEPC vertex detector. This paper presents a detailed laboratory-based characterization of the JadePix-3 sensor, focusing on the previously under-explored effects of substrate reverse bias voltage on key performance metrics: charge collection efficiency, average cluster size, and hit efficiency of laser. Systematic testing demonstrated that JadePix-3 operates reliably under reverse bias, exhibiting a reduced input capacitance, an expanded depletion region, enhanced charge collection efficiency, and a lower fake-hit rate. These findings confirm the sensor's potential for high-precision particle tracking and vertexing at the CEPC while offering valuable references for future iterational R\&D of the JadePix series.

physics.ins-det

Obstruction Theory for Bigraded Differential Algebras

We develop an obstruction theory for Hirsch extensions of cbba's with twisted coefficients. This leads to a variety of applications, including a structural theorem for minimal cbba's, a construction of relative minimal models with twisted coefficients, as well as a proof of uniqueness. These results are further employed to study automorphism groups of minimal cbba's and to characterize formality in terms of grading automorphisms.

math.AT

Beyond Tree Models: A Hybrid Model of KAN and gMLP for Large-Scale Financial Tabular Data

Tabular data plays a critical role in real-world financial scenarios. Traditionally, tree models have dominated in handling tabular data. However, financial datasets in the industry often encounter some challenges, such as data heterogeneity, the predominance of numerical features and the large scale of the data, which can range from tens of millions to hundreds of millions of records. These challenges can lead to significant memory and computational issues when using tree-based models. Consequently, there is a growing need for neural network-based solutions that can outperform these models. In this paper, we introduce TKGMLP, an hybrid network for tabular data that combines shallow Kolmogorov Arnold Networks with Gated Multilayer Perceptron. This model leverages the strengths of both architectures to improve performance and scalability. We validate TKGMLP on a real-world credit scoring dataset, where it achieves state-of-the-art results and outperforms current benchmarks. Furthermore, our findings demonstrate that the model continues to improve as the dataset size increases, making it highly scalable. Additionally, we propose a novel feature encoding method for numerical data, specifically designed to address the predominance of numerical features in financial datasets. The integration of this feature encoding method within TKGMLP significantly improves prediction accuracy. This research not only advances table prediction technology but also offers a practical and effective solution for handling large-scale numerical tabular data in various industrial applications.

cs.LG

Characterization of differential K-theory by hexagon diagram

Using a canonical topology on differential K-theory induced from the Frechét space topology on differential forms and the discrete topology on topological K-theory, we prove that differential K-theory is uniquely determined by the character diagram up to a unique natural equivalence, thus giving an affirmative answer to a question asked by Simons and Sullivan in \cite{SS10}. We further deduce rigidity results including that there is a unique way of realizing $\RR/\ZZ$-K-theory as the flat theory, strengthening the results of \cite{BS10}.

math.AT