SearcharxivSearch

arXiv subjects

Yaqi Zhao

Publications and source records attributed to Yaqi Zhao.

At least 19 recordsLinked to original sources

A New Sample of $\sim$ 100 Intermediate-mass Black Holes Reaching $z \approx 1$

We present a systematic search for intermediate-mass black hole (IMBH) active galactic nuclei (AGNs) at $0.5 < z \lesssim 1$ using DESI DR1 spectroscopy.We identify 98 broad-line IMBH AGNs with black hole masses $M_\mathrm{BH}<10^6$ $M_{\odot}$ through quantitative spectral decomposition and broad-H$β$ selection. This sample spans $M_\mathrm{BH}=10^{5.5}-10^{6.0}$ $M_{\odot}$ and Eddington ratios from 1.2 to 9.0 extending systematic IMBH AGN searches to intermediate redshift. Compared with a consistently selected $z<0.6$ IMBH sample, the $0.5<z<1$ sources reach higher broad-H$β$ luminosities and substantially higher Eddington ratios, indicating more extreme accretion states at earlier cosmic times. They also show broader [O III] profiles and stronger blueshifted wing components, with these kinematic differences persisting after matching in black hole mass and Eddington ratio. These results reveal two distinct signatures of evolution among IMBH AGNs at $z<1$: ability of IMBHs to reach increasingly extreme accretion states toward higher redshift, and systematic changes in their ionized-gas kinematics. The former demonstrates that rapid, including super-Eddington, growth of IMBHs can persist to relatively late cosmic times, while the latter may indicate evolution in the ionized-gas environment and associated outflow activity of actively growing IMBHs. Together, these findings provide new constraints on the evolutionary pathways of IMBHs at $z<1$, and show that seed-mass black holes can continue to undergo rapid growth well after the cosmic dawn.

astro-ph.GA

Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels

Prompt learning has gained significant attention as a parameter-efficient approach for adapting large pre-trained vision-language models to downstream tasks. However, when only partial labels are available, its performance is often limited by label ambiguity and insufficient supervisory information. To address this issue, we propose Holistic Optimal Label Selection (HopS), leveraging the generalization ability of pre-trained feature encoders through two complementary strategies. First, we design a local density-based filter that selects the top frequent labels from the nearest neighbors' candidate sets and uses the softmax scores to identify the most plausible label, capturing structural regularities in the feature space. Second, we introduce a global selection objective based on optimal transport that maps the uniform sampling distribution to the candidate label distributions across a batch. By minimizing the expected transport cost, it can determine the most likely label assignments. These two strategies work together to provide robust label selection from both local and global perspectives. Extensive experiments on eight benchmark datasets show that HopS consistently improves performance under partial supervision and outperforms all baselines. Those results highlight the merit of holistic label selection and offer a practical solution for prompt learning in weakly supervised settings.

cs.CV

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning

Vision-language models such as CLIP achieve strong visual-textual alignment, but often suffer from overfitting and limited interpretability when adapted through continuous prompt learning. While discrete prompt optimization improves interpretability, it usually depends on large external models, leading to high computational costs and limited scalability. In this paper, we propose Interpretable Prompt Learning (IPL), a hybrid framework that alternates between discrete semantic token selection and continuous prompt optimization. Specifically, IPL formulates semantic token selection as an approximate submodular optimization problem, encouraging tokens that are both human-understandable and semantically diverse. It further adopts an alternating optimization strategy to integrate discrete token selection with continuous prompt tuning, improving interpretability while preserving adaptability to downstream tasks. Our framework is plug-and-play, allowing seamless integration with existing prompt learning methods. Extensive experiments on multiple benchmarks show that IPL consistently improves both interpretability and accuracy across five representative prompt learning methods, providing an effective and scalable extension to existing frameworks.

cs.CV

Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models

Large language models store not only isolated facts but also rules that support reasoning across symbolic expressions, natural language explanations, and concrete instances. Yet most model editing methods are built for fact-level knowledge, assuming that a target edit can be achieved through a localized intervention. This assumption does not hold for rule-level knowledge, where a single rule must remain consistent across multiple interdependent forms. We investigate this problem through a mechanistic study of rule-level knowledge editing. To support this study, we extend the RuleEdit benchmark from 80 to 200 manually verified rules spanning mathematics and physics. Fine-grained causal tracing reveals a form-specific organization of rule knowledge in transformer layers: formulas and descriptions are concentrated in earlier layers, while instances are more associated with middle layers. These results suggest that rule knowledge is not uniformly localized, and therefore cannot be reliably edited by a single-layer or contiguous-block intervention. Based on this insight, we propose Distributed Multi-Layer Editing (DMLE), which applies a shared early-layer update to formulas and descriptions and a separate middle-layer update to instances. While remaining competitive on standard editing metrics, DMLE achieves substantially stronger rule-level editing performance. On average, it improves instance portability and rule understanding by 13.91 and 50.19 percentage points, respectively, over the strongest baseline across GPT-J-6B, Qwen2.5-7B, Qwen2-7B, and LLaMA-3-8B. The code is available at https://github.com/Pepper66/DMLE.

cs.CL

UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous representation. We empirically demonstrate that reducing channel dimension is significantly more effective than spatial downsampling for both reconstruction and generation. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. Furthermore, we validate that the transfusion architecture surpasses query-based designs in convergence and consistency. Experiments demonstrate that UniCom achieves state-of-the-art generation performance among unified models. Notably, by preserving rich semantic priors, it delivers exceptional controllability in image editing and maintains image consistency even without relying on VAE.

cs.CV

Negativity Percolation in Continuous-Variable Quantum Networks

Quantum networks (QNs) have been predominantly driven by discrete-variable (DV) architectures. Yet, optical platforms naturally generate Gaussian states--the common states of continuous-variable (CV) systems, making CV-based QNs an attractive route toward scalable, chip-integrated quantum computation and communication. To bridge the gap between well-studied DV entanglement percolation theories and their CV counterpart, we introduce a Gaussian-to-Gaussian entanglement distribution scheme that deterministically transports two-mode squeezed vacuum states across large CV networks. Analysis of the scheme's collective behavior using statistical-physics methods reveals a new form of entanglement percolation--negativity percolation theory (NegPT)--characterized by a bounded entanglement measure called the ratio negativity. We discover that NegPT exhibits a mixed-order phase transition, marked simultaneously by both an abrupt change in global entanglement and a long-range correlation between nodes. This distinctive behavior places CV-based QNs in a new universality class, fundamentally distinct from DV systems. Additionally, the abruptness of this transition introduces a critical vulnerability of CV-based QNs: conventional feedback mechanism becomes inherently unstable near the threshold, highlighting practical implications for stabilizing large-scale CV-based QNs. Our results unify statistical models for CV-based entanglement distribution and uncover previously unexplored critical phenomena unique to CV systems, providing valuable insights and guidelines essential for developing robust, feedback-stabilized QNs.

quant-ph

Quasinormal modes of tensor perturbation in Kaluza-Klein black hole for Einstein-Gauss-Bonnet gravity

In Einstein-Gauss-Bonnet gravity, we study the quasi-normal modes (QNMs) of the tensor perturbation for the so-called Maeda-Dadhich black hole which locally has a topology $\mathcal{M}^n \simeq M^4 \times \mathcal{K}^{n-4}$. Our discussion is based on the tensor perturbation equation derived in~\cite{Cao:2021sty}, where the Kodama-Ishibashi gauge invariant formalism for Einstein gravity theory has been generalized to the Einstein-Gauss-Bonnet gravity theory. With the help of characteristic tensors for the constant curvature space $\mathcal{K}^{n-4}$, we investigate the effect of extra dimensions and obtain the scalar equation in four dimensional spacetime, which is quite different from the Klein-Gordon equation. Using the asymptotic iteration method and the numerical integration method with the Kumaresan-Tufts frequency extraction method, we numerically calculate the QNM frequencies. In our setups, characteristic frequencies depend on six distinct factors. They are the spacetime dimension $n$, the Gauss-Bonnet coupling constant $α$, the black hole mass parameter $μ$, the black hole charge parameter $q$, and two ``quantum numbers" $l$, $γ$. Without loss of generality, the impact of each parameter on the characteristic frequencies is investigated while fixing other five parameters. Interestingly, the dimension of compactification part has no significant impact on the lifetime of QNMs.

gr-qc

Demystifying Numerosity in Diffusion Models -- Limitations and Remedies

Numerosity remains a challenge for state-of-the-art text-to-image generation models like FLUX and GPT-4o, which often fail to accurately follow counting instructions in text prompts. In this paper, we aim to study a fundamental yet often overlooked question: Can diffusion models inherently generate the correct number of objects specified by a textual prompt simply by scaling up the dataset and model size? To enable rigorous and reproducible evaluation, we construct a clean synthetic numerosity benchmark comprising two complementary datasets: GrayCount250 for controlled scaling studies, and NaturalCount6 featuring complex naturalistic scenes. Second, we empirically show that the scaling hypothesis does not hold: larger models and datasets alone fail to improve counting accuracy on our benchmark. Our analysis identifies a key reason: diffusion models tend to rely heavily on the noise initialization rather than the explicit numerosity specified in the prompt. We observe that noise priors exhibit biases toward specific object counts. In addition, we propose an effective strategy for controlling numerosity by injecting count-aware layout information into the noise prior. Our method achieves significant gains, improving accuracy on GrayCount250 from 20.0\% to 85.3\% and on NaturalCount6 from 74.8\% to 86.3\%, demonstrating effective generalization across settings.

cs.CV

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual inputs, yet modality alignment remains one of the most challenging aspects. Current MLLMs typically rely on simple adapter architectures and pretraining approaches to bridge vision encoders with large language models (LLM), guided by image-level supervision. We identify this paradigm often leads to suboptimal alignment between modalities, significantly constraining the LLM's ability to properly interpret and reason with visual features particularly for smaller language models. This limitation degrades overall performance-particularly for smaller language models where capacity constraints are more pronounced and adaptation capabilities are limited. To address this fundamental limitation, we propose Supervised Embedding Alignment (SEA), a token-level supervision alignment method that enables more precise visual-text alignment during pretraining. SEA introduces minimal computational overhead while preserving language capabilities and substantially improving cross-modal understanding. Our comprehensive analyses reveal critical insights into the adapter's role in multimodal integration, and extensive experiments demonstrate that SEA consistently improves performance across various model sizes, with smaller models benefiting the most (average performance gain of 7.61% for Gemma-2B). This work establishes a foundation for developing more effective alignment strategies for future multimodal systems.

cs.CV

A nonconvex entanglement monotone determining the characteristic length of entanglement distribution in continuous-variable quantum networks

Quantum networks (QNs) promise to enhance the performance of various quantum technologies in the near future by distributing entangled states over long distances. The first step towards this is to develop novel entanglement measures that are both informative and computationally tractable at large scales. While numerous such entanglement measures exist for discrete-variable (DV) systems, a comprehensive exploration for experimentally preferred continuous-variable (CV) systems is lacking. Here, we introduce a class of CV entanglement measures, among which we identify a nonconvex entanglement monotone -- the ratio negativity, which possesses a simple, scalable form that determines the exponential decay of optimal entanglement swapping on a chain of pure Gaussian states. This characterization opens avenues for leveraging statistical physics tools to analyze swapping-protocol-based CV QNs.

quant-ph

Elasticity of bidisperse attractive particle systems

Bidisperse particle systems are common in both natural and engineered materials, and it is known to influence packing, flow, and stability. However, their direct effect on elastic properties, particularly in systems with attractive interactions, remains poorly understood. Gaining insight into this relationship is important for designing soft particle-based materials with desired mechanical response. In this work, we study how particle size ratio and composition affect the shear modulus of attractive particle systems. Using coarse-grained molecular simulations, we analyze systems composed of two particle sizes at fixed total packing fraction and find that the shear modulus increases systematically with bidispersity. To explain this behavior, we develop two asymptotic models following limiting cases: one where a percolated network of large particles is stiffened by small particles, and another where a small-particle network is modified by embedded large particles. Both models yield closed-form expressions that capture the qualitative trends observed in simulations, including the dependence of shear modulus on size ratio and relative volume fraction. Our results demonstrate that bidispersity can enhance elastic stiffness through microstructural effects, independently of overall density, offering a simple strategy to design particle-based materials with tunable mechanical properties.

cond-mat.soft

Baichuan-Omni-1.5 Technical Report

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve fluent and high-quality interaction across modalities without compromising the capabilities of any modality, we prioritized optimizing three key aspects. First, we establish a comprehensive data cleaning and synthesis pipeline for multimodal data, obtaining about 500B high-quality data (text, audio, and vision). Second, an audio-tokenizer (Baichuan-Audio-Tokenizer) has been designed to capture both semantic and acoustic information from audio, enabling seamless integration and enhanced compatibility with MLLM. Lastly, we designed a multi-stage training strategy that progressively integrates multimodal alignment and multitask fine-tuning, ensuring effective synergy across all modalities. Baichuan-Omni-1.5 leads contemporary models (including GPT4o-mini and MiniCPM-o 2.6) in terms of comprehensive omni-modal capabilities. Notably, it achieves results comparable to leading models such as Qwen2-VL-72B across various multimodal medical benchmarks.

cs.CL

Towards Precise Scaling Laws for Video Diffusion Transformers

Achieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in language models to predict performance, their existence and accurate derivation in visual generation models remain underexplored. In this paper, we systematically analyze scaling laws for video diffusion transformers and confirm their presence. Moreover, we discover that, unlike language models, video diffusion models are more sensitive to learning rate and batch size, two hyperparameters often not precisely modeled. To address this, we propose a new scaling law that predicts optimal hyperparameters for any model size and compute budget. Under these optimal settings, we achieve comparable performance and reduce inference costs by 40.1% compared to conventional scaling methods, within a compute budget of 1e10 TFlops. Furthermore, we establish a more generalized and precise relationship among validation loss, any model size, and compute budget. This enables performance prediction for non-optimal model sizes, which may also be appealed under practical inference cost constraints, achieving a better trade-off.

cs.CV

Baichuan-Omni Technical Report

The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart. In this paper, we introduce Baichuan-omni, the first open-source 7B Multimodal Large Language Model (MLLM) adept at concurrently processing and analyzing modalities of image, video, audio, and text, while delivering an advanced multimodal interactive experience and strong performance. We propose an effective multimodal training schema starting with 7B model and proceeding through two stages of multimodal alignment and multitask fine-tuning across audio, image, video, and text modal. This approach equips the language model with the ability to handle visual and audio data effectively. Demonstrating strong performance across various omni-modal and multimodal benchmarks, we aim for this contribution to serve as a competitive baseline for the open-source community in advancing multimodal understanding and real-time interaction.

cs.AI

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Does seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the large language model (LLM). Specifically, the VE's representation of visual information may not fully align with LLM's cognitive framework, leading to a mismatch where visual features exceed the language model's interpretive range. To address this, we investigate how variations in VE representations influence LVLM comprehension, especially when the LLM faces VE-Unknown data-images whose ambiguous visual representations challenge the VE's interpretive precision. Accordingly, we construct a multi-granularity landmark dataset and systematically examine the impact of VE-Known and VE-Unknown data on interpretive abilities. Our results show that VE-Unknown data limits LVLM's capacity for accurate understanding, while VE-Known data, rich in distinctive features, helps reduce cognitive misalignment. Building on these insights, we propose Entity-Enhanced Cognitive Alignment (EECA), a method that employs multi-granularity supervision to generate visually enriched, well-aligned tokens that not only integrate within the LLM's embedding space but also align with the LLM's cognitive framework. This alignment markedly enhances LVLM performance in landmark recognition. Our findings underscore the challenges posed by VE-Unknown data and highlight the essential role of cognitive alignment in advancing multimodal systems.

cs.CV

Stochastic gravitational wave background from the collisions of dark matter halos

We investigate the effect of the dark matter (DM) halos collisions, namely collisions of galaxies and galaxy clusters, through gravitational bremsstrahlung, on the stochastic gravitational wave background. We first calculate the gravitational wave signal of a single collision event, assuming point masses and linear perturbation theory. Then we proceed to the calculation of the energy spectrum of the collective effect of all dark matter collisions in the Universe. Concerning the DM halo collision rate, we show that it is given by the product of the number density of DM halos, which is calculated by the extended Press-Schechter (EPS) theory, with the collision rate of a single DM halo, which is given by simulation results, with a function of the linear growth rate of matter density through cosmological evolution. Hence, integrating over all mass and distance ranges, we finally extract the spectrum of the stochastic gravitational wave background created by DM halos collisions. As we show, the resulting contribution to the stochastic gravitational wave background is of the order of $h_{c} \approx 10^{-29}$ in the band of $f \approx 10^{-15} Hz$. However, in very low frequency band, it is larger. With current observational sensitivity it cannot be detected.

astro-ph.CO

The effective field theory approach to the strong coupling issue in $f(T)$ gravity

We investigate the scalar perturbations and the possible strong coupling issues of $f(T)$ around a cosmological background, applying the effective field theory (EFT) approach. We revisit the generalized EFT framework of modified teleparallel gravity and apply it by considering both linear and second-order perturbations for $f(T)$ theory. No new scalar mode is present in linear and second-order perturbations in $f(T)$ gravity, which suggests a strong coupling problem. However, based on the ratio of cubic to quadratic Lagrangians, we provide a simple estimation of the strong coupling scale, a result which shows that the strong coupling problem can be avoided at least for some modes. In conclusion, perturbation behaviors that at first appear problematic may not inevitably lead to a strong coupling problem, as long as the relevant scale is comparable with the cutoff scale $M$ of the applicability of the theory.

gr-qc

Quasinormal modes of black holes in f(T) gravity

We calculate the quasinormal modes (QNM) frequencies of a test massless scalar field and an electromagnetic field around static black holes in $f(T)$ gravity. Focusing on quadratic $f(T)$ modifications, which is a good approximation for every realistic $f(T)$ theory, we first extract the spherically symmetric solutions using the perturbative method, imposing two ans$\ddot{\text{a}}$tze for the metric functions, which suitably quantify the deviation from the Schwarzschild solution. Moreover, we extract the effective potential, and then calculate the QNM frequency of the obtained solutions. Firstly, we numerically solve the Schr$\ddot{\text{o}}$dinger-like equation using the discretization method, and we extract the frequency and the time evolution of the dominant mode applying the function fit method. Secondly, we perform a semi-analytical calculation by applying the WKB method with the Pade approximation. We show that the results for $f(T)$ gravity are different compared to General Relativity, and in particular we obtain a different slope and period of the field decay behavior for different model parameter values. Hence, under the light of gravitational-wave observations of increasing accuracy from binary systems, the whole analysis could be used as an additional tool to test General Relativity and examine whether torsional gravitational modifications are possible.

gr-qc