SearcharxivSearch

arXiv subjects

Heng Hao

Publications and source records attributed to Heng Hao.

At least 19 recordsLinked to original sources

Query-Aware Flow Diffusion for Graph-Based RAG with Retrieval Guarantees

Graph-based Retrieval-Augmented Generation (RAG) systems leverage interconnected knowledge structures to capture complex relationships that flat retrieval struggles with, enabling multi-hop reasoning. Yet most existing graph-based methods suffer from (i) heuristic designs lacking theoretical guarantees for subgraph quality or relevance and/or (ii) the use of static exploration strategies that ignore the query's holistic meaning, retrieving neighborhoods or communities regardless of intent. We propose Query-Aware Flow Diffusion RAG (QAFD-RAG), a training-free framework that dynamically adapts graph traversal to each query's holistic semantics. The central innovation is query-aware traversal: during graph exploration, edges are dynamically weighted by how well their endpoints align with the query's embedding, guiding flow along semantically relevant paths while avoiding structurally connected but irrelevant regions. These query-specific reasoning subgraphs enable the first statistical guarantees for query-aware graph retrieval, showing that QAFD-RAG recovers relevant subgraphs with high probability under mild signal-to-noise conditions. The algorithm converges exponentially fast, with complexity scaling with the retrieved subgraph size rather than the full graph. Experiments on question answering and text-to-SQL tasks demonstrate consistent improvements over state-of-the-art graph-based RAG methods.

cs.IR

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning

Text-to-SQL models allow users to interact with a database more easily by generating executable SQL statements from natural-language questions. Despite recent successes on simpler databases and questions, current Text-to-SQL methods still suffer from low execution accuracy on industry-scale databases and complex questions involving domain-specific business logic. We present \emph{PaVeRL-SQL}, a framework that combines \emph{Partial-Match Rewards} and \emph{Verbal Reinforcement Learning} to drive self-improvement in reasoning language models (RLMs) for Text-to-SQL. To handle practical use cases, we adopt two pipelines: (1) a newly designed in-context learning framework with group self-evaluation (verbal-RL), using capable open- and closed-source large language models (LLMs) as backbones; and (2) a chain-of-thought (CoT) RL pipeline with a small backbone model (OmniSQL-7B) trained with a specially designed reward function and two-stage RL. These pipelines achieve state-of-the-art (SOTA) results on popular Text-to-SQL benchmarks -- Spider, Spider 2.0, and BIRD. For the industrial-level Spider2.0-SQLite benchmark, the verbal-RL pipeline achieves an execution accuracy 7.4\% higher than SOTA, and the CoT pipeline is 1.4\% higher. RL training with mixed SQL dialects yields strong, threefold gains, particularly for dialects with limited training data. Overall, \emph{PaVeRL-SQL} delivers reliable, SOTA Text-to-SQL under realistic industrial constraints. The code is available at https://github.com/PaVeRL-SQL/PaVeRL-SQL.

cs.AI

Bayesian Active Learning for Semantic Segmentation

Fully supervised training of semantic segmentation models is costly and challenging because each pixel within an image needs to be labeled. Therefore, the sparse pixel-level annotation methods have been introduced to train models with a subset of pixels within each image. We introduce a Bayesian active learning framework based on sparse pixel-level annotation that utilizes a pixel-level Bayesian uncertainty measure based on Balanced Entropy (BalEnt) [84]. BalEnt captures the information between the models' predicted marginalized probability distribution and the pixel labels. BalEnt has linear scalability with a closed analytical form and can be calculated independently per pixel without relational computations with other pixels. We train our proposed active learning framework for Cityscapes, Camvid, ADE20K and VOC2012 benchmark datasets and show that it reaches supervised levels of mIoU using only a fraction of labeled pixels while outperforming the previous state-of-the-art active learning models with a large margin.

cs.CV

Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation

Assessing response quality to instructions in language models is vital but challenging due to the complexity of human language across different contexts. This complexity often results in ambiguous or inconsistent interpretations, making accurate assessment difficult. To address this issue, we propose a novel Uncertainty-aware Reward Model (URM) that introduces a robust uncertainty estimation for the quality of paired responses based on Bayesian approximation. Trained with preference datasets, our uncertainty-enabled proxy not only scores rewards for responses but also evaluates their inherent uncertainty. Empirical results demonstrate significant benefits of incorporating the proposed proxy into language model training. Our method boosts the instruction following capability of language models by refining data curation for training and improving policy optimization objectives, thereby surpassing existing methods by a large margin on benchmarks such as Vicuna and MT-bench. These findings highlight that our proposed approach substantially advances language model training and paves a new way of harnessing uncertainty within language models.

cs.CL

Self-Supervised Contrastive Representation Learning for 3D Mesh Segmentation

3D deep learning is a growing field of interest due to the vast amount of information stored in 3D formats. Triangular meshes are an efficient representation for irregular, non-uniform 3D objects. However, meshes are often challenging to annotate due to their high geometrical complexity. Specifically, creating segmentation masks for meshes is tedious and time-consuming. Therefore, it is desirable to train segmentation networks with limited-labeled data. Self-supervised learning (SSL), a form of unsupervised representation learning, is a growing alternative to fully-supervised learning which can decrease the burden of supervision for training. We propose SSL-MeshCNN, a self-supervised contrastive learning method for pre-training CNNs for mesh segmentation. We take inspiration from traditional contrastive learning frameworks to design a novel contrastive learning algorithm specifically for meshes. Our preliminary experiments show promising results in reducing the heavy labeled data requirement needed for mesh segmentation by at least 33%.

cs.CV

Highly Efficient Representation and Active Learning Framework and Its Application to Imbalanced Medical Image Classification

We propose a highly data-efficient active learning framework for image classification. Our novel framework combines: (1) unsupervised representation learning of a Convolutional Neural Network and (2) the Gaussian Process (GP) method, in sequence to achieve highly data and label efficient classifications. Moreover, both elements are less sensitive to the prevalent and challenging class imbalance issue, thanks to the (1) feature learned without labels and (2) the Bayesian nature of GP. The GP-provided uncertainty estimates enable active learning by ranking samples based on the uncertainty and selectively labeling samples showing higher uncertainty. We apply this novel combination to the severely imbalanced case of COVID-19 chest X-ray classification and the Nerthus colonoscopy classification. We demonstrate that only . 10% of the labeled data is needed to reach the accuracy from training all available labels. We also applied our model architecture and proposed framework to a broader class of datasets with expected success.

cs.CV

PatchNet: Unsupervised Object Discovery based on Patch Embedding

We demonstrate that frequently appearing objects can be discovered by training randomly sampled patches from a small number of images (100 to 200) by self-supervision. Key to this approach is the pattern space, a latent space of patterns that represents all possible sub-images of the given image data. The distance structure in the pattern space captures the co-occurrence of patterns due to the frequent objects. The pattern space embedding is learned by minimizing the contrastive loss between randomly generated adjacent patches. To prevent the embedding from learning the background, we modulate the contrastive loss by color-based object saliency and background dissimilarity. The learned distance structure serves as object memory, and the frequent objects are simply discovered by clustering the pattern vectors from the random patches sampled for inference. Our image representation based on image patches naturally handles the position and scale invariance property that is crucial to multi-object discovery. The method has been proven surprisingly effective, and successfully applied to finding multiple human faces and bodies from natural images.

cs.CV

Inter-comparison of Radio-Loudness Criteria for Type 1 AGNs in the XMM-COSMOS Survey

Limited studies have been performed on the radio-loud fraction in X-ray selected type 1 AGN samples. The consistency between various radio-loudness definitions also needs to be checked. We measure the radio-loudness of the 407 type 1 AGNs in the XMM-COSMOS quasar sample using nine criteria from the literature (six defined in the rest-frame and three defined in the observed frame): $R_L=\log(L_{5GHz}/L_B)$, $q_{24}=\log(L_{24μm}/L_{1.4GHz})$, $R_{uv}=\log(L_{5GHz}/L_{2500Å})$, $R_{i}=\log(L_{1.4GHz}/L_i)$, $R_X=\log(νL_ν(5GHz)/L_X)$, $P_{5GHz}=\log(P_{5GHz}(W/Hz/Sr))$, $R_{L,obs}=\log(f_{1.4GHz}/f_B)$ (observed frame), $R_{i,obs}=\log(f_{1.4GHz}/f_i)$ (observed frame), and $q_{24, obs}=\log(f_{24μm}/f_{1.4GHz})$ (observed frame). Using any single criterion defined in the rest-frame, we find a low radio-loud fraction of $\lesssim 5\%$ in the XMM-COSMOS type 1 AGN sample, except for $R_{uv}$. Requiring that any two criteria agree reduces the radio-loud fraction to $\lesssim 2\%$ for about 3/4 of the cases. The low radio-loud fraction cannot be simply explained by the contribution of the host galaxy luminosity and reddening. The $P_{5GHz}=\log(P_{5GHz}(W/Hz/Sr))$ gives the smallest radio-loud fraction. Two of the three radio-loud fractions from the criteria defined in the observed frame without k-correction ($R_{L,obs}$ and $R_{i,obs}$) are much larger than the radio-loud fractions from other criteria.

astro-ph.GA

Spectral Energy Distributions of Type 1 AGN in XMM-COSMOS Survey II - Shape Evolution

The mid-infrared to ultraviolet (0.1 -- 10 $μm$) spectral energy distribution (SED) shapes of 407 X-ray-selected radio-quiet type 1 AGN in the wide-field ``Cosmic Evolution Survey" (COSMOS) have been studied for signs of evolution. For a sub-sample of 200 radio-quiet quasars with black hole mass estimates and host galaxy corrections, we studied their mean SEDs as a function of a broad range of redshift, bolometric luminosity, black hole mass and Eddington ratio, and compared them with the Elvis et al. (1994, E94) type 1 AGN mean SED. We found that the mean SEDs in each bin are closely similar to each other, showing no statistical significant evidence of dependence on any of the analyzed parameters. We also measured the SED dispersion as a function of these four parameters, and found no significant dependencies. The dispersion of the XMM-COSMOS SEDs is generally larger than E94 SED dispersion in the ultraviolet, which might be due to the broader ``window function'' for COSMOS quasars, and their X-ray based selection.

astro-ph.GA

A Quasar-Galaxy Mixing Diagram: Quasar Spectral Energy Distribution Shapes in the Optical to Near-Infrared

We define a quasar-galaxy mixing diagram using the slopes of their spectral energy distributions (SEDs) from 1μm to 3000Å and from 1μm to 3μm in the rest frame. The mixing diagram can easily distinguish among quasar-dominated, galaxy-dominated and reddening-dominated SED shapes. By studying the position of the 413 XMM selected Type 1 AGN in the wide-field "Cosmic Evolution Survey" (COSMOS) in the mixing diagram, we find that a combination of the Elvis et al. (1994, hereafter E94) quasar SED with various contributions from galaxy emission and some dust reddening is remarkably effective in describing the SED shape from 0.3-3μm for large ranges of redshift, luminosity, black hole mass and Eddington ratio of type 1 AGN. In particular, the location in the mixing diagram of the highest luminosity AGN is very close (within 1σ) to that of the E94 SED. The mixing diagram can also be used to estimate the host galaxy fraction and reddening in quasar. We also show examples of some outliers which might be AGN in different evolutionary stages compared to the majority of AGN in the quasar-host galaxy co-evolution cycle.

astro-ph.GA

The Distribution of AGN Covering Factors

We review our knowledge of the most basic properties of the AGN obscuring region - its location, scale, symmetry, and mean covering factor - and discuss new evidence on the distribution of covering factors in a sample of ~9000 quasars with WISE, UKIDSS, and SDSS photometry. The obscuring regions of AGN may be in some ways more complex than we thought - multi-scale, not symmetric, chaotic - and in some ways simpler - with no dependence on luminosity, and a covering factor distribution that may be determined by the simplest of considerations - e.g. random misalignments.

astro-ph.CO

Hot-Dust-Poor Quasars in Mid-Infrared and Optically Selected Samples

We show that the Hot-Dust-Poor (HDP) quasars, originally found in the X-ray selected XMM-COSMOS type 1 AGN sample, are just as common in two samples selected at optical/infrared wavelengths: the Richards et al. Spitzer/SDSS sample ($8.7%\pm2.2%$), and the PG-quasar dominated sample of Elvis et al. ($9.5%\pm 5.0%$). The properties of the HDP quasars in these two samples are consistent with the XMM-COSMOS sample, except that, at the $99% (\sim 2.5σ)$ significance, a larger proportion of the HDP quasars in the Spitzer/SDSS sample have weak host galaxy contributions, probably due to the selection criteria used. Either the host-dust is destroyed (dynamically or by radiation), or is offset from the central black hole due to recoiling. Alternatively, the universality of HDP quasars in samples with different selection methods and the continuous distribution of dust covering factor in type 1 AGNs, suggest that the range of SEDs could be related to the range of tilts in warped fueling disks, as in the model of Lawrence and Elvis (2010), with HDP quasars having relatively small warps.

astro-ph.CO

Accretion Rate and the Physical Nature of Unobscured Active Galaxies

We show how accretion rate governs the physical properties of a sample of unobscured broad-line, narrow-line, and lineless active galactic nuclei (AGNs). We avoid the systematic errors plaguing previous studies of AGN accretion rate by using accurate accretion luminosities (L_int) from well-sampled multiwavelength SEDs from the Cosmic Evolution Survey (COSMOS), and accurate black hole masses derived from virial scaling relations (for broad-line AGNs) or host-AGN relations (for narrow-line and lineless AGNs). In general, broad emission lines are present only at the highest accretion rates (L_int/L_Edd > 0.01), and these rapidly accreting AGNs are observed as broad-line AGNs or possibly as obscured narrow-line AGNs. Narrow-line and lineless AGNs at lower specific accretion rates (L_int/L_Edd < 0.01) are unobscured and yet lack a broad line region. The disappearance of the broad emission lines is caused by an expanding radiatively inefficient accretion flow (RIAF) at the inner radius of the accretion disk. The presence of the RIAF also drives L_int/L_Edd < 10^-2 narrow-line and lineless AGNs to 10 times higher ratios of radio to optical/UV emission than L_int/L_Edd > 0.01 broad-line AGNs, since the unbound nature of the RIAF means it is easier to form a radio outflow. The IR torus signature also tends to become weaker or disappear from L_int/L_Edd < 0.01 AGNs, although there may be additional mid-IR synchrotron emission associated with the RIAF. Together these results suggest that specific accretion rate is an important physical "axis" of AGN unification, described by a simple model.

astro-ph.CO

Hot-Dust-Poor Type 1 Active Galactic Nuclei in the COSMOS Survey

We report a sizable class of type 1 active galactic nuclei (AGNs) with unusually weak near-infrared (1-3μm) emission in the XMM-COSMOS type 1 AGN sample. The fraction of these "hot-dust-poor" AGNs increases with redshift from 6% at lowredshift (z < 2) to 20% at moderate high redshift (2 < z < 3.5). There is no clear trend of the fraction with other parameters: bolometric luminosity, Eddington ratio, black hole mass, and X-ray luminosity. The 3μm emission relative to the 1μm emission is a factor of 2-4 smaller than the typical Elvis et al. AGN spectral energy distribution (SED), which indicates a "torus" covering factor of 2%-29%, a factor of 3-40 smaller than required by unified models. The weak hot dust emission seems to expose an extension of the accretion disk continuum in some of the source SEDs. We estimate the outer edge of their accretion disks to lie at $(0.3-2.0) /times 10^4$ Schwarzschild radii, ~10-23 times the gravitational stability radii. Formation scenarios for these sources are discussed.

astro-ph.CO

Observational Limits on Type 1 AGN Accretion Rate in COSMOS

We present black hole masses and accretion rates for 182 Type 1 AGN in COSMOS. We estimate masses using the scaling relations for the broad Hb, MgII, and CIV emission lines in the redshift ranges 0.16<z<0.88, 1<z<2.4, and 2.7<z<4.9. We estimate the accretion rate using an Eddington ratio L_I/L_Edd estimated from optical and X-ray data. We find that very few Type 1 AGN accrete below L_I/L_Edd ~ 0.01, despite simulations of synthetic spectra which show that the survey is sensitive to such Type 1 AGN. At lower accretion rates the BLR may become obscured, diluted or nonexistent. We find evidence that Type 1 AGN at higher accretion rates have higher optical luminosities, as more of their emission comes from the cool (optical) accretion disk with respect to shorter wavelengths. We measure a larger range in accretion rate than previous works, suggesting that COSMOS is more efficient at finding low accretion rate Type 1 AGN. However the measured range in accretion rate is still comparable to the intrinsic scatter from the scaling relations, suggesting that Type 1 AGN accrete at a narrow range of Eddington ratio, with L_I/L_Edd ~ 0.1.

astro-ph.CO

The Chandra COSMOS Survey, I: Overview and Point Source Catalog

The Chandra COSMOS Survey (C-COSMOS) is a large, 1.8 Ms, Chandra} program that has imaged the central 0.5 sq.deg of the COSMOS field (centered at 10h, +02deg) with an effective exposure of ~160ksec, and an outer 0.4sq.deg. area with an effective exposure of ~80ksec. The limiting source detection depths are 1.9e-16 erg cm(-2) s(-1) in the Soft (0.5-2 keV) band, 7.3e(-16) erg cm^-2 s^-1 in the Hard (2-10 keV) band, and 5.7e(-16) erg cm(-2) s(-1) in the Full (0.5-10 keV) band. Here we describe the strategy, design and execution of the C-COSMOS survey, and present the catalog of 1761 point sources detected at a probability of being spurious of <2e(-5) (1655 in the Full, 1340 in the Soft, and 1017 in the Hard bands). By using a grid of 36 heavily (~50%) overlapping pointing positions with the ACIS-I imager, a remarkably uniform (to 12%) exposure across the inner 0.5 sq.deg field was obtained, leading to a sharply defined lower flux limit. The widely different PSFs obtained in each exposure at each point in the field required a novel source detection method, because of the overlapping tiling strategy, which is described in a companion paper. (Puccetti et al. Paper II). This method produced reliable sources down to a 7-12 counts, as verified by the resulting logN-logS curve, with sub-arcsecond positions, enabling optical and infrared identifications of virtually all sources, as reported in a second companion paper (Civano et al. Paper III). The full catalog is described here in detail, and is available on-line.

astro-ph.CO

The COSMOS AGN Spectroscopic Survey I: XMM Counterparts

We present optical spectroscopy for an X-ray and optical flux-limited sample of 677 XMM-Newton selected targets covering the 2 deg^2 COSMOS field, with a yield of 485 high-confidence redshifts. The majority of the spectra were obtained over three seasons (2005-2007) with the IMACS instrument on the Magellan (Baade) telescope. We also include in the sample previously published Sloan Digital Sky Survey spectra and supplemental observations with MMT/Hectospec. We detail the observations and classification analyses. The survey is 90% complete to flux limits of f_{0.5-10 keV}>8 x 10^-16 erg cm^-2 s^-1 and i_AB+<22, where over 90% of targets have high-confidence redshifts. Making simple corrections for incompleteness due to redshift and spectral type allows for a description of the complete population to $i_AB+<23. The corrected sample includes 57% broad emission line (Type 1, unobscured) AGN at 0.13 3 x 10^42 erg s^-1) to z<1, of both optically obscured and unobscured types. We find statistically significant evidence that the obscured to unobscured AGN ratio at z<1 increases with redshift and decreases with luminosity.

astro-ph

Detecting a physical difference between the CDM halos in simulation and in nature

Numerical simulation is an important tool to help us understand the process of structure formation in the universe. However many simulation results of cold dark matter (CDM) halos on small scale are inconsistent with observations: the central density profile is too cuspy and there are too many substructures. Here we point out that these two problems may be connected with a hitherto unrecognized bias in the simulation halos. Although CDM halos in nature and in simulation are both virialized systems of collisionless CDM particles, gravitational encounter cannot be neglected in the simulation halos because they contain much less particles. We demonstrate this by two numerical experiments, showing that there is a difference on the microcosmic scale between the natural and simulation halos. The simulation halo is more akin to globular clusters where gravitational encounter is known to lead to such drastic phenomena as core collapse. And such artificial core collapse process appears to link the two problems together in the bottom-up scenario of structure formation in the $Λ$CDM universe. The discovery of this bias also has implications on the applicability of the Jeans Theorem in Galactic Dynamics.

astro-ph