Searcharxiv⌕ Search

arXiv subjects

Yu Hu

Publications and source records attributed to Yu Hu.

At least 91 records · Page 5Linked to original sources

SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding

Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND

cs.CL↗

Learning Linear Polytree Structural Equation Models

We are interested in the problem of learning the directed acyclic graph (DAG) when data are generated from a linear structural equation model (SEM) and the causal structure can be characterized by a polytree. Under the Gaussian polytree models, we study sufficient conditions on the sample sizes for the well-known Chow-Liu algorithm to exactly recover both the skeleton and the equivalence class of the polytree, which is uniquely represented by a CPDAG. On the other hand, necessary conditions on the required sample sizes for both skeleton and CPDAG recovery are also derived in terms of information-theoretic lower bounds, which match the respective sufficient conditions and thereby give a sharp characterization of the difficulty of these tasks. We also consider the problem of inverse correlation matrix estimation under the linear polytree models, and establish the estimation error bound in terms of the dimension and the total number of v-structures. We also consider an extension of group linear polytree models, in which each node represents a group of variables. Our theoretical findings are illustrated by comprehensive numerical simulations, and experiments on benchmark data also demonstrate the robustness of polytree learning when the true graphical structures can only be approximated by polytrees.

stat.ML↗

Energy in critical collapse

We study the energy issue in critical collapse of a spherically symmetric scalar field. It is found that in critical collapse, the contribution from the material energy is greater than that from the gravitational energy. The quantity $m/r$ plays an important role in identifying the formation of apparent horizon in gravitational collapse, where $m$ is the Misner-Sharp mass and $r$ the areal radius. We observe that in critical collapse, the maximum value of $m/r$ fluctuates between $2/15$ and $4/15$. This denotes a large gap between critical collapse and black hole formation for which the criterion is $m/r=1/2$.

gr-qc↗

New results on the dynamics of critical collapse

We study the dynamics of the critical collapse of a spherically symmetric scalar field. Approximate analytic expressions for the metric functions and matter field in the large-radius region are obtained. In the central region, owing to the boundary conditions, the equation of motion for the scalar field is reduced to the flat-spacetime form.

gr-qc↗

Disturbance Rejection-Guarded Learning for Vibration Suppression of Two-Inertia Systems

Model uncertainty presents significant challenges in vibration suppression of multi-inertia systems, as these systems often rely on inaccurate nominal mathematical models due to system identification errors or unmodeled dynamics. An observer, such as an extended state observer (ESO), can estimate the discrepancy between the inaccurate nominal model and the true model, thus improving control performance via disturbance rejection. The conventional observer design is memoryless in the sense that once its estimated disturbance is obtained and sent to the controller, the datum is discarded. In this research, we propose a seamless integration of ESO and machine learning. On one hand, the machine learning model attempts to model the disturbance. With the assistance of prior information about the disturbance, the observer is expected to achieve faster convergence in disturbance estimation. On the other hand, machine learning benefits from an additional assurance layer provided by the ESO, as any imperfections in the machine learning model can be compensated for by the ESO. We validated the effectiveness of this novel learning-for-control paradigm through simulation and physical tests on two-inertial motion control systems used for vibration studies.

eess.SY↗

A Non-parametric Reconstruction of the Hubble Parameter $H(z)$ Based on Radial Basis Function Neural Networks

Accurately measuring the Hubble parameter is vital for understanding the expansion history and properties of the universe. In this paper, we propose a new method that supplements the covariance between redshift pairs to improve the reconstruction of the Hubble parameter using the OHD dataset. Our approach utilizes a cosmological model-independent radial basis function neural network (RBFNN) to describe the Hubble parameter as a function of redshift effectively. Our experiments show that this method results in a reconstructed Hubble parameter of $H_0 = 67.1\pm9.7~\mathrm{km~s^{-1}~Mpc^{-1}}$ , which is more noise-resistant and fits better with the $Λ$CDM model at high redshifts. Providing the covariance between redshift pairs in subsequent observations will significantly improve the reliability and accuracy of Hubble parametric data reconstruction. Future applications of this method could help overcome the limitations of previous methods and lead to new advances in our understanding of the universe.

astro-ph.CO↗

Measurements of $p-Λ$ and $d-Λ$ correlations in 3 GeV Au+Au collisions at STAR

Heavy-ion collisions provide a unique opportunity to explore nucleon-hyperon (N-Y) interactions through two-particle correlations. The $p-Λ$ and $d-Λ$ correlations shed light on both N-Y two-body and N-N-Y three-body interactions, which is crucial for understanding neutron star properties. We present the high precision measurement of $p-Λ$ and the first measurement of $d-Λ$ correlation with $\sqrt{s_{_{\rm NN}}}=$ 3 GeV Au+Au collisions at STAR. Using the Lednicky-Lyuboshitz formalism, we characterized emission source size, the scattering length ($f_0$), and the effective range ($d_0$) of $p-Λ$ and $d-Λ$ interactions. Using the $f_0$ and $d_0$ extracted from two spin states in $d-Λ$ correlation, the parameters from the doublet state indicate the hypertriton binding energy is consistent with the current average of world measurements.

nucl-ex↗

Technical note: ShinyAnimalCV: open-source cloud-based web application for object detection, segmentation, and three-dimensional visualization of animals using computer vision

Computer vision (CV), a non-intrusive and cost-effective technology, has furthered the development of precision livestock farming by enabling optimized decision-making through timely and individualized animal care. The availability of affordable two- and three-dimensional camera sensors, combined with various machine learning and deep learning algorithms, has provided a valuable opportunity to improve livestock production systems. However, despite the availability of various CV tools in the public domain, applying these tools to animal data can be challenging, often requiring users to have programming and data analysis skills, as well as access to computing resources. Moreover, the rapid expansion of precision livestock farming is creating a growing need to educate and train animal science students in CV. This presents educators with the challenge of efficiently demonstrating the complex algorithms involved in CV. Thus, the objective of this study was to develop ShinyAnimalCV, an open-source cloud-based web application. This application provides a user-friendly interface for performing CV tasks, including object segmentation, detection, three-dimensional surface visualization, and extraction of two- and three-dimensional morphological features. Nine pre-trained CV models using top-view animal data are included in the application. ShinyAnimalCV has been deployed online using cloud computing platforms. The source code of ShinyAnimalCV is available on GitHub, along with detailed documentation on training CV models using custom data and deploying ShinyAnimalCV locally to allow users to fully leverage the capabilities of the application. ShinyAnimalCV can contribute to CV research and teaching in the animal science community.

cs.CV↗

Parameter estimation for Einstein-dilaton-Gauss-Bonnet gravity with ringdown signals

Future space-based gravitational-wave detectors will detect gravitational waves with high sensitivity in the millihertz frequency band, which provides more opportunities to test theories of gravity than ground-based ones. The study of quasinormal modes (QNMs) and their application to testing gravity theories have been an important aspect in the field of gravitational physics. In this study, we investigate the capability of future space-based gravitational wave detectors such as LISA, TaiJi, and TianQin to constrain the dimensionless deviating parameter for Einstein-dilaton-Gauss-Bonnet (EdGB) gravity with ringdown signals from the merger of binary black holes. The ringdown signal is modeled by the two strongest QNMs in EdGB gravity. Taking into account time-delay interferometry, we calculate the signal-to-noise ratio (SNR) of different space-based detectors for ringdown signals to analyze their capabilities. The Fisher information matrix is employed to analyze the accuracy of parameter estimation, with particular focus on the dimensionless deviating parameter for EdGB gravity. The impact of the parameters of gravitational wave sources on the estimation accuracy of the dimensionless deviating parameter has also been studied. We find that the constraint ability of EdGB gravity is limited because the uncertainty of the dimensionless deviating parameter increases with the decrease of the dimensionless deviating parameter. LISA and TaiJi has more advantages to constrain the dimensionless deviating parameter to a more accurate level for the massive black hole, while TianQin is more suitable for less massive black holes. Bayesian inference method is used to perform parameter estimation on simulated data, which verifies the reliability of the conclusion.

gr-qc↗

A General Model-Based Extended State Observer with Built-In Zero Dynamics

A general model-based extended state observer (GMB-ESO) is proposed for single-input single-output linear time-invariant systems with a given state space model, where the total disturbance, a lump sum of model uncertainties and external disturbances, is defined as an extended state in the same manner as in the original formulation of ESO. The conditions for the existence of such an observer, however, are shown for the first time as 1) the original plant is observable; and 2) there is no invariant zero between the plant output and the total disturbance. Then, the finite-step convergence and error characteristics of GMB-ESO are shown by exploiting its inherent connection to the well-known unknown input observer (UIO). Furthermore, it is shown that, with the relative degree of the plant greater than one and the observer eigenvalues all placed at the origin, GMB-ESO produces the identical disturbance estimation as that of UIO. Finally, an improved GMB-ESO with built-in zero dynamics is proposed for those plants with zero dynamics, which is a problem that has not been addressed in all existing ESO designs.

eess.SY↗

FeatureBooster: Boosting Feature Descriptors with a Lightweight Neural Network

We introduce a lightweight network to improve descriptors of keypoints within the same image. The network takes the original descriptors and the geometric properties of keypoints as the input, and uses an MLP-based self-boosting stage and a Transformer-based cross-boosting stage to enhance the descriptors. The boosted descriptors can be either real-valued or binary ones. We use the proposed network to boost both hand-crafted (ORB, SIFT) and the state-of-the-art learning-based descriptors (SuperPoint, ALIKE) and evaluate them on image matching, visual localization, and structure-from-motion tasks. The results show that our method significantly improves the performance of each task, particularly in challenging cases such as large illumination changes or repetitive patterns. Our method requires only 3.2ms on desktop GPU and 27ms on embedded GPU to process 2000 features, which is fast enough to be applied to a practical system. The code and trained weights are publicly available at github.com/SJTU-ViSYS/FeatureBooster.

cs.CV↗

Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback

Large language models (LLMs), such as ChatGPT, are able to generate human-like, fluent responses for many downstream tasks, e.g., task-oriented dialog and question answering. However, applying LLMs to real-world, mission-critical applications remains challenging mainly due to their tendency to generate hallucinations and their inability to use external knowledge. This paper proposes a LLM-Augmenter system, which augments a black-box LLM with a set of plug-and-play modules. Our system makes the LLM generate responses grounded in external knowledge, e.g., stored in task-specific databases. It also iteratively revises LLM prompts to improve model responses using feedback generated by utility functions, e.g., the factuality score of a LLM-generated response. The effectiveness of LLM-Augmenter is empirically validated on two types of scenarios, task-oriented dialog and open-domain question answering. LLM-Augmenter significantly reduces ChatGPT's hallucinations without sacrificing the fluency and informativeness of its responses. We make the source code and models publicly available.

cs.CL↗

Testing the polarization of gravitational wave background with LISA-TianQin network

While general relativity predicts only two tensor modes for gravitational wave polarization, general metric theories of gravity allows up to four additional modes, including two vector and two scalar modes. Observing the polarization modes of gravitational waves could provide a direct test of the modified gravity. The stochastic gravitational wave background (SGWB), which may be detected by space-based laser-interferometric detectors at design sensitivity, will provide an opportunity to directly measure alternative polarization. In this paper, we investigate the performance of the LISA-TianQin network for detecting alternative polarizations of stochastic backgrounds, and propose a method to separate different polarization modes. First, we generalize the small antenna approximation to compute the overlap reduction functions for SGWB with arbitrary polarization, which is suitable for any time-delay interferometry combination. Then we analyze the detection capability of LISA-TianQin for SGWB with different polarizations. Based on the LISA-TianQin orbital characteristics, we propose a method to distinguish different polarization modes from their mixed data. Compared with ground-based detectors, the LISA-TianQin network is more capable of resolving polarizations of SGWB. In particular, the LISA-TianQin network has the potential to resolve two scalar modes that ground-based detectors cannot.

gr-qc↗

Few-shot 3D LiDAR Semantic Segmentation for Autonomous Driving

In autonomous driving, the novel objects and lack of annotations challenge the traditional 3D LiDAR semantic segmentation based on deep learning. Few-shot learning is a feasible way to solve these issues. However, currently few-shot semantic segmentation methods focus on camera data, and most of them only predict the novel classes without considering the base classes. This setting cannot be directly applied to autonomous driving due to safety concerns. Thus, we propose a few-shot 3D LiDAR semantic segmentation method that predicts both novel classes and base classes simultaneously. Our method tries to solve the background ambiguity problem in generalized few-shot semantic segmentation. We first review the original cross-entropy and knowledge distillation losses, then propose a new loss function that incorporates the background information to achieve 3D LiDAR few-shot semantic segmentation. Extensive experiments on SemanticKITTI demonstrate the effectiveness of our method.

cs.RO↗

PA&DA: Jointly Sampling PAth and DAta for Consistent NAS

Based on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during training. And we further find that large gradient variance occurs during supernet training, which degrades the supernet ranking consistency. To mitigate this issue, we propose to explicitly minimize the gradient variance of the supernet training by jointly optimizing the sampling distributions of PAth and DAta (PA&DA). We theoretically derive the relationship between the gradient variance and the sampling distributions, and reveal that the optimal sampling probability is proportional to the normalized gradient norm of path and training data. Hence, we use the normalized gradient norm as the importance indicator for path and training data, and adopt an importance sampling strategy for the supernet training. Our method only requires negligible computation cost for optimizing the sampling distributions of path and data, but achieves lower gradient variance during supernet training and better generalization performance for the supernet, resulting in a more consistent NAS. We conduct comprehensive comparisons with other improved approaches in various search spaces. Results show that our method surpasses others with more reliable ranking performance and higher accuracy of searched architectures, showing the effectiveness of our method. Code is available at https://github.com/ShunLu91/PA-DA.

cs.CV↗

Uniform tensor clustering by jointly exploring sample affinities of various orders

Conventional clustering methods based on pairwise affinity usually suffer from the concentration effect while processing huge dimensional features yet low sample sizes data, resulting in inaccuracy to encode the sample proximity and suboptimal performance in clustering. To address this issue, we propose a unified tensor clustering method (UTC) that characterizes sample proximity using multiple samples' affinity, thereby supplementing rich spatial sample distributions to boost clustering. Specifically, we find that the triadic tensor affinity can be constructed via the Khari-Rao product of two affinity matrices. Furthermore, our early work shows that the fourth-order tensor affinity is defined by the Kronecker product. Therefore, we utilize arithmetical products, Khatri-Rao and Kronecker products, to mathematically integrate different orders of affinity into a unified tensor clustering framework. Thus, the UTC jointly learns a joint low-dimensional embedding to combine various orders. Finally, a numerical scheme is designed to solve the problem. Experiments on synthetic datasets and real-world datasets demonstrate that 1) the usage of high-order tensor affinity could provide a supplementary characterization of sample proximity to the popular affinity matrix; 2) the proposed method of UTC is affirmed to enhance clustering by exploiting different order affinities when processing high-dimensional data.

cs.LG↗

Search for the Chiral Effect using isobar collisions and BES-II data from STAR

In these proceedings we discuss the recent precision measurements of charge separation difference between Ru+Ru and Zr+Zr collisions at $\sqrt{s_{\rm NN}}=200$ GeV by STAR collaboration. The measurements indicate that the magnitude of the difference in the charge separation attributable to the magnetic fields between the two systems is smaller than previously expected. We also present charge separation measurements on the Chiral Magnetic Effect search from the RHIC BES-II experiment using the Event Plane Detectors (EPD) from Au+Au collisions at $\sqrt{s_{\rm NN}} =$ 27 GeV.

nucl-ex↗

Univoque bases of real numbers: simply normal bases, irregular bases and multiple rationals

Given a positive integer $M$ and a real number $x\in(0,1]$, we call $q\in(1,M+1]$ a univoque simply normal base of $x$ if there exists a unique simply normal sequence $(d_i)\in\{0,1,\ldots,M\}^\mathbb N$ such that $x=\sum_{i=1}^\infty d_i q^{-i}$. Similarly, a base $q\in(1,M+1]$ is called a univoque irregular base of $x$ if there exists a unique sequence $(d_i)\in\{0,1,\ldots, M\}^\mathbb N$ such that $x=\sum_{i=1}^\infty d_i q^{-i}$ and the sequence $(d_i)$ has no digit frequency. Let $\mathcal U_{SN}(x)$ and $\mathcal U_{I_r}(x)$ be the sets of univoque simply normal bases and univoque irregular bases of $x$, respectively. In this paper we show that for any $x\in(0,1]$ both $\mathcal U_{SN}(x)$ and $\mathcal U_{I_r}(x)$ have full Hausdorff dimension. Furthermore, given finitely many rationals $x_1, x_2, \ldots, x_n\in(0,1]$ so that each $x_i$ has a finite expansion in base $M+1$, we show that there exists a full Hausdorff dimensional set of $q\in(1,M+1]$ such that each $x_i$ has a unique expansion in base $q$.

math.DS↗