SearcharxivSearch

arXiv subjects

Kun Dong

Publications and source records attributed to Kun Dong.

15 recordsLinked to original sources

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in visual grounding. Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM architecture. Hi-GAR complements this representation with a geometry-based reward for Group Relative Policy Optimization (GRPO), using box overlap and coordinate accuracy at multiple scales. Controlled comparisons under matched training conditions show that Hi-Token improves localization throughout the evaluated IoU range. Hi-GAR further reduces low-overlap predictions and is used only during training. Experiments on three VLM backbones and the RefCOCO family show consistent gains across models and benchmarks. Hi-R1 achieves higher values than strong specialist baselines on most reported metrics. Analyses of token frequency, digit boundaries, object scale, and IoU distributions explain the effects of coordinate representation and reward training. The results show that structured coordinate generation provides an effective approach to generative visual grounding. Project page: https://xyzzzh.github.io/Hi-Token/

cs.CV

A Recursive Hybrid Tetrahedron Method for Brillouin-zone Integration

A recursive extension of the hybrid tetrahedron method for Brillouin-zone integration is proposed, allowing iterative tetrahedron refinement and significantly reducing the error from the linear tetrahedron method. The Brillouin-zone integral is expressed as a weighted sum on the initial grid, with integral weights collected recursively from the finest grid. Our method is capable of simultaneously handling multiple singularities in the integrand and thus may provide practical solutions to various Brillouin-zone integral tasks encountered in realistic calculations, including the computation of response and spectral function with superior sampling convergence. We demonstrate its effectiveness through numerical calculations of the density response functions of two model Hamiltonians and one real material system, the face-centered cubic cobalt.

cond-mat.mtrl-sci

ReWiTe: Realistic Wide-angle and Telephoto Dual Camera Fusion Dataset via Beam Splitter Camera Rig

The fusion of images from dual camera systems featuring a wide-angle and a telephoto camera has become a hotspot problem recently. By integrating simultaneously captured wide-angle and telephoto images from these systems, the resulting fused image achieves a wide field of view (FOV) coupled with high-definition quality. Existing approaches are mostly deep learning methods, and predominantly rely on supervised learning, where the training dataset plays a pivotal role. However, current datasets typically adopt a data synthesis approach generate input pairs of wide-angle and telephoto images alongside ground-truth images. Notably, the wide-angle inputs are synthesized rather than captured using real wide-angle cameras, and the ground-truth image is captured by wide-angle camera whose quality is substantially lower than that of input telephoto images captured by telephoto cameras. To address these limitations, we introduce a novel hardware setup utilizing a beam splitter to simultaneously capture three images, i.e. input pairs and ground-truth images, from two authentic cellphones equipped with wide-angle and telephoto dual cameras. Specifically, the wide-angle and telephoto images captured by cellphone 2 serve as the input pair, while the telephoto image captured by cellphone 1, which is calibrated to match the optical path of the wide-angle image from cellphone 2, serves as the ground-truth image, maintaining quality on par with the input telephoto image. Experiments validate the efficacy of our newly introduced dataset, named ReWiTe, significantly enhances the performance of various existing methods for real-world wide-angle and telephoto dual image fusion tasks.

cs.CV

FoodSAM: Any Food Segmentation

In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, called FoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. We release our code at https://github.com/jamesjg/FoodSAM.

cs.CV

Relay Hindsight Experience Replay: Self-Guided Continual Reinforcement Learning for Sequential Object Manipulation Tasks with Sparse Rewards

Exploration with sparse rewards remains a challenging research problem in reinforcement learning (RL). Especially for sequential object manipulation tasks, the RL agent always receives negative rewards until completing all sub-tasks, which results in low exploration efficiency. To solve these tasks efficiently, we propose a novel self-guided continual RL framework, RelayHER (RHER). RHER first decomposes a sequential task into new sub-tasks with increasing complexity and ensures that the simplest sub-task can be learned quickly by utilizing Hindsight Experience Replay (HER). Secondly, we design a multi-goal & multi-task network to learn these sub-tasks simultaneously. Finally, we propose a Self-Guided Exploration Strategy (SGES). With SGES, the learned sub-task policy will guide the agent to the states that are helpful to learn more complex sub-task with HER. By this self-guided exploration and relay policy learning, RHER can solve these sequential tasks efficiently stage by stage. The experimental results show that RHER significantly outperforms vanilla-HER in sample-efficiency on five singleobject and five complex multi-object manipulation tasks (e.g., Push, Insert, ObstaclePush, Stack, TStack, etc.). The proposed RHER has also been applied to learn a contact-rich push task on a physical robot from scratch, and the success rate reached 10/10 with only 250 episodes.

cs.RO

On-the-Fly Rectification for Robust Large-Vocabulary Topic Inference

Across many data domains, co-occurrence statistics about the joint appearance of objects are powerfully informative. By transforming unsupervised learning problems into decompositions of co-occurrence statistics, spectral algorithms provide transparent and efficient algorithms for posterior inference such as latent topic analysis and community detection. As object vocabularies grow, however, it becomes rapidly more expensive to store and run inference algorithms on co-occurrence statistics. Rectifying co-occurrence, the key process to uphold model assumptions, becomes increasingly more vital in the presence of rare terms, but current techniques cannot scale to large vocabularies. We propose novel methods that simultaneously compress and rectify co-occurrence statistics, scaling gracefully with the size of vocabulary and the dimension of latent space. We also present new algorithms learning latent variables from the compressed statistics, and verify that our methods perform comparably to previous approaches on both textual and non-textual data.

cs.CL

Anti-electrostatic hydrogen bonding between anions of ionic liquids: A density functional theory study

Hydrogen bonds (HBs) play a crucial role in the physicochemical properties of ionic liquids (ILs). At present, HBs between cations and anions (Ca-An) or between cations (Ca-Ca) in ILs have been reported extensively. Here, we provided DFT evidences for the exists of HBs between anions (An-An) in the IL 1-(2-hydroxyethyl)-3-methylimidazolium 4-(2-hydroxyethyl)imidazolide [HEMIm][HEIm]. The thermodynamics stabilities of anionic, cationic, and H2O dimers together with ionic pairs were studied by potential energy scans. The results show that the cation-anion pair is the most stable one, while the HB in anionic dimer possesses similar thermodynamics stability to the water dimer. The further geometric, spectral and electronic structure analyses demonstrate that the inter-anionic HB meets the general theoretical criteria of traditional HBs. The strength order of four HBs in complexes is cation-anion pair > H2O dimer = cationic dimer > anionic dimer. The energy decomposition analysis indicates that induction and dispersion interactions are the crucial factors to overcome strong Coulomb repulsions, forming inter-anionic HBs. Lastly, the presence of inter-anionic HBs in ionic cluster has been confirmed by a global minimum search for a system containing two ionic pairs. Even though hydroxyl-functionalized cations are more likely to form HBs with anions, there still have inter-anionic HBs between hydroxyl groups in the low-lying structures. Our studies broaden the understanding of HBs in ionic liquids and support the recently proposed concept of anti-electrostatic HBs.

physics.chem-ph

Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty

Efficient and effective learning is one of the ultimate goals of the deep reinforcement learning (DRL), although the compromise has been made in most of the time, especially for the application of robot manipulations. Learning is always expensive for robot manipulation tasks and the learning effectiveness could be affected by the system uncertainty. In order to solve above challenges, in this study, we proposed a simple but powerful reward shaping method, namely Dense2Sparse. It combines the advantage of fast convergence of dense reward and the noise isolation of the sparse reward, to achieve a balance between learning efficiency and effectiveness, which makes it suitable for robot manipulation tasks. We evaluated our Dense2Sparse method with a series of ablation experiments using the state representation model with system uncertainty. The experiment results show that the Dense2Sparse method obtained higher expected reward compared with the ones using standalone dense reward or sparse reward, and it also has a superior tolerance of system uncertainty.

cs.LG

Network Density of States

Spectral analysis connects graph structure to the eigenvalues and eigenvectors of associated matrices. Much of spectral graph theory descends directly from spectral geometry, the study of differentiable manifolds through the spectra of associated differential operators. But the translation from spectral geometry to spectral graph theory has largely focused on results involving only a few extreme eigenvalues and their associated eigenvalues. Unlike in geometry, the study of graphs through the overall distribution of eigenvalues - the spectral density - is largely limited to simple random graph models. The interior of the spectrum of real-world graphs remains largely unexplored, difficult to compute and to interpret. In this paper, we delve into the heart of spectral densities of real-world graphs. We borrow tools developed in condensed matter physics, and add novel adaptations to handle the spectral signatures of common graph motifs. The resulting methods are highly efficient, as we illustrate by computing spectral densities for graphs with over a billion edges on a single compute node. Beyond providing visually compelling fingerprints of graphs, we show how the estimation of spectral densities facilitates the computation of many common centrality measures, and use spectral densities to estimate meaningful information about graph structure that cannot be inferred from the extremal eigenpairs alone.

cs.SI

Scaling Gaussian Process Regression with Derivatives

Gaussian processes (GPs) with derivatives are useful in many applications, including Bayesian optimization, implicit surface reconstruction, and terrain reconstruction. Fitting a GP to function values and derivatives at $n$ points in $d$ dimensions requires linear solves and log determinants with an ${n(d+1) \times n(d+1)}$ positive definite matrix -- leading to prohibitive $\mathcal{O}(n^3d^3)$ computations for standard direct methods. We propose iterative solvers using fast $\mathcal{O}(nd)$ matrix-vector multiplications (MVMs), together with pivoted Cholesky preconditioning that cuts the iterations to convergence by several orders of magnitude, allowing for fast kernel learning and prediction. Our approaches, together with dimensionality reduction, enables Bayesian optimization with derivatives to scale to high-dimensional problems and large evaluation budgets.

cs.LG

Scalable Log Determinants for Gaussian Process Kernel Learning

For applications as varied as Bayesian neural networks, determinantal point processes, elliptical graphical models, and kernel learning for Gaussian processes (GPs), one must compute a log determinant of an $n \times n$ positive definite matrix, and its derivatives - leading to prohibitive $\mathcal{O}(n^3)$ computations. We propose novel $\mathcal{O}(n)$ approaches to estimating these quantities from only fast matrix vector multiplications (MVMs). These stochastic approximations are based on Chebyshev, Lanczos, and surrogate models, and converge quickly even for kernel matrices that have challenging spectra. We leverage these approximations to develop a scalable Gaussian process approach to kernel learning. We find that Lanczos is generally superior to Chebyshev for kernel learning, and that a surrogate approach can be highly efficient and accurate with popular kernels.

stat.ML

Interpolative Separable Density Fitting through Centroidal Voronoi Tessellation With Applications to Hybrid Functional Electronic Structure Calculations

The recently developed interpolative separable density fitting (ISDF) decomposition is a powerful way for compressing the redundant information in the set of orbital pairs, and has been used to accelerate quantum chemistry calculations in a number of contexts. The key ingredient of the ISDF decomposition is to select a set of non-uniform grid points, so that the values of the orbital pairs evaluated at such grid points can be used to accurately interpolate those evaluated at all grid points. The set of non-uniform grid points, called the interpolation points, can be automatically selected by a QR factorization with column pivoting (QRCP) procedure. This is the computationally most expensive step in the construction of the ISDF decomposition. In this work, we propose a new approach to find the interpolation points based on the centroidal Voronoi tessellation (CVT) method, which offers a much less expensive alternative to the QRCP procedure when ISDF is used in the context of hybrid functional electronic structure calculations. The CVT method only uses information from the electron density, and can be efficiently implemented using a K-Means algorithm. We find that this new method achieves comparable accuracy to the ISDF-QRCP method, at a cost that is negligible in the overall hybrid functional calculations. For instance, for a system containing $1000$ silicon atoms simulated using the HSE06 hybrid functional on $2000$ computational cores, the cost of QRCP-based method for finding the interpolation points costs $434.2$ seconds, while the CVT procedure only takes $3.2$ seconds. We also find that the ISDF-CVT method also enhances the smoothness of the potential energy surface in the context of \emph{ab initio} molecular dynamics (AIMD) simulations with hybrid functionals.

physics.comp-ph

Study of the Spin-weighted Spheroidal Wave Equation in the Case of s=3/2

In this paper, we use the means of super-symmetric quantum mechanics to study of the Spin-weighted Spheroidal Wave in the case of s=3/2. We obtain some interesting results: the first-five terms of the super-potential, the general form of the super-potential. The ground eigen-function and eigenvalue of the equation are also given. According these results, we make use of the shape invariance property to compute the exited eigenvalues and eigen-functions. These results help us to understand the Spin-weighted Spheroidal Wave and show that it is integral.

quant-ph

Study of the Spin-weighted Spheroidal Equation in the Case of s=1

We present series study of using the method of super-symmetric quantum mechanics(SUSYQM) solving the spin-weighted spheroidal wave equation. In this paper, we obtain the first four terms of super-potential of the spin-weighted spheroidal wave equation in the case of s=1. These results may help summary the general form for the n-th term of the super-potential, which is proved correct by means of induction. We finally compute the ground eigenvalues and ground eigenfunction. All the results may be of significative for studies of electromagnetic radiation processes near rotating black holes and compute radiation reaction in curved space-time.

math-ph

The Spin-weighted Spheroidal Wave functions in the Case of s=1/2

The spin-weighted spheroidal equations in the case s=1/2 is thoroughly studied in the paper by means of the perturbation method in supersymmetry quantum mechanics. The first-five terms of the super-potential in the series of the parameter beta are given. The general form of the nth term of the superpotential is also obtained, which could derived from the previous terms W_{k}, k<n. From the results, it is easy to give the ground eigenfunction of the equation. Furthermore, the shape-invariance property is investigated in the series form of the parameter beta and is proven kept in this series form for the equations. This nice property guarantee one could obtain the excited eigenfunctions in the series form from the ground eigenfunctions by the method in supersymmetry quantum mechanics. This shows the perturbation method method in supersymmetry quantum mechanics could solve the spin-weight spheroidal wave equations completely in the series form of the small parameter beta.

math-ph