SearcharxivSearch

arXiv subjects

Shuo Xin

Publications and source records attributed to Shuo Xin.

17 recordsLinked to original sources

AInsteinBench: Benchmarking Coding Agents on Scientific Repositories

We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research.

cs.SE

Relativistic scalar dark matter drag forces on a black hole binary

Dark matter around black holes can induce drag forces through dynamical friction and accretion, potentially affecting the orbital evolution and gravitational wave emission of binary systems. While dynamical friction from scalar field dark matter has been studied in the relativistic regime for single black holes, the case of a binary black hole (BBH) has remained unexplored. As a first step, we present a series of two-dimensional general-relativistic simulations of BBH in a wind tunnel for an asymptotically homogeneous scalar field background. We extract the drag forces, torque, mass and charge accretion acting on the binary, and analyze their dependence on the binary separation, velocity and the scalar field parameters. We find that the binary's drag is not a simple superposition of two isolated black holes; the presence of a companion modifies the gravitational wake and yields significant nonlinearities. This additional force and torque can (in principle) modify the inspiral and induce a dephasing of the gravitational wave signal.

gr-qc

A3FR: Agile 3D Gaussian Splatting with Incremental Gaze Tracked Foveated Rendering in Virtual Reality

Virtual reality (VR) significantly transforms immersive digital interfaces, greatly enhancing education, professional practices, and entertainment by increasing user engagement and opening up new possibilities in various industries. Among its numerous applications, image rendering is crucial. Nevertheless, rendering methodologies like 3D Gaussian Splatting impose high computational demands, driven predominantly by user expectations for superior visual quality. This results in notable processing delays for real-time image rendering, which greatly affects the user experience. Additionally, VR devices such as head-mounted displays (HMDs) are intricately linked to human visual behavior, leveraging knowledge from perception and cognition to improve user experience. These insights have spurred the development of foveated rendering, a technique that dynamically adjusts rendering resolution based on the user's gaze direction. The resultant solution, known as gaze-tracked foveated rendering, significantly reduces the computational burden of the rendering process. Although gaze-tracked foveated rendering can reduce rendering costs, the computational overhead of the gaze tracking process itself can sometimes outweigh the rendering savings, leading to increased processing latency. To address this issue, we propose an efficient rendering framework called~\textit{A3FR}, designed to minimize the latency of gaze-tracked foveated rendering via the parallelization of gaze tracking and foveated rendering processes. For the rendering algorithm, we utilize 3D Gaussian Splatting, a state-of-the-art neural rendering technique. Evaluation results demonstrate that A3FR can reduce end-to-end rendering latency by up to $2\times$ while maintaining visual quality.

cs.GR

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today's video proliferation era. Multi-modal video summarization that accomodates user input has become a research hotspot. However, current multi-modal video summarization methods suffer from two limitations. First, existing methods inadequately fuse information from different modalities and cannot effectively utilize modality-unique features. Second, most multi-modal methods focus on video and text modalities, neglecting the audio modality, despite the fact that audio information can be very useful in certain types of videos. In this paper we propose CFSum, a transformer-based multi-modal video summarization framework with coarse-fine fusion. CFSum exploits video, text, and audio modal features as input, and incorporates a two-stage transformer-based feature fusion framework to fully utilize modality-unique information. In the first stage, multi-modal features are fused simultaneously to perform initial coarse-grained feature fusion, then, in the second stage, video and audio features are explicitly attended with the text representation yielding more fine-grained information interaction. The CFSum architecture gives equal importance to each modality, ensuring that each modal feature interacts deeply with the other modalities. Our extensive comparative experiments against prior methods and ablation studies on various datasets confirm the effectiveness and superiority of CFSum.

cs.CV

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

We present OmniVLM, a sub-billion-parameter vision-language model for efficient on-device inference. OmniVLM introduces a token compression mechanism that reduces visual token sequence length from 729 to 81 tokens, significantly reducing computational overhead while preserving visual-semantic fidelity. Through a multi-stage training pipeline of pretraining, supervised fine-tuning, and minimal-edit Direct Preference Optimization (DPO), OmniVLM matches the performance of larger models. On multiple benchmarks including ScienceQA, POPE, and MMMU, OmniVLM outperforms existing baselines like nanoLLAVA within a 968M-parameter footprint. Empirical results on the same laptop demonstrate 9.1x faster time-to-first-token (0.75s vs 6.82s) and 1.5x higher decoding speed (29.41 vs 19.20 tokens/s) compared to nanoLLAVA, enabling efficient deployment on edge devices. The model weights can be accessed on huggingface: https://huggingface.co/NexaAIDev/OmniVLM-968M, and the inference examples can be find in Appendix B.

cs.CV

Visual Object Tracking across Diverse Data Modalities: A Review

Visual Object Tracking (VOT) is an attractive and significant research area in computer vision, which aims to recognize and track specific targets in video sequences where the target objects are arbitrary and class-agnostic. The VOT technology could be applied in various scenarios, processing data of diverse modalities such as RGB, thermal infrared and point cloud. Besides, since no one sensor could handle all the dynamic and varying environments, multi-modal VOT is also investigated. This paper presents a comprehensive survey of the recent progress of both single-modal and multi-modal VOT, especially the deep learning methods. Specifically, we first review three types of mainstream single-modal VOT, including RGB, thermal infrared and point cloud tracking. In particular, we conclude four widely-used single-modal frameworks, abstracting their schemas and categorizing the existing inheritors. Then we summarize four kinds of multi-modal VOT, including RGB-Depth, RGB-Thermal, RGB-LiDAR and RGB-Language. Moreover, the comparison results in plenty of VOT benchmarks of the discussed modalities are presented. Finally, we provide recommendations and insightful observations, inspiring the future development of this fast-growing literature.

cs.CV

A Two-Loop Four-Point Form Factor at Function Level

Recently, the maximally-helicity-violating four-point form factor for the chiral stress-energy tensor in planar $\mathcal{N}=4$ super Yang-Mills was computed to three loops at the level of the symbol associated with multiple polylogarithms. It exhibits {\it antipodal self-duality}, or invariance under the combined action of a kinematic map and reversing the ordering of letters in the symbol. Here we lift the two-loop form factor from symbol level to function level. We provide an iterated representation of the function's derivatives (coproducts). In order to do so, we find a three-parameter limit of the five-parameter phase space where the symbol's letters are all rational. We also use function-level information about dihedral symmetries and the soft, collinear, and factorization limits, as well as limits governed by the form-factor operator product expansion (FFOPE). We provide plots of the remainder function on several kinematic slices, and show that the result is compatible with the FFOPE data. We further verify that antipodal self-duality is valid at two loops beyond the level of the symbol.

hep-th

Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models

This paper presents Dolphin, a novel decoder-decoder architecture for energy-efficient processing of long contexts in language models. Our approach addresses the significant energy consumption and latency challenges inherent in on-device models. Dolphin employs a compact 0.5B parameter decoder to distill extensive contextual information into a memory embedding, substantially reducing the input length for the primary 7B parameter decoder model. Inspired by vision-language models, we repurpose the image embedding projector to encode long textual contexts, effectively treating extended context as a distinct modality. This innovative method enables processing of substantially longer contexts without the typical computational overhead associated with extended input sequences. Empirical evaluations demonstrate a 10-fold improvement in energy efficiency and a 5-fold reduction in latency compared to conventional full-length context processing methods without losing quality of the response. Our work contributes to the development of more sustainable and scalable language models for on-device applications, addressing the critical need for energy-efficient and responsive AI technologies in resource-constrained environments while maintaining the accuracy to understand long contexts. This research has implications for the broader field of natural language processing, particularly in the domain of efficient model design for resource-limited settings. By enabling more sophisticated AI capabilities on edge devices, Dolphin paves the way for advanced language processing in a wide range of applications where computational resources are at a premium. The Dolphin model is publicly available at https://huggingface.co/NexaAIDev/Dolphin.

cs.CL

Dark magnetohydrodynamics: Black hole accretion in superradiant dark photon clouds

Black holes threaded by massive vector fields can be subject to a superradiant instability, growing a cloud of massive vector particles around it. In this work, we consider what happens if such a dark matter candidate field mimicking a dark photon interacts with an accretion flow onto the black hole. By including a kinetic mixing term with the standard model photon, we extend the commonly used equations of general-relativistic magnetohydrodynamics to a dark photon constituent. The coupling to the dark photon then appears as an effective dynamo term together with a dark Lorentz force acting on the accreting matter. We numerically study the interactions between the superradiant dark photon cloud and the inner accretion flow by solving the coupled system in full numerical relativity. By parameterically varying the mixing parameter between dark and standard model sector, we provide a first investigation of how the accretion flow could be modified. Depending on the coupling strength, our solutions exhibit increased wind launching, as well as oscillation modes in the disk.

astro-ph.HE

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers, which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore, recent studies have utilized prompt tuning to transfer pre-trained RGB-based trackers to multimodal data. However, the modality gap limits pre-trained knowledge recall, and the dominance of the RGB modality persists, preventing the full utilization of information from other modalities. To address these issues, we propose a novel symmetric multimodal tracking framework called SDSTrack. We introduce lightweight adaptation for efficient fine-tuning, which directly transfers the feature extraction ability from RGB to other domains with a small number of trainable parameters and integrates multimodal features in a balanced, symmetric manner. Furthermore, we design a complementary masked patch distillation strategy to enhance the robustness of trackers in complex environments, such as extreme weather, poor imaging, and sensor failure. Extensive experiments demonstrate that SDSTrack outperforms state-of-the-art methods in various multimodal tracking scenarios, including RGB+Depth, RGB+Thermal, and RGB+Event tracking, and exhibits impressive results in extreme conditions. Our source code is available at https://github.com/hoqolo/SDSTrack.

cs.CV

On the Origin of Acoustic Spin and Elastic Spin: Uncovering Hidden Wave Spin of Scalar Fields with Higher-Order Derivative Lagrangian

Scalar field should have no spin angular momentum according to conventional understandings in classical field theory. Yet, recent studies demonstrate the undoubted existence of wave spin endowed by acoustic and elastic longitudinal waves, which are of irrotational curl-free nature without vorticity and can be described by scalar fields. Here, to solve this seeming discrepancy, we uncover the origin of wave spin in scalar fields beyond traditional formalism by clarifying that the presence of higher order derivatives in scalar field Lagrangians can give rise to non-vanishing spin. For scalar fields with only first order derivative, we can make the hidden wave spin emerge, by constructing a latent field that leads to the original field through a time derivative so that is of second order. We exemplify the wave spin for elastic and acoustic fields, as well as for dissipative media, following Noether's theorem in higher-order derivative Lagrangian. The results would prompt people to build more comprehensive and fundamental understandings of structural wave spin in classical fields.

physics.class-ph

Gravitational-wave echoes from spinning exotic compact objects: numerical waveforms from the Teukolsky equation

We present numerical waveforms of gravitational-wave echoes from spinning exotic compact objects (ECOs) that result from binary black hole coalescence. We obtain these echoes by solving the Teukolsky equation for the $\psi_4$ associated with gravitational waves that propagate toward the horizon of a Kerr spacetime, and process the subsequent reflections of the horizon-going wave by the surface of the ECO, which lies right above the Kerr horizon. The trajectories of the infalling objects are modified from Kerr geodesics, such that the gravitational waves propagating toward future null infinity match those from merging black holes with comparable masses. In this way, the corresponding echoes approximate to those from comparable-mass mergers. For boundary conditions at the ECO surface, we adopt recent work using the membrane paradigm, which relates $\psi_0$ associated with the horizon-going wave and $\psi_4$ of the wave that leaves the ECO surface. We obtain $\psi_0$ of the horizon-going wave from $\psi_4$ using the Teukolsky-Starobinsky relation. The echoes we obtain turn out to be significantly weaker than those from previous studies that generate echo waveforms by modeling the ringdown part of binary black hole coalescence waveforms as originating from the past horizon.

gr-qc

Very extreme mass-ratio bursts in the Galaxy and neighbouring galaxies in relation to space-borne detectors

Two recent papers\citep{xmri1, xmri2} revealed that in our Galaxy there are very extreme-mass-ratio inspirals composed by brown dwarfs and the supermassive black hole at the center of the Galaxy. The event rates estimated in these papers are very considerable for future space-borne detectors. In addition, there are plunge events during the formation of inspiraling orbits. In this work, we calculate the gravitational waves from compact objects (brown dwarf, primordial black hole and etc.) plunging into or being scattered by the central supermassive black hole. We find that for space-borne detectors the signal-to-noise ratios of these bursts are quite high. The event rates are estimated as $\sim$ $0.01 {\rm{yr}^{-1}}$ for the Galaxy. If we are lucky, this kind of very extreme-mass-ratio bursts will offer a unique chance to reveal the nearest supermassive black hole and nuclei dynamics. The event rate can be as large as 4 $\sim$ 8 ${\rm yr^{-1}}$ in 10 Mpc, and because the signal is strong enough for observations by space-borne detectors, we have a good chance of being able to probe the nature of neighboring black holes.

gr-qc

Testing dispersion of gravitational waves from eccentric extreme-mass-ratio inspirals

In general relativity, there is no dispersion in gravitational waves, while some modified gravity theories predict dispersion phenomena in the propagation of gravitational waves. In this paper, we demonstrate that this dispersion will induce an observable deviation of waveforms if the orbits have large eccentricities. The mechanism is that the waveform modes with different frequencies will be emitted at the same time due to the existence of eccentricity. During the propagation, because of the dispersion, the arrival time of different modes will be different, then produce the deviation and dephasing of waveforms compared with general relativity. This kind of dispersion phenomena related with extreme-mass-ratio inspirals could be observed by space-borne detectors, and the constraint on the graviton mass could be improved . Moreover, we find that the dispersion effect may also be constrained by ground detectors better than the current result if a highly eccentric intermediate-mass-ratio inspirals be observed.

gr-qc

The Sloan Digital Sky Survey Reverberation Mapping Project: Accretion and Broad Emission Line Physics from a Hypervariable Quasar

We analyze extensive spectroscopic and photometric data of the hypervariable quasar SDSS J131424+530527 (RMID 017) at z=0.456, an optical "changing look" quasar from the Sloan Digital Sky Survey Reverberation Mapping project that increased in optical luminosity by a factor of 10 between 2014 and 2017. The observed broad emission lines all respond in luminosity and width to the changing optical continuum, as expected for photoionization in a stratified, virialized broad emission line region. The luminosity changes therefore result from intrinsic changes in accretion power rather than variable obscuration. The variability is continuous and apparently stochastic, disfavoring an origin as a discrete event such as a tidal disruption flare or microlensing event. It is coordinated on day timescales with blue leading red, consistent with reprocessing powering the entire optical SED. We show that this process cannot work in a standard thin disk geometry on energetic grounds, and would instead require a large covering factor reprocessor. Disk instability models could potentially also explain the data, provided that the instability sets in near the inner radius of a geometrically thick accretion disk.

astro-ph.GA

Testing General Relativity with X-ray reflection spectroscopy: The Konoplya-Rezzolla-Zhidenko parametrization

X-ray reflection spectroscopy is a promising technique for testing general relativity in the strong field regime, as it can be used to test the Kerr black hole hypothesis. In this context, the parametrically deformed black hole metrics proposed by Konoplya, Rezzolla \& Zhidenko (Phys. Rev. D93, 064015, 2016) form an important class of non-Kerr black holes. We implement this class of black hole metrics in \textsc{relxill\_nk}, which is a framework we have developed for testing for non-Kerr black holes using X-ray reflection spectroscopy. We perform a qualitative analysis of the effect of the leading order strong-field deformation parameters on typical observables like the innermost stable circular orbits and the reflection spectra. We also present the first X-ray constraints on some of the deformation parameters of this metric, using \textit{Suzaku} data from the supermassive black hole in Ark~564, and compare them with those obtained (or expected) from other observational techniques like gravitational waves and black hole imaging.

gr-qc

Gravitational wave emission under general parametrized metric from extreme mass ratio inspirals

Future space-borne interferometers will be able to detect gravitational waves at $10^{-3}$ to $10^{-1}$ Hz. At this band extreme-mass-ratio inspirals (EMRIs) can be promising gravitational wave sources. In this paper, we investigate possibility of testing Kerr hypothesis against a parametrized non-Kerr metric by matching EMRI signals. However, EMRIs from either equatorial orbits or inclined orbits suffer from the "confusion problem". Our results show that, within the time scale before radiation flux plays an important role, small and moderate deviations from the Kerr spacetime($|δ_i|<1$) can be discerned only when spin parameter is high. In most cases, the EMRI waveforms related with a non-Kerr metric can be mimicked by the waveform templates produced with a Kerr black hole.

gr-qc