SearcharxivSearch

arXiv subjects

Hang Yu

Publications and source records attributed to Hang Yu.

At least 127 records · Page 7Linked to original sources

Uncovering Stealth Bias in LISA observations of Double White Dwarf Binaries due to Tidal Coupling

Double white dwarfs are important gravitational wave sources for LISA, as they are some of the most numerous compact systems in our universe. Here we consider finite-sized effects due to tidal interactions, as they are expected to have a measurable impact on these systems. Previous studies suggested that tidal effects would allow the individual masses to be measured, but there was a subtle error in those analyses. Using a fully Bayesian analysis we find that while tidal effects do not allow us to constrain the individual masses, they do yield informative lower bounds on the total mass of the system. Including tidal effects is crucial to the accuracy of our estimation of the chirp and total mass. Neglecting tidal effects leads to significant biases towards higher chirp masses, and we see that the lower bound of the total masses is biased towards a higher value as well. For many systems observed by LISA, tidal effects can lead to a "stealth" bias, since only the first derivative of the frequency can be measured. To separate tidal effects from the usual point particle decay we need to be able to measure the change in the second derivative of the frequency cause by the tides. This can only be done for high frequency systems observed with high signal-to-noise. The bias, if not accounted for, can have significant astrophysical implications; for example, it could lead to an incorrect estimation of the population of potential Type IA supernovae progenitors.

astro-ph.HE

DUPLEX: Dual GAT for Complex Embedding of Directed Graphs

Current directed graph embedding methods build upon undirected techniques but often inadequately capture directed edge information, leading to challenges such as: (1) Suboptimal representations for nodes with low in/out-degrees, due to the insufficient neighbor interactions; (2) Limited inductive ability for representing new nodes post-training; (3) Narrow generalizability, as training is overly coupled with specific tasks. In response, we propose DUPLEX, an inductive framework for complex embeddings of directed graphs. It (1) leverages Hermitian adjacency matrix decomposition for comprehensive neighbor integration, (2) employs a dual GAT encoder for directional neighbor modeling, and (3) features two parameter-free decoders to decouple training from particular tasks. DUPLEX outperforms state-of-the-art models, especially for nodes with sparse connectivity, and demonstrates robust inductive capability and adaptability across various tasks. The code is available at https://github.com/alipay/DUPLEX.

cs.LG

SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy

Text-to-SQL conversion is a critical innovation, simplifying the transition from complex SQL to intuitive natural language queries, especially significant given SQL's prevalence in the job market across various roles. The rise of Large Language Models (LLMs) like GPT-3.5 and GPT-4 has greatly advanced this field, offering improved natural language understanding and the ability to generate nuanced SQL statements. However, the potential of open-source LLMs in Text-to-SQL applications remains underexplored, with many frameworks failing to leverage their full capabilities, particularly in handling complex database queries and incorporating feedback for iterative refinement. Addressing these limitations, this paper introduces SQLfuse, a robust system integrating open-source LLMs with a suite of tools to enhance Text-to-SQL translation's accuracy and usability. SQLfuse features four modules: schema mining, schema linking, SQL generation, and a SQL critic module, to not only generate but also continuously enhance SQL query quality. Demonstrated by its leading performance on the Spider Leaderboard and deployment by Ant Group, SQLfuse showcases the practical merits of open-source LLMs in diverse business contexts.

cs.CL

Dynamical tides during the inspiral of rapidly spinning neutron stars: Solutions beyond mode resonance

We investigate the dynamical tide in a gravitational wave (GW)-driven coalescing binary involving a neutron star (NS). The NS is assumed to spin rapidly, with its spin axis anti-aligned with the orbit. Such an NS may exist if the binary forms dynamically in a dense environment, and it can lead to a strong tide because the f-mode can be resonantly excited during the inspiral. We present a new analytical solution for the f-mode resonance by decomposing the tide into a resummed equilibrium component and a dynamical component that is excited only around resonance. This solution simplifies numerical implementations by avoiding the subtraction of two diverging terms. It also extends the solution's validity to frequencies beyond mode resonance. When the dynamical tide back reacts on the orbit, the commonly adopted effective Love number is insufficient because it does not capture the tidal torque on the orbit that dominates the back reaction during mode resonance. An additional dressing factor originating from the imaginary part of the Love number is introduced to model the torque. The dissipative interaction between the NS and the orbital mass multipoles is computed including the dynamical tide. Orbital phase shifts caused by the $l=3$ and $l=2$ f-modes can reach 0.5 and 10 radians at their respective resonances if the NS has a spin rate of 850 Hz. Because of the large impact of the dynamical tide, a linearized analytical description becomes insufficient. After mode excitation, the orbit cannot remain quasi-circular, and the eccentricity excited by the dynamical tide can approach $e\simeq 0.1$, leading to non-monotonic frequency evolution which breaks the stationary phase approximation commonly adopted by frequency-domain waveform constructions. The GW radiation from the excited f-mode alone can be detected with a signal-to-noise ratio exceeding unity with the next-generation detectors.

gr-qc

Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene

The unsupervised 3D object detection is to accurately detect objects in unstructured environments with no explicit supervisory signals. This task, given sparse LiDAR point clouds, often results in compromised performance for detecting distant or small objects due to the inherent sparsity and limited spatial resolution. In this paper, we are among the early attempts to integrate LiDAR data with 2D images for unsupervised 3D detection and introduce a new method, dubbed LiDAR-2D Self-paced Learning (LiSe). We argue that RGB images serve as a valuable complement to LiDAR data, offering precise 2D localization cues, particularly when scarce LiDAR points are available for certain objects. Considering the unique characteristics of both modalities, our framework devises a self-paced learning pipeline that incorporates adaptive sampling and weak model aggregation strategies. The adaptive sampling strategy dynamically tunes the distribution of pseudo labels during training, countering the tendency of models to overfit easily detected samples, such as nearby and large-sized objects. By doing so, it ensures a balanced learning trajectory across varying object scales and distances. The weak model aggregation component consolidates the strengths of models trained under different pseudo label distributions, culminating in a robust and powerful final model. Experimental evaluations validate the efficacy of our proposed LiSe method, manifesting significant improvements of +7.1% AP$_{BEV}$ and +3.4% AP$_{3D}$ on nuScenes, and +8.3% AP$_{BEV}$ and +7.4% AP$_{3D}$ on Lyft compared to existing techniques.

cs.CV

Are WASP-107-like Systems Consistent with High-eccentricity Migration?

WASP-107 b seems to be a poster child of the long-suspected high-eccentricity migration scenario. It is on a 5.7-day, polar orbit. The planet is Jupiter-like in radius but Neptune-like in mass with exceptionally low density. WASP-107 c is on a 1100-day, $e=0.28$ orbit with at least Saturn mass. Planet b may still have a residual eccentricity of $0.06\pm 0.04$: the ongoing tidal dissipation leads to the observed internally heated atmosphere and hydrodynamic atmospheric erosion. We present a population synthesis study coupling octopole Lidov-Kozai oscillations with various short-range forces, while simultaneously accounting for the radius inflation and tidal disruption of the planet. We find that a high-eccentricity migration scenario can successfully explain nearly all observed system properties. Our simulations further suggest that the initial location of WASP-107 b at the onset of migration is likely within the snowline ($<0.5\,{\rm AU}$). More distant initial orbits usually lead to tidal disruption or orbit crossing. WASP-107 b most likely lost no more than 20% of its mass during the high-eccentricity migration, i.e. it did not form as a Jupiter-mass object. More vigorous tidally-induced mass loss leads to disruption of the planet during migration. We predict that the current-day mutual inclination between the planets b and c is substantial: at least 25-55$^\circ$ which may be tested with future Gaia astrometric observations. Knowing the current-day mutual inclination may further constrain the initial orbit of planet b. We suggest that the proposed high-eccentricity migration scenario of WASP-107 may be applicable to HAT-P-11, GJ-3470, HAT-P-18, and GJ-436 which have similar orbital architectures.

astro-ph.EP

Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

In this work we systematically review the recent advancements in software engineering with language models, covering 70+ models, 40+ evaluation tasks, 180+ datasets, and 900 related works. Unlike previous works, we integrate software engineering (SE) with natural language processing (NLP) by discussing the perspectives of both sides: SE applies language models for development automation, while NLP adopts SE tasks for language model evaluation. We break down code processing models into general language models represented by the GPT family and specialized models that are specifically pretrained on code, often with tailored objectives. We discuss the relations and differences between these models, and highlight the historical transition of code modeling from statistical models and RNNs to pretrained Transformers and LLMs, which is exactly the same course that had been taken by NLP. We also go beyond programming and review LLMs' application in other software engineering activities including requirement engineering, testing, deployment, and operations in an endeavor to provide a global view of NLP in SE, and identify key challenges and potential future directions in this domain. We keep the survey open and updated on GitHub at https://github.com/codefuse-ai/Awesome-Code-LLM.

cs.CL

D2LLM: Decomposed and Distilled Large Language Models for Semantic Search

The key challenge in semantic search is to create models that are both accurate and efficient in pinpointing relevant sentences for queries. While BERT-style bi-encoders excel in efficiency with pre-computed embeddings, they often miss subtle nuances in search tasks. Conversely, GPT-style LLMs with cross-encoder designs capture these nuances but are computationally intensive, hindering real-time applications. In this paper, we present D2LLMs-Decomposed and Distilled LLMs for semantic search-that combines the best of both worlds. We decompose a cross-encoder into an efficient bi-encoder integrated with Pooling by Multihead Attention and an Interaction Emulation Module, achieving nuanced understanding and pre-computability. Knowledge from the LLM is distilled into this model using contrastive, rank, and feature imitation techniques. Our experiments show that D2LLM surpasses five leading baselines in terms of all metrics across three tasks, particularly improving NLI task performance by at least 6.45%. The source code is available at https://github.com/codefuse-ai/D2LLM.

cs.CL

Complex scaling in finite volume

Quantum resonances, i.e., metastable states with a finite lifetime, play an important role in nuclear physics and other domains. Describing this phenomenon theoretically is generally a challenging task. In this work, we combine two established techniques to address this challenge. Complex scaling makes it possible to calculate resonances with bound-state-like methods. Finite-volume simulations exploit the fact that the infinite-volume properties of quantum systems are encoded in how discrete energy levels change as one varies the size of the volume. We apply complex scaling to systems in finite periodic boxes and derive the volume dependence of states in this scenario, demonstrating with explicit examples how one can use these relations to infer infinite-volume resonance energies and lifetimes.

nucl-th

Channel Access Methods for RF-Powered IoT Networks: A Survey

Many Internet of Things (IoT) networks with Radio Frequency (RF) powered devices operate over a shared medium. They thus require a channel access protocol. Unlike conventional networks where devices have unlimited energy, in an RF-powered IoT network, devices must first harvest RF energy in order to transmit or/and receive data. To this end, this survey presents the {\em first} comprehensive review of prior works that employ contention-based and contention-free protocols in IoT networks with one or more {\em dedicated} energy sources. Specifically, these protocols work in conjunction with RF-energy sources to deliver energy delivery or/and data. In this respect, this survey covers protocols based on Aloha, Carrier Sense Multiple Access (CSMA), polling, and dynamic Time Division Multiple Access (TDMA). Further, it covers successive interference cancellation protocols. It highlights key issues and challenges addressed by prior works, and provides a qualitative comparison of these works. Lastly, it identifies gaps in the literature and presents a list of future research directions.

cs.NI

MAVRL: Learn to Fly in Cluttered Environments with Varying Speed

Many existing obstacle avoidance algorithms overlook the crucial balance between safety and agility, especially in environments of varying complexity. In our study, we introduce an obstacle avoidance pipeline based on reinforcement learning. This pipeline enables drones to adapt their flying speed according to the environmental complexity. Moreover, to improve the obstacle avoidance performance in cluttered environments, we propose a novel latent space. The latent space in this representation is explicitly trained to retain memory of previous depth map observations. Our findings confirm that varying speed leads to a superior balance of success rate and agility in cluttered environments. Additionally, our memory-augmented latent representation outperforms the latent representation commonly used in reinforcement learning. Finally, after minimal fine-tuning, we successfully deployed our network on a real drone for enhanced obstacle avoidance.

cs.RO

BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis

Bases have become an integral part of modern deep learning-based models for time series forecasting due to their ability to act as feature extractors or future references. To be effective, a basis must be tailored to the specific set of time series data and exhibit distinct correlation with each time series within the set. However, current state-of-the-art methods are limited in their ability to satisfy both of these requirements simultaneously. To address this challenge, we propose BasisFormer, an end-to-end time series forecasting architecture that leverages learnable and interpretable bases. This architecture comprises three components: First, we acquire bases through adaptive self-supervised learning, which treats the historical and future sections of the time series as two distinct views and employs contrastive learning. Next, we design a Coef module that calculates the similarity coefficients between the time series and bases in the historical view via bidirectional cross-attention. Finally, we present a Forecast module that selects and consolidates the bases in the future view based on the similarity coefficients, resulting in accurate future predictions. Through extensive experiments on six datasets, we demonstrate that BasisFormer outperforms previous state-of-the-art methods by 11.04\% and 15.78\% respectively for univariate and multivariate forecasting tasks. Code is available at: \url{https://github.com/nzl5116190/Basisformer}

cs.LG

CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model

Code Large Language Models (Code LLMs) have gained significant attention in the industry due to their wide applications in the full lifecycle of software engineering. However, the effectiveness of existing models in understanding non-English inputs for multi-lingual code-related tasks is still far from well studied. This paper introduces CodeFuse-13B, an open-sourced pre-trained code LLM. It is specifically designed for code-related tasks with both English and Chinese prompts and supports over 40 programming languages. CodeFuse achieves its effectiveness by utilizing a high quality pre-training dataset that is carefully filtered by program analyzers and optimized during the training process. Extensive experiments are conducted using real-world usage scenarios, the industry-standard benchmark HumanEval-x, and the specially designed CodeFuseEval for Chinese prompts. To assess the effectiveness of CodeFuse, we actively collected valuable human feedback from the AntGroup's software development process where CodeFuse has been successfully deployed. The results demonstrate that CodeFuse-13B achieves a HumanEval pass@1 score of 37.10%, positioning it as one of the top multi-lingual code LLMs with similar parameter sizes. In practical scenarios, such as code generation, code translation, code comments, and testcase generation, CodeFuse performs better than other models when confronted with Chinese prompts.

cs.SE

Detecting gravitational lensing in hierarchical triples in galactic nuclei with space-borne gravitational-wave observatories

Stellar-mass binary black holes (BBHs) may merge in the vicinity of a supermassive black hole (SMBH). It is suggested that the gravitational-wave (GW) emitted by a BBH has a high probability to be lensed by the SMBH if the BBH's orbit around the SMBH (i.e., the outer orbit) has a period of less than a year and is less than the duration of observation of the BBH by a space-borne GW observatory. For such a BBH + SMBH triple system, the de Sitter precession of the BBH's orbital plane is also significant. In this work, we thus study GW waveforms emitted by the BBH and then modulated by the SMBH due to effects including Doppler shift, de Sitter precession, and gravitational lensing. We show specifically that for an outer orbital period of 0.1 yr and an SMBH mass of $10^7 M_\odot$, there is a 3\%-10\% chance for the standard, strong lensing signatures to be detectable by space-borne GW detectors such as LISA and/or TianGO. For more massive lenses ($\gtrsim 10^8 M_\odot$) and more compact outer orbits with periods <0.1 yr, retro-lensing of the SMBH might also have a 1%-level chance of detection. Furthermore, by combining the lensing effects and the dynamics of the outer orbit, we find the mass of the central SMBH can be accurately determined with a fraction error of $\sim 10^{-4}$. This is much better than the case of static lensing because the degeneracy between the lens' mass and the source's angular position is lifted by the outer orbital motion. Including lensing effects also allows the de Sitter precession to be detectable at a precession period 3 times longer than the case without lensing. Lastly, we demonstrate that one can check the consistency between the SMBH's mass determined from the orbital dynamics and the one inferred from gravitational lensing, which serves as a test on theories behind both phenomena. The statistical error on the deviation can be constrained to a 1% level.

gr-qc

Accurate and Efficient Waveform Model for Precessing Binary Black Holes

We present IMRPhenomXODE, a new phenomenological frequency-domain waveform approximant for gravitational wave (GW) signals from precessing binary black holes (BBHs) with generic spin configurations. We build upon the success of IMRPhenomXPHM [G. Pratten et al., Phys. Rev. D 103, 104056 (2021), which is one of the most widely adopted waveform approximants in GW data analyses that include spin precession, and introduce two additional significant improvements. First, we employ an efficient technique to numerically solve the (next-to)$^4$-leading-order post-Newtonian precession equations, which allows us to accurately determine the evolution of the orientation of the orbital angular momentum $\boldsymbol{\hat{L}}_{\rm N}$ even in cases with complicated precession dynamics, such as transitional precession. Second, we recalibrate the phase of GW modes in the frame coprecessing with $\boldsymbol{\hat{L}}_{\rm N}$ against SEOBNRv4PHM [S. Ossokine et al., Phys. Rev. D 102, 044055 (2020)] to capture effects due to precession such as variations in the spin components aligned with $\boldsymbol{\hat{L}}_{\rm N}$. By incorporating these new features, IMRPhenomXODE achieves matches with SEOBNRv4PHM that are better than 99% for most BBHs with mass ratios $q \geq 1/6$ and with arbitrary spin configurations. In contrast, the mismatch between IMRPhenomXPHM and SEOBNRv4PHM often exceeds 10% for a BBH with $q\lesssim 1/2$ and large in-plane or antialigned spin components. Our implementation is also computationally efficient, with waveform evaluation times that can even be shorter than those of IMRPhenomXPHM for BBH signals with long durations and hence high frequency resolutions. The accuracy and efficiency of IMRPhenomXODE position it as a valuable tool for GW event searches, parameter estimation analyses, and the inference of underlying population properties.

gr-qc

Charged-particle bound states in periodic boxes

We consider the binding energy of a two-body system with a repulsive Coulomb interaction in a finite periodic volume. We define the finite-volume Coulomb potential as the usual Coulomb potential, except that the distance is defined as the shortest separation between the two bodies in the periodic volume. We investigate this problem in one and three-dimensional periodic boxes and derive the asymptotic behavior of the volume dependence for bound states with zero angular momentum in terms of Whittaker functions. We benchmark our results against numerical calculations and show how the method can be used to extract asymptotic normalization coefficients for charged-particle bound states. The results we derive here have immediate applications for calculations of atomic nuclei in finite periodic volumes for the case where the leading finite-volume correction is associated with two charged clusters.

nucl-th

MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Code LLMs have emerged as a specialized research field, with remarkable studies dedicated to enhancing model's coding capabilities through fine-tuning on pre-trained models. Previous fine-tuning approaches were typically tailored to specific downstream tasks or scenarios, which meant separate fine-tuning for each task, requiring extensive training resources and posing challenges in terms of deployment and maintenance. Furthermore, these approaches failed to leverage the inherent interconnectedness among different code-related tasks. To overcome these limitations, we present a multi-task fine-tuning framework, MFTcoder, that enables simultaneous and parallel fine-tuning on multiple tasks. By incorporating various loss functions, we effectively address common challenges in multi-task learning, such as data imbalance, varying difficulty levels, and inconsistent convergence speeds. Extensive experiments have conclusively demonstrated that our multi-task fine-tuning approach outperforms both individual fine-tuning on single tasks and fine-tuning on a mixed ensemble of tasks. Moreover, MFTcoder offers efficient training capabilities, including efficient data tokenization modes and PEFT fine-tuning, resulting in significantly improved speed compared to traditional fine-tuning methods. MFTcoder seamlessly integrates with several mainstream open-source LLMs, such as CodeLLama and Qwen. Leveraging the CodeLLama foundation, our MFTcoder fine-tuned model, \textsc{CodeFuse-CodeLLama-34B}, achieves an impressive pass@1 score of 74.4\% on the HumaneEval benchmark, surpassing GPT-4 performance (67\%, zero-shot). MFTCoder is open-sourced at \url{https://github.com/codefuse-ai/MFTCOder}

cs.LG

Measuring Supermassive Black Hole Properties via Gravitational Radiation from Eccentrically Orbiting Stellar Mass Black Hole Binaries

There may exist stellar-mass binary black holes (BBH) which merge while orbiting nearby a supermassive black hole (SMBH). In such a triple system, the SMBH will modulate the gravitational waveform of the BBH through orbital Doppler shift and de Sitter precession of the angular momentum. Future space-based GW observatories focused on the milli- and decihertz band will be uniquely poised to observe these waveform modulations, as the GW frequency from stellar-mass BBHs varies slowly in this band while modulation effects accumulate. In this work, we apply the Fisher information matrix formalism to estimate how well space-borne GW detectors can measure properties of BBH+SMBH hierarchical triples using the GW from orbiting BBH. We extend previous work by considering the more realistic case of an eccentric orbit around the SMBH, and notably include the effects of orbital pericenter precession. We find that for detector concepts such as LISA, B-DECIGO, and TianGO, we can extract the SMBH mass and semimajor axis of the orbit with a fractional uncertainty below the 0.1% level over a wide range of triple system parameters. Furthermore, we find that the effects of pericenter precession and orbital eccentricity significantly improve our ability to measure this system. We also find that while LISA could measure these systems, the decihertz detector concepts B-DECIGO and TianGO would enable better sensitivity to the triple's parameters.

gr-qc