SearcharxivSearch

arXiv subjects

Yifeng Li

Publications and source records attributed to Yifeng Li.

At least 19 recordsLinked to original sources

A Pilot Study of Mildly Recycled Pulsars: A Case Study of PSR J2338+4818

Mildly recycled pulsars are neutron stars partially spun up through relatively short mass-transfer phases, typically with massive carbon-oxygen (CO) or oxygen-neon-magnesium (ONeMg) white dwarf companions. PSR J2338+4818, a mildly recycled pulsar, was discovered with the Five-hundred-meter Aperture Spherical Telescope (FAST). As a pilot study on the formation and evolutionary pathways of mildly recycled pulsars, we present the updated timing solution for PSR J2338+4818 and examine its single pulses and scintillation properties. Aided by the sensitivity of FAST, the single pulses of PSR J2338+4818 were systematically studied. 27,228 single pulses with S/N > 7 have been detected in our observations. For the FAST ultra-wideband observation on MJD 61045, the receiver was still in the technical commissioning phase, and then only a preliminary single-pulse search was performed. Pulse nulling was examined using a Markov Chain Monte Carlo (MCMC) method, but no evidence for nulling was found. The possible long-term nulling reported by previous studies did not occur in any of our observations in either the 1.0 to 1.5 GHz band or the 300 to 600 MHz band. Interstellar scintillation is evident in our observations. The measured scintillation timescales and bandwidths range from 2.93 to 25.26 minutes and 1.68 to 27.41 MHz, respectively. In all observations, no clear scintillation arc was found in the secondary spectra of PSR J2338+4818.

astro-ph.HE

The FAST Discovery of a binary millisecond pulsar PSR~J1647-0156B (M12B) with a candidate cross matching algorithm

We propose a pulsar candidate cross matching algorithm to sift radio pulsar search candidates from repeated observations of the same sky location such as globular clusters, high energy sources, or supernova remnants. Our method uses both the candidate spin period ($P$) and dispersion measure (DM) value; if two or more candidates from different observations have similar spin periods to within 1\%, and dispersion measure values within 10\%, they are likely to correspond to the same candidate detection. We have demonstrated the effectiveness of our method through the discovery of the pulsar M12B with the Five-hundred-meter Aperture Spherical radio Telescope (FAST). This pulsar has a spin period of 2.76\,ms and a dispersion measure of $42.70 \pm 0.05\,\mathrm{cm}^{-3}~\mathrm{pc}$. This pulsar has a profile with three peaks, being faint, showing scintillation. It is in an approximately 0.53-day orbit. Our discovery indicates that more pulsars might be effectively discovered if the algorithm is applied to the search results from other archival globular cluster observations.

astro-ph.HE

Intuitive control of supernumerary robotic limbs through a tactile-encoded neural interface

Brain-computer interfaces (BCIs) promise to extend human movement capabilities by enabling direct neural control of supernumerary effectors, yet integrating augmented commands with multiple degrees of freedom without disrupting natural movement remains a key challenge. Here, we propose a tactile-encoded BCI that leverages sensory afferents through a novel tactile-evoked P300 paradigm, allowing intuitive and reliable decoding of supernumerary motor intentions even when superimposed with voluntary actions. The interface was evaluated in a multi-day experiment comprising of a single motor recognition task to validate baseline BCI performance and a dual task paradigm to assess the potential influence between the BCI and natural human movement. The brain interface achieved real-time and reliable decoding of four supernumerary degrees of freedom, with significant performance improvements after only three days of training. Importantly, after training, performance did not differ significantly between the single- and dual-BCI task conditions, and natural movement remained unimpaired during concurrent supernumerary control. Lastly, the interface was deployed in a movement augmentation task, demonstrating its ability to command two supernumerary robotic arms for functional assistance during bimanual tasks. These results establish a new neural interface paradigm for movement augmentation through stimulation of sensory afferents, expanding motor degrees of freedom without impairing natural movement.

cs.RO

GR-3 Technical Report

We report our recent progress towards building generalist robot policies, the development of GR-3. GR-3 is a large-scale vision-language-action (VLA) model. It showcases exceptional capabilities in generalizing to novel objects, environments, and instructions involving abstract concepts. Furthermore, it can be efficiently fine-tuned with minimal human trajectory data, enabling rapid and cost-effective adaptation to new settings. GR-3 also excels in handling long-horizon and dexterous tasks, including those requiring bi-manual manipulation and mobile movement, showcasing robust and reliable performance. These capabilities are achieved through a multi-faceted training recipe that includes co-training with web-scale vision-language data, efficient fine-tuning from human trajectory data collected via VR devices, and effective imitation learning with robot trajectory data. In addition, we introduce ByteMini, a versatile bi-manual mobile robot designed with exceptional flexibility and reliability, capable of accomplishing a wide range of tasks when integrated with GR-3. Through extensive real-world experiments, we show GR-3 surpasses the state-of-the-art baseline method, $\pi_0$, on a wide variety of challenging tasks. We hope GR-3 can serve as a step towards building generalist robots capable of assisting humans in daily life.

cs.RO

FAST Observation and Results for Core Collapse Globular Cluster M15 and NGC 6517

Radio astronomy is part of radio science that developed rapidly in recent decades. In the research of radio astronomy, pulsars have always been an enduring popular research target. To find and observe more pulsars, large radio telescopes have been built all over the world. In this paper, we present our studies on pulsars in M15 and NGC 6517 with FAST, including monitoring pulsars in M15 and new pulsar discoveries in NGC 6517. All the previously known pulsars in M15 were detected without no new discoveries. Among them, M15C was still detectable by FAST, while it is assumed to fade out due to precession [1]. In NGC 6517, new pulsars were continues to be discovered and all of them are tend to be isolated pulsars. Currently, the number of pulsars in NGC 6517 is 17, much more than the predicted before [2].

astro-ph.HE

Sound insulation performance of multi-layer membrane-type acoustic metamaterials based on orthogonal experiments

The challenge of achieving effective sound insulation using metamaterials persists in the field. In this research endeavor, a novel three-layer membrane-type acoustic metamaterial is introduced as a potential solution. Through the application of orthogonal experiments, remarkable sound insulation capabilities are demonstrated within the frequency spectrum of 100-1200 Hz. The sound insulation principle of membrane-type acoustic metamaterial is obtained through the analysis of eigenmodes at the peak and trough points, combined with sound transmission loss. In addition, an orthogonal experiment is utilized to pinpoint the critical factors that impact sound insulation performance. By using relative bandwidth as the classification criterion, the optimal combination of influencing factors is determined, thereby improving the sound transmission loss of the multi-layer membrane-type acoustic metamaterial structure and broadening the sound insulation bandwidth. This study not only contributes a fresh and practical approach to insulation material design but also offers valuable insights into advancing sound insulation technology.

physics.app-ph

Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index

As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a novel metric that combines accuracy scores with token consumption, adjusted by a user-specified cost concern level to reflect different resource constraints. Our study examines 6 advanced prompting techniques, including Chain-of-Thought, Self-Consistency, and Tree of Thoughts, across 10 widely-used language models and 4 diverse datasets. We demonstrate that approaches such as Self-Consistency often provide statistically insignificant gains while becoming cost-prohibitive. For example, on high-performing models like Claude 3.5 Sonnet, the EPI of simpler techniques like Chain-of-Thought (0.72) surpasses more complex methods like Self-Consistency (0.64) at slight cost concern levels. Our findings suggest a reevaluation of complex prompting strategies in resource-constrained scenarios, potentially reshaping future research priorities and improving cost-effectiveness for end-users.

cs.CL

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

We present GR-2, a state-of-the-art generalist robot agent for versatile and generalizable robot manipulation. GR-2 is first pre-trained on a vast number of Internet videos to capture the dynamics of the world. This large-scale pre-training, involving 38 million video clips and over 50 billion tokens, equips GR-2 with the ability to generalize across a wide range of robotic tasks and environments during subsequent policy learning. Following this, GR-2 is fine-tuned for both video generation and action prediction using robot trajectories. It exhibits impressive multi-task learning capabilities, achieving an average success rate of 97.7% across more than 100 tasks. Moreover, GR-2 demonstrates exceptional generalization to new, previously unseen scenarios, including novel backgrounds, environments, objects, and tasks. Notably, GR-2 scales effectively with model size, underscoring its potential for continued growth and application. Project page: \url{https://gr2-manipulation.github.io}.

cs.RO

DyGMamba: Efficiently Modeling Long-Term Temporal Dependency on Continuous-Time Dynamic Graphs with State Space Models

Learning useful representations for continuous-time dynamic graphs (CTDGs) is challenging, due to the concurrent need to span long node interaction histories and grasp nuanced temporal details. In particular, two problems emerge: (1) Encoding longer histories requires more computational resources, making it crucial for CTDG models to maintain low computational complexity to ensure efficiency; (2) Meanwhile, more powerful models are needed to identify and select the most critical temporal information within the extended context provided by longer histories. To address these problems, we propose a CTDG representation learning model named DyGMamba, originating from the popular Mamba state space model (SSM). DyGMamba first leverages a node-level SSM to encode the sequence of historical node interactions. Another time-level SSM is then employed to exploit the temporal patterns hidden in the historical graph, where its output is used to dynamically select the critical information from the interaction history. We validate DyGMamba experimentally on the dynamic link prediction task. The results show that our model achieves state-of-the-art in most cases. DyGMamba also maintains high efficiency in terms of computational resources, making it possible to capture long temporal dependencies with a limited computation budget.

cs.LG

Timing and Scintillation Studies of Pulsars in Globular Cluster M3 (NGC 5272) with FAST

We present the phase-connected timing solutions of all the five pulsars in globular cluster (GC) M3 (NGC 5272), namely PSRs M3A to F (PSRs J1342+2822A to F), with the exception of PSR M3C, from FAST archival data. In these timing solutions, those of PSRs M3E, and F are obtained for the first time. We find that PSRs M3E and F have low mass companions, and are in circular orbits with periods of 7.1 and 3.0 days, respectively. For PSR M3C, we have not detected it in all the 41 observations. We found no X-ray counterparts for these pulsars in archival Chandra images in the band of 0.2-20 keV. We noticed that the pulsars in M3 seem to be native. From the Auto-Correlation Function (ACF) analysis of the M3A's and M3B's dynamic spectra, the scintillation timescale ranges from $7.0\pm0.3$ min to $60.0\pm0.6$ min, and the scintillation bandwidth ranges from $4.6\pm0.2$ MHz to $57.1\pm1.1$ MHz. The measured scintillation bandwidths from the dynamic spectra indicate strong scintillation, and the scattering medium is anisotropic. From the secondary spectra, we captured a scintillation arc only for PSR M3B with a curvature of $649\pm23 {\rm m}^{-1} {\rm mHz}^{-2}$.

astro-ph.HE

Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge

Large Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning. However, applying this reasoning to multimodal domains, where understanding text and images together is essential, remains a substantial challenge. To address this, we introduce WinoVis, a novel dataset specifically designed to probe text-to-image models on pronoun disambiguation within multimodal contexts. Utilizing GPT-4 for prompt generation and Diffusion Attentive Attribution Maps (DAAM) for heatmap analysis, we propose a novel evaluation framework that isolates the models' ability in pronoun disambiguation from other visual processing challenges. Evaluation of successive model versions reveals that, despite incremental advancements, Stable Diffusion 2.0 achieves a precision of 56.7% on WinoVis, only marginally surpassing random guessing. Further error analysis identifies important areas for future research aimed at advancing text-to-image models in their ability to interpret and interact with the complex visual world.

cs.CL

Towards Unified Interactive Visual Grounding in The Wild

Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user input by active information gathering. Previous approaches often rely on predefined templates to ask disambiguation questions, resulting in performance reduction in realistic interactive scenarios. In this paper, we propose TiO, an end-to-end system for interactive visual grounding in human-robot interaction. Benefiting from a unified formulation of visual dialogue and grounding, our method can be trained on a joint of extensive public data, and show superior generality to diversified and challenging open-world scenarios. In the experiments, we validate TiO on GuessWhat?! and InViG benchmarks, setting new state-of-the-art performance by a clear margin. Moreover, we conduct HRI experiments on the carefully selected 150 challenging scenes as well as real-robot platforms. Results show that our method demonstrates superior generality to diversified visual and language inputs with a high success rate. Codes and demos are available at https://github.com/jxu124/TiO.

cs.RO

The discovery of three pulsars in the globular cluster M15 with the FAST

We present the discovery of three pulsars in the Globular Cluster (GC) M15 (NGC 7078) by the Five-hundred-meter Aperture Spherical radio Telescope (FAST). PSR J2129+1210J (M15J) is a millisecond pulsar with a spin period of 11.84 ms and a dispersion measure of 66.68 pc cm-3. Both PSR J2129+1210K and L (M15K and L) are long-period pulsars with spin periods of 1928 ms and 3961 ms, respectively. M15L is the GC pulsar with the longest spin period known. The timing solutions of M15A to M15H are updated. As predicted by Ridolfi et al.(2018), the flux density of M15C keeps decreasing and the latest detection in our dataset was on December 20th, 2022. We have also detected M15I's signal for the first time since its discovery. Current timing suggests that it is an isolated pulsar.

astro-ph.HE

Generative Category-Level Shape and Pose Estimation with Semantic Primitives

Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for object pose estimation are still not satisfactory due to the diversity of object shapes. In this paper, we propose a novel framework for category-level object shape and pose estimation from a single RGB-D image. To handle the intra-category variation, we adopt a semantic primitive representation that encodes diverse shapes into a unified latent space, which is the key to establish reliable correspondences between observed point clouds and estimated shapes. Then, by using a SIM(3)-invariant shape descriptor, we gracefully decouple the shape and pose of an object, thus supporting latent shape optimization of target objects in arbitrary poses. Extensive experiments show that the proposed method achieves SOTA pose estimation performance and better generalization in the real-world dataset. Code and video are available at https://zju3dv.github.io/gCasp.

cs.CV

Tri-Attention: Explicit Context-Aware Attention Mechanism for Natural Language Processing

In natural language processing (NLP), the context of a word or sentence plays an essential role. Contextual information such as the semantic representation of a passage or historical dialogue forms an essential part of a conversation and a precise understanding of the present phrase or sentence. However, the standard attention mechanisms typically generate weights using query and key but ignore context, forming a Bi-Attention framework, despite their great success in modeling sequence alignment. This Bi-Attention mechanism does not explicitly model the interactions between the contexts, queries and keys of target sequences, missing important contextual information and resulting in poor attention performance. Accordingly, a novel and general triple-attention (Tri-Attention) framework expands the standard Bi-Attention mechanism and explicitly interacts query, key, and context by incorporating context as the third dimension in calculating relevance scores. Four variants of Tri-Attention are generated by expanding the two-dimensional vector-based additive, dot-product, scaled dot-product, and bilinear operations in Bi-Attention to the tensor operations for Tri-Attention. Extensive experiments on three NLP tasks demonstrate that Tri-Attention outperforms about 30 state-of-the-art non-attention, standard Bi-Attention, contextual Bi-Attention approaches and pretrained neural language models1.

cs.CL

CAIBC: Capturing All-round Information Beyond Color for Text-based Person Retrieval

Given a natural language description, text-based person retrieval aims to identify images of a target person from a large-scale person image database. Existing methods generally face a \textbf{color over-reliance problem}, which means that the models rely heavily on color information when matching cross-modal data. Indeed, color information is an important decision-making accordance for retrieval, but the over-reliance on color would distract the model from other key clues (e.g. texture information, structural information, etc.), and thereby lead to a sub-optimal retrieval performance. To solve this problem, in this paper, we propose to \textbf{C}apture \textbf{A}ll-round \textbf{I}nformation \textbf{B}eyond \textbf{C}olor (\textbf{CAIBC}) via a jointly optimized multi-branch architecture for text-based person retrieval. CAIBC contains three branches including an RGB branch, a grayscale (GRS) branch and a color (CLR) branch. Besides, with the aim of making full use of all-round information in a balanced and effective way, a mutual learning mechanism is employed to enable the three branches which attend to varied aspects of information to communicate with and learn from each other. Extensive experimental analysis is carried out to evaluate our proposed CAIBC method on the CUHK-PEDES and RSTPReid datasets in both \textbf{supervised} and \textbf{weakly supervised} text-based person retrieval settings, which demonstrates that CAIBC significantly outperforms existing methods and achieves the state-of-the-art performance on all the three tasks.

cs.CV

Look Before You Leap: Improving Text-based Person Retrieval by Learning A Consistent Cross-modal Common Manifold

The core problem of text-based person retrieval is how to bridge the heterogeneous gap between multi-modal data. Many previous approaches contrive to learning a latent common manifold mapping paradigm following a \textbf{cross-modal distribution consensus prediction (CDCP)} manner. When mapping features from distribution of one certain modality into the common manifold, feature distribution of the opposite modality is completely invisible. That is to say, how to achieve a cross-modal distribution consensus so as to embed and align the multi-modal features in a constructed cross-modal common manifold all depends on the experience of the model itself, instead of the actual situation. With such methods, it is inevitable that the multi-modal data can not be well aligned in the common manifold, which finally leads to a sub-optimal retrieval performance. To overcome this \textbf{CDCP dilemma}, we propose a novel algorithm termed LBUL to learn a Consistent Cross-modal Common Manifold (C$^{3}$M) for text-based person retrieval. The core idea of our method, just as a Chinese saying goes, is to `\textit{san si er hou xing}', namely, to \textbf{Look Before yoU Leap (LBUL)}. The common manifold mapping mechanism of LBUL contains a looking step and a leaping step. Compared to CDCP-based methods, LBUL considers distribution characteristics of both the visual and textual modalities before embedding data from one certain modality into C$^{3}$M to achieve a more solid cross-modal distribution consensus, and hence achieve a superior retrieval accuracy. We evaluate our proposed method on two text-based person retrieval datasets CUHK-PEDES and RSTPReid. Experimental results demonstrate that the proposed LBUL outperforms previous methods and achieves the state-of-the-art performance.

cs.CV

Electrolyte Flow Rate Control for Vanadium Redox Flow Batteries using the Linear Parameter Varying Framework

In this article, an electrolyte flow rate control approach is developed for an all-vanadium redox flow battery (VRB) system based on the linear parameter varying (LPV) framework. The electrolyte flow rate is regulated to provide a trade-off between stack voltage efficiency and pumping energy losses, so as to achieve optimal battery energy efficiency. The nonlinear process model is embedded in a linear parameter varying state-space description and a set of state feedback controllers are designed to handle fluctuations in current during both charging and discharging. Simulation studies have been conducted under different operating conditions to demonstrate the performance of the proposed approach. This control approach was further implemented on a laboratory scale VRB system.

eess.SY