Searcharxiv⌕ Search

arXiv subjects

Long Li

Publications and source records attributed to Long Li.

At least 73 records · Page 4Linked to original sources

On existence of a variational regularization parameter under Morozov's discrepancy principle

Morozov's discrepancy principle is commonly adopted in Tikhonov regularization for choosing the regularization parameter. Nevertheless, for a general non-linear inverse problem, the discrepancy $\|F(x_α^δ)-y^δ\|_Y$ does not depend continuously on $α$ and it is questionable whether there exists a regularization parameter $α$ such that $τ_1δ\leq \|F(x_α^δ)-y^δ\|_Y\leq τ_2 δ$ $(1\le τ_1<τ_2)$. In this paper, we prove the existence of $α$ under Morozov's discrepancy principle if $τ_2\ge (3+2γ)τ_1$, where $γ>0$ is a parameter in a tangential cone condition for the nonlinear operator $F$. Furthermore, we present results on the convergence of the regularized solutions under Morozov's discrepancy principle. Numerical results are reported on the efficiency of the proposed approach.

math.NA↗

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions. To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios. Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types. To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation framework with four types of questions and design a novel multi-scale temporal accuracy metric for open-ended temporal evaluation. Based on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems.

cs.CV↗

Hydrogen-poor Superluminous Supernovae with Bumpy Light Curves Powered by Precessing Magnetars

Recent observations and statistical studies have revealed that a significant fraction of hydrogen-poor superluminous supernovae (SLSNe-I) exhibit light curves that deviate from the smooth evolution predicted by the magnetar-powered model, instead showing one or more bumps after the primary peak. However, the formation mechanisms of these post-peak bumps remain a matter of debate. Furthermore, previous studies employing the magnetar-powered model have typically assumed a fixed magnetic inclination angle and neglected the effects of magnetar precession. However, recent research has shown that the precession of newborn magnetars forming during the collapse of massive stars causes the magnetic inclination angle to evolve over time, thereby influencing magnetic dipole radiation. In this paper, therefore, we incorporate the effects of magnetar precession into the magnetar-powered model to develop the precessing magnetar-powered model. Using this model, we successfully reproduce the multi-band light curves of 6 selected representative SLSNe-I with post-peak bumps. Moreover, the derived model parameters fall within the typical parameter range for SLSNe-I. By combining the precessing magnetars in SLSNe-I and long GRBs, we find that the ellipticity of magnetars is related to the dipole magnetic field strength, which may suggest a common origin for the two phenomena. Our work provides a potential explanation for the origin of post-peak bumps in SLSNe-I and offers evidence for the early precession of newborn magnetars formed in supernova explosions.

astro-ph.HE↗

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a large-scale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities.

cs.CV↗

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code are available at https://github.com/Rh-Dang/ECBench.

cs.CV↗

Charge calibration of MALTA2, a radiation hard depleted monolithic active pixel sensor

MALTA2 is a depleted monolithic active pixel sensor (DMAPS) designed for tracking at high rates and typically low detection threshold of $\sim150\,\mathrm{e^-}$. A precise knowledge of the threshold is crucial to understanding the charge collection in the pixel and specifying the environment for sensor application. A simple procedure is developed to calibrate the threshold to unit electrons making use of a dedicated charge injection circuit and an Fe-55 source with dominant charge deposition of $1600\, \mathrm{e^-}$. The injection voltage is determined which corresponds to the injection under Fe-55 exposure and is the basis for charge calibration. The charge injection circuit incorporates a capacitance with design value of $\mathrm{C_{inj}}=$ 230 aF. Experimentally, the average capacitance value for non-irradiated samples is found to be $\mathrm{C_{inj,exp}}=$ 257 aF. The 12 % divergence motivates the need for the presented precise calibration procedure, which is proposed to be performed for each MALTA2 sensor.

physics.ins-det↗

Signature of Triaxially Precessing Magnetars in Gamma-ray Burst X-Ray Afterglows

The X-ray afterglows of some gamma-ray bursts (GRBs) exhibit plateaus, which can be explained by the internal dissipation of a newborn millisecond magnetar wind. In the early phase of these newborn magnetars, the magnetic inclination angle undergoes periodic changes due to precession, leading to periodic modulation of the injection luminosity due to magnetic dipole radiation. This may result in quasi-periodic oscillations (QPOs) on the plateaus. In this paper, we identify four GRBs with regular flux variations on their X-ray afterglow plateaus from Swift/XRT data before November 2023, three of which exhibit periodicity. Based on the likelihood of supporting a precessing magnetar as the central engine, we classify them into three categories: Gold (GRB 060202 and GRB 180620A), Silver (GRB 050730), and Bronze (GRB 210610A). We invoke a model of magnetic dipole radiation emitted by a triaxially freely precessing magnetar whose spin-down is dominated by electromagnetic radiation, to fit the light curves. Our model successfully reproduces the light curves of these four GRBs, including the regular flux variations on the plateaus and their periodicity (if present). Our work provides further evidence for early precession in newborn millisecond magnetars in GRBs.

astro-ph.HE↗

Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Effective research ideation is a critical step for scientific research. However, the exponential increase in scientific literature makes it challenging for researchers to stay current with recent advances and identify meaningful research directions. Recent developments in large language models~(LLMs) suggest a promising avenue for automating the generation of novel research ideas. However, existing methods for idea generation either trivially prompt LLMs or directly expose LLMs to extensive literature without indicating useful information. Inspired by the research process of human researchers, we propose a Chain-of-Ideas~(CoI) agent, an LLM-based agent that organizes relevant literature in a chain structure to effectively mirror the progressive development in a research domain. This organization facilitates LLMs to capture the current advancements in research, thereby enhancing their ideation capabilities. Furthermore, we propose Idea Arena, an evaluation protocol that can comprehensively evaluate idea generation methods from different perspectives, aligning closely with the preferences of human researchers. Experimental results indicate that the CoI agent consistently outperforms other methods and shows comparable quality as humans in research idea generation. Moreover, our CoI agent is budget-friendly, with a minimum cost of \$0.50 to generate a candidate idea and its corresponding experimental design.

cs.AI↗

How Do Humans Write Code? Large Models Do It the Same Way Too

Program-of-Thought (PoT) replaces natural language-based Chain-of-Thought (CoT) as the most popular method in Large Language Models (LLMs) mathematical reasoning tasks by utilizing external tool calls to circumvent computational errors. However, our evaluation of the GPT-4 and Llama series reveals that using PoT introduces more reasoning errors, such as incorrect formulas or flawed logic, compared to CoT. To address this issue, we propose Human-Think Language (HTL), which leverages a suite of strategies that help integrate PoT and CoT, encompassing: (1) a new generation paradigm that uses full CoT reasoning to control code generation. (2) Focus Attention, that directs model attention to the CoT reasoning during PoT to generate more logical code. (3) reinforcement learning that utilizes the accuracy of both CoT and PoT responses as rewards to prevent repetitive reasoning steps in LLMs when solving difficult math problems. Our method achieves an average improvement of 6.5% on the Llama-Base model and 4.3% on the Mistral-Base model across 8 mathematical calculation datasets. It also shows significant effectiveness on five out-of-domain datasets by controlling the model's information flow, exhibiting strong transferability. Additionally, HTL shows the most significant improvement in non-mathematical natural language inference task, contributing to a unified reasoning task framework

cs.AI↗

CONDA: Condensed Deep Association Learning for Co-Salient Object Detection

Inter-image association modeling is crucial for co-salient object detection. Despite satisfactory performance, previous methods still have limitations on sufficient inter-image association modeling. Because most of them focus on image feature optimization under the guidance of heuristically calculated raw inter-image associations. They directly rely on raw associations which are not reliable in complex scenarios, and their image feature optimization approach is not explicit for inter-image association modeling. To alleviate these limitations, this paper proposes a deep association learning strategy that deploys deep networks on raw associations to explicitly transform them into deep association features. Specifically, we first create hyperassociations to collect dense pixel-pair-wise raw associations and then deploys deep aggregation networks on them. We design a progressive association generation module for this purpose with additional enhancement of the hyperassociation calculation. More importantly, we propose a correspondence-induced association condensation module that introduces a pretext task, i.e. semantic correspondence estimation, to condense the hyperassociations for computational burden reduction and noise elimination. We also design an object-aware cycle consistency loss for high-quality correspondence estimations. Experimental results in three benchmark datasets demonstrate the remarkable effectiveness of our proposed method with various training settings.

cs.CV↗

Identifying the Origin of FRB-associated X-ray Bursts with X-ray Polarization

The origin of extraordinary X-ray burst (XRB) associated with a fast radio burst (FRB) like FRB 20200428D is still unclear, though several models such as the emission of a trapped fireball modified by resonant cyclotron scattering, the outflow from a polar trapped-expanding fireball, and the synchrotron radiation of a far-away relativistic shock, have been proposed. To determine which model is true, we study possible X-ray polarization signature for each model, inspired by the importance of radio polarization in identifying FRB origin. We first numerically simulate or calculate the XRB spectrum for each model and fit it to the observed data, then compute the corresponding polarization signal based on the fit. We find that these three models predict different polarization patterns in terms of phase/time and energy variations. The differences can be used to test the models with future X-ray polarization observations.

astro-ph.HE↗

Symbolic Learning Enables Self-Evolving Agents

The AI community has been exploring a pathway to artificial general intelligence (AGI) by developing "language agents", which are complex large language models (LLMs) pipelines involving both prompting techniques and tool usage methods. While language agents have demonstrated impressive capabilities for many real-world tasks, a fundamental limitation of current language agents research is that they are model-centric, or engineering-centric. That's to say, the progress on prompts, tools, and pipelines of language agents requires substantial manual engineering efforts from human experts rather than automatically learning from data. We believe the transition from model-centric, or engineering-centric, to data-centric, i.e., the ability of language agents to autonomously learn and evolve in environments, is the key for them to possibly achieve AGI. In this work, we introduce agent symbolic learning, a systematic framework that enables language agents to optimize themselves on their own in a data-centric way using symbolic optimizers. Specifically, we consider agents as symbolic networks where learnable weights are defined by prompts, tools, and the way they are stacked together. Agent symbolic learning is designed to optimize the symbolic network within language agents by mimicking two fundamental algorithms in connectionist learning: back-propagation and gradient descent. Instead of dealing with numeric weights, agent symbolic learning works with natural language simulacrums of weights, loss, and gradients. We conduct proof-of-concept experiments on both standard benchmarks and complex real-world tasks and show that agent symbolic learning enables language agents to update themselves after being created and deployed in the wild, resulting in "self-evolving agents".

cs.CL↗

Unveiling the Impact of Multi-Modal Interactions on User Engagement: A Comprehensive Evaluation in AI-driven Conversations

Large Language Models (LLMs) have significantly advanced user-bot interactions, enabling more complex and coherent dialogues. However, the prevalent text-only modality might not fully exploit the potential for effective user engagement. This paper explores the impact of multi-modal interactions, which incorporate images and audio alongside text, on user engagement in chatbot conversations. We conduct a comprehensive analysis using a diverse set of chatbots and real-user interaction data, employing metrics such as retention rate and conversation length to evaluate user engagement. Our findings reveal a significant enhancement in user engagement with multi-modal interactions compared to text-only dialogues. Notably, the incorporation of a third modality significantly amplifies engagement beyond the benefits observed with just two modalities. These results suggest that multi-modal interactions optimize cognitive processing and facilitate richer information comprehension. This study underscores the importance of multi-modality in chatbot design, offering valuable insights for creating more engaging and immersive AI communication experiences and informing the broader AI community about the benefits of multi-modal interactions in enhancing user engagement.

cs.CL↗

All-sky Guide Star Catalog for CSST

The China Space Station Telescope (CSST) is a two-meter space telescope with multiple back-end instruments. The Fine Guidance Sensor (FGS) is an essential subsystem of the CSST Precision Image Stability System to ensure the required absolute pointing accuracy and line-of-sight stabilization. In this study, we construct the Main Guide Star Catalog for FGS. To accomplish this, we utilize the information about the FGS and object information from the Gaia Data Release 3. We provide an FGS instrument magnitude and exclude variables, binaries, and high proper motion stars from the catalog to ensure uniform FGS guidance capabilities. Subsequently, we generate a HEALPix index, which provides a hierarchical tessellation of the celestial sphere, and employ the Voronoi algorithm to achieve a homogeneous distribution of stars across the catalog. This distribution ensures adequate coverage and sampling of the sky. The performance of the CSST guide star catalog was assessed by simulating the field of view of the FGS according to the CSST mock survey strategy catalog. The analysis of the results indicates that this catalog provides adequate coverage and accuracy. The catalog's performance meets the FGS requirements, ensuring the functioning of the FGS and its guidance capabilities.

astro-ph.IM↗

SN 2019tua : A Type IIb Supernova with Multiple Bumps in the Light Curves

We present photometric and spectroscopic observations and analysis of the type IIb supernova (SN) SN 2019tua, which exhibits multiple bumps in its declining light curves between 40 and 65 days after discovery. SN 2019tua shows a time to peak of about 25 days similar to other type IIb SNe. Our observations indicate a decrease in its brightness of about 1 magnitude in the 60 days after the peak. At about days 50, and 60, its multiband light curves exhibit bumpy behavior. The complex luminosity evolution of SN 2019tua could not be well modeled with a single currently popular energy source model, e.g., radioactive decay of $^{56}$Ni, magnetar, interaction between the ejecta and a circumstellar shell. Even though the magnetar model has a smaller $χ^2 / \text{dof}$ value, the complex changes in SN 2019tua's brightness suggest that more than one physical process might be involved. We propose a hybrid CSM interaction plus $^{56}$Ni model to explain the bolometric light curve (LC) of SN 2019tua. The fitting results show that the ejecta mass $M_{\rm ej} \approx 2.4~M_\odot$, the total CSM mass $M_{\rm CSM} \approx 1.0~M_\odot$, and the $^{56}$Ni mass $M_{\rm Ni} \approx 0.4~M_\odot$. The total kinetic energy of the ejecta is $E_k\approx 0.5 \times 10^{51}\rm~erg$. Pre-existing multiple shells suggest that the progenitor of SN 2019tua experienced mass ejections within approximately $\sim6 - 44$ years prior to the explosion.

astro-ph.HE↗

Opening Gaps in the Spectrum of Strictly Ergodic Jacobi and CMV Matrices

We prove that dynamically defined Jacobi and CMV matrices associated with generic continuous sampling functions have all gaps predicted by the Gap Labelling Theorem open. We also give a mechanism for generic gap opening for quasi-periodic analytic sampling functions in the subcritical region following from the analyticity of resonance tongue boundaries for both Jacobi and CMV matrices.

math.SP↗

AT2023lli: A Tidal Disruption Event with Prominent Optical Early Bump and Delayed Episodic X-ray Emission

High-cadence, multiwavelength observations have continuously revealed the diversity of tidal disruption events (TDEs), thus greatly advancing our knowledge and understanding of TDEs. In this work, we conducted an intensive optical-UV and X-ray follow-up campaign of TDE AT2023lli, and found a remarkable month-long bump in its UV/optical light curve nearly two months prior to maximum brightness. The bump represents the longest separation time from the main peak among known TDEs to date. The main UV/optical outburst declines as $t^{-4.10}$, making it one of the fastest decaying optically selected TDEs. Furthermore, we detected sporadic X-ray emission 30 days after the UV/optical peak, accompanied by a reduction in the period of inactivity. It is proposed that the UV/optical bump could be caused by the self-intersection of the stream debris, whereas the primary peak is generated by the reprocessed emission of the accretion process. In addition, our results suggest that episodic X-ray radiation during the initial phase of decline may be due to the patched obscurer surrounding the accretion disk, a phenomenon associated with the inhomogeneous reprocessing process. The double TDE scenario, in which two stars are disrupted in sequence, is also a possible explanation for producing the observed early bump and main peak. We anticipate that the multicolor light curves of TDEs, especially in the very early stages, and the underlying physics can be better understood in the near future with the assistance of dedicated surveys such as the deep high-cadence survey of the 2.5-meter Wide Field Survey Telescope (WFST).

astro-ph.HE↗