Searcharxiv⌕ Search

arXiv subjects

Guo Chen

Publications and source records attributed to Guo Chen.

At least 37 records · Page 2Linked to original sources

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently. This token-by-token decoding mismatches the coupled structure of box geometry and creates a practical inference bottleneck due to strictly sequential generation. We introduce LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). By decoding geometric elements such as bounding boxes and points as atomic units in a single step, LocateAnything preserves intra-box geometric coherence and unlocks substantial parallelism. We show that PBD improves both decoding throughput and localization accuracy. We further develop a scalable data engine and curate LocateAnything-Data, a large-scale dataset with more than 138 million training samples, substantially increasing data diversity for high-precision localization. Extensive evaluations show that LocateAnything advances the speed-accuracy frontier, achieving significantly higher decoding throughput while improving high-IoU localization quality across diverse benchmarks. The results highlight the complementary benefits of Parallel Box Decoding and large-scale training data in enabling efficient and precise unified visual grounding and detection.

cs.CV↗

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The challenge is that base pre-trained models are optimized for next-token prediction and often fail to follow instructions or produce well-formed answers under standard prompting and direct decoding. As a result, benchmark performance can conflate model capability with decoding-induced failures to produce task-oriented outputs, while exposing such behavior often relies on costly post-training. Recent decodingonly approaches attempt to reshape output distributions, but such methods can be inefficient and brittle across open-ended tasks. To address these limitations, we propose Energy-Based Decoding (EBD), a training-free, reward-guided framework for activating task-oriented behaviors from frozen pre-trained LLMs across both open-ended and objective tasks. EBD augments decoding with an external lightweight reward model, steering generations toward high-utility responses while anchoring them to the pre-trained model prior through a reward-tilted target distribution. We show that EBD shifts base-model outputs toward more instructionfollowing behavior, increasing behavioral similarity to post-trained counterparts and enabling a fairer inference-time evaluation of accessible pre-trained-model behavior. Empirically, EBD outperforms baselines across five models and six benchmarks, improving Qwen3-8B-Base on AlpacaEval2.0 from 8.8 to 44.5, reducing Mistral-7B Math500 latency by 18.9x relative to prior decoding work, and remaining robust to reward-model size.

cs.CL↗

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services

Host-GPU data movement has become a latency-critical bottleneck in LLM serving, surfacing in common paths such as model-weight movement and KV cache offload/fetch. Today, each host-GPU copy is effectively confined to the PCIe path of the target GPU, even though modern multi-GPU servers contain additional PCIe links on peer GPUs and high bandwidth GPU interconnects. This leaves substantial intra-server I/O capacity unused. To address this issue, we present Multipath Memory Access (MMA), a software-defined multipath memory access system for host--GPU data transfer. To the best of our knowledge, MMA is the first software-defined system to enable efficient multipath host--GPU data transfer within a single multi-GPU server. MMA expands a single host--GPU copy across available direct and relay paths without hardware, driver, or application changes. It preserves CUDA stream semantics with a dependency-preserving Dummy Task, coordinates distributed micro-transfer completion through a lightweight synchronization mechanism, and uses queue backpressure to route traffic without explicit link-state feedback. On an 8-GPU NVIDIA H20 server, MMA achieves 245 GB/s peak host-to-GPU bandwidth, a 4.62x improvement over native CUDA copies, and reduces TTFT for KV cache fetching by 1.14-2.38x and model wake-up/switching latency by 1.12-2.48x.

cs.DC↗

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.

cs.LG↗

Optimal Decentralized Dynamic Energy Management over Asynchronous Peer-to-Peer Transactive Networks via Operator Splitting

Peer-to-peer (P2P) energy management facilitates decentralized resource allocation among prosumers, improving local hosting capacity for renewables and minimizing energy expenditures while ensuring data privacy through distributed coordination. However, conventional P2P energy management methods are confined to synchronous scheduling paradigms, creating synchronization bottlenecks that fundamentally conflict with the dynamic and decentralized nature of P2P energy management tasks. To bridge this gap, this paper focuses on resolving a class of dynamic energy management problems over asynchronous P2P (Asyn-P2P) transactive networks. We first recast the dynamic energy management problems into a saddle-point problem, and then propose a synchronous decentralized dynamic energy management algorithm, dubbed Syn-DYNA,based on operator splitting theory. To eliminate the global synchronization clock in Syn-DYNA, we introduce a random activation scheme, together with local buffers for latest state tracking, to develop an asynchronous variant of Syn-DYNA, namely Asyn-DYNA. Based on monotone operator theory, theoretical analysis proves a non-asymptotic linear convergence rate for Syn-DYNA and establishes the almost sure convergence ofAsyn-DYNA. Numerical experiments validate effectiveness of Syn-DYNA and Asyn-DYNA algorithms by tackling a dynamic energy management task over P2P transactive networks.

eess.SY↗

Closeby Habitable Exoplanet Survey (CHES). V. Planetary Parameters Derived from Angular Separation Variations

The Closeby Habitable Exoplanet Survey (CHES) aims to achieve microarcsecond-level astrometry of about one hundred nearby FGK-type stars within 10 parsecs to detect Earth-like planets. Such precision exceeds the capability of absolute astrometry relying on Gaia catalogs, whose positional accuracy degrades over time due to error propagation from stellar motion and epoch offsets, limiting their use in microarcsecond-level detection. Traditional relative astrometry depends on positional components along right ascension and declination, requiring precise knowledge of field rotation and satellite attitude, which introduces additional errors. To address this, we propose a new relative measurement model based solely on variations in the length of angular separation between the target and reference stars, independent of direction. The model incorporates effects such as proper motion, parallax, radial velocity, light aberration, gravitational lensing, and planetary perturbations, enabling reconstruction of planetary orbits and masses. This approach enhances measurement stability and precision, providing a framework that is not entirely dependent on the Gaia catalog and suitable for CHES and other future high-accuracy astrometric missions.

astro-ph.EP↗

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

While Video Large Language Models (Video-LLMs) have shown significant potential in multimodal understanding and reasoning tasks, how to efficiently select the most informative frames from videos remains a critical challenge. Existing methods attempt to optimize frame sampling by reducing inter-frame redundancy or employing unsupervised event localization. However, these approaches often fall short in handling complex instruction-following tasks and scenarios that demand precise temporal modeling, resulting in limited performance in both semantic alignment and temporal reasoning. To address the above challenges, we introduce Instructed Temporal Grounding for Videos (VideoITG), a framework aiming to adaptively customize frame sampling strategies based on user instructions. Specifically, we design the VidThinker pipeline, which automates annotation by generating instruction-conditioned captions, retrieving relevant video segments, and selecting key frames to enable efficient supervision. Using VidThinker, we build the VideoITG-40K dataset with 40K videos and 500K temporal grounding annotations. Our plug-and-play VideoITG model leverages Video-LLMs' visual-language alignment and reasoning for discriminative frame selection. VideoITG consistently boosts the performance on multiple multimodal video understanding benchmarks, demonstrating its effectiveness and potential.

cs.CV↗

Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys

The KV cache in self-attention has emerged as a major bottleneck in long-context and large-batch inference for LLMs. Existing approaches often treat sparsity prediction and compression as separate modules, relying on auxiliary index structures to select relevant tokens, and on complex quantization schemes to reduce memory usage. This fragmented design introduces redundant overhead and limits scalability. In this paper, we propose a novel paradigm: treating the compressed key representation not merely as storage, but as a self-indexing structure that directly enables efficient sparse attention. By designing a sign-based 1-bit vector quantization (VQ) scheme, our method unifies compression and retrieval in a single, hardware-friendly format. This approach eliminates the need for external indices or learning-based predictors, offering a lightweight yet robust solution for memory-constrained inference. All components are designed to be hardware-efficient and easy to implement. By implementing custom CUDA kernels, our method integrates seamlessly with FlashAttention, minimizing additional runtime and memory overhead. Experimental results demonstrate that our approach delivers both effectiveness and efficiency.

cs.LG↗

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To bridge this gap, we introduce MM-Lifelong, a dataset designed for Multimodal Lifelong Understanding. Comprising 181.1 hours of footage, it is structured across Day, Week, and Month scales to capture varying temporal densities. Extensive evaluations reveal two critical failure modes in current paradigms: end-to-end MLLMs suffer from a Working Memory Bottleneck due to context saturation, while representative agentic baselines experience Global Localization Collapse when navigating sparse, month-long timelines. To address this, we propose the Recursive Multimodal Agent (ReMA), which employs dynamic memory management to iteratively update a recursive belief state, significantly outperforming existing methods. Finally, we establish dataset splits designed to isolate temporal and domain biases, providing a rigorous foundation for future research in supervised learning and out-of-distribution generalization.

cs.CV↗

Fed-GAME: Personalized Federated Learning with Graph Attention Mixture-of-Experts For Time-Series Forecasting

Federated learning (FL) on graphs shows promise for distributed time-series forecasting. Yet, existing methods rely on static topologies and struggle with client heterogeneity. We propose Fed-GAME, a framework that models personalized aggregation as message passing over a learnable dynamic implicit graph. The core is a decoupled parameter difference-based update protocol, where clients transmit parameter differences between their fine-tuned private model and a shared global model. On the server, these differences are decomposed into two streams: (1) averaged difference used to updating the global model for consensus (2) the selective difference fed into a novel Graph Attention Mixture-of-Experts (GAME) aggregator for fine-grained personalization. In this aggregator, shared experts provide scoring signals while personalized gates adaptively weight selective updates to support personalized aggregation. Experiments on two real-world electric vehicle charging datasets demonstrate that Fed-GAME outperforms state-of-the-art personalized FL baselines.

cs.LG↗

$σ$ bands driven high-temperature superconductivity in hydrogenated hexagonal BC$_3$ monolayer

Material with metallic $σ$-bonding bands is expected to be a high-temperature superconductor, due to the sensitivity of $σ$ electrons to lattice vibration. Based on the first-principles calculations, electronic structures of hydrogenated BC$_3$ monolayers (H$_n$-B$_2$C$_6$ with $n$=1-8) are systematically investigated. At high coverage of hydrogen, the monolayer stabilizes in chair-like $sp^3$-hybridized configurations, leading to the metallization of $σ$ bands, especially in H$_7$-B$_2$C$_6$ and H$_8$-B$_2$C$_6$. This metallicity originates from the electron deficiency of boron, compared with insulating graphane. Utilizing Wannier interpolation, the electron-phonon coupling strengths for metallic phases of H$_n$-B$_2$C$_6$ are determined. As expected, strong couplings are identified between the conducting $σ$ electrons and low-frequency phonon modes. By solving the anisotropic Eliashberg equations, we confirm that H$_7$-B$_2$C$_6$ and H$_8$-B$_2$C$_6$ are single-gap superconductors with critical temperature being 87 K, exceeding the boiling point of liquid nitrogen. Considering that monolayer BC$_3$ has been synthesized in experiment, our results demonstrate that hydrogenation of two-dimensional BC$_3$ provides a viable pathway to achieve high-temperature superconductivity at ambient pressure.

cond-mat.supr-con↗

TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation

In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model with significantly reduced parameters and computational costs: Time-frequency Interleaved Gain Extraction and Reconstruction network (TIGER). TIGER leverages prior knowledge to divide frequency bands and compresses frequency information. We employ a multi-scale selective attention module to extract contextual features while introducing a full-frequency-frame attention module to capture both temporal and frequency contextual information. Additionally, to more realistically evaluate the performance of speech separation models in complex acoustic environments, we introduce a dataset called EchoSet. This dataset includes noise and more realistic reverberation (e.g., considering object occlusions and material properties), with speech from two speakers overlapping at random proportions. Experimental results showed that models trained on EchoSet had better generalization ability than those trained on other datasets compared to the data collected in the physical world, which validated the practical value of the EchoSet. On EchoSet and real-world data, TIGER significantly reduces the number of parameters by 94.3% and the MACs by 95.3% while achieving performance surpassing the state-of-the-art (SOTA) model TF-GridNet.

cs.SD↗

Thermal emission spectra of the ultra-hot Jupiter WASP-33 b

Observations of exoplanetary atmospheres provide critical insights into their chemical composition, formation and evolution history. Ultra-hot Jupiters serve as excellent targets for atmospheric characterization; studies of these planets may yield key understanding of gas giant's formation and evolution history. We present a thermal emission study of WASP-33 b's dayside atmosphere, based on two secondary eclipse observations with CFHT/WIRCam in two specific narrow band filters, namely the CO and CH4$_{\rm on}$ filters, and archival data with HST/WFC3 and Spitzer. Stellar pulsations of the host star induce some quasi-periodic photometric variations, particularly in the CH4$_{\rm on}$ band, which are modelled and corrected in the high-precision differential light curves. An eclipse depth of $1565.2^{+228.6}_{-237.5}$ ppm and $914.3^{+56.1}_{-57.0}$ ppm is determined for the CO and CH4$_{\rm on}$ bands, respectively. Combined with HST/WFC3 and Spitzer data, our joint retrieval of WASP-33 b's dayside atmosphere reveals a high metallicity ([Fe/H] $= 1.52^{+0.35}_{-0.52}$), high C/O ratio (C/O $= 0.78^{+0.03}_{-0.04}$), and a thermal inversion layer, suggesting a formation history involving metal-rich gas accretion. We confirm the presence of the molecules H$_{2}$O, H$^{-}$ and CO, and report a tentative detection of TiO in the dayside atmosphere of WASP-33 b. Future higher precision observations with JWST may provide better understand constraints on the chemical abundances of oxygen and refractory element abundances to better WASP-33 b's formation and evolutionary pathway.

astro-ph.EP↗

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall, pre-training is shifting from more tokens to better tokens. However, existing methods either rely on heuristic static filters that ignore training dynamics, or use dynamic yet optimizer-agnostic criteria based on raw gradients. We propose OPUS (Optimizer-induced Projected Utility Selection), a dynamic data selection framework that defines utility in the optimizer-induced update space. OPUS scores candidates by projecting their effective updates, shaped by modern optimizers, onto a target direction derived from a stable, in-distribution proxy. To ensure scalability, we employ Ghost technique with CountSketch for computational efficiency, and Boltzmann sampling for data diversity, incurring only 4.7\% additional compute overhead. OPUS achieves remarkable results across diverse corpora, quality tiers, optimizers, and model scales. In pre-training of GPT-2 Large/XL on FineWeb and FineWeb-Edu with 30B tokens, OPUS outperforms industrial-level baselines and even full 200B-token training. Moreover, when combined with industrial-level static filters, OPUS further improves pre-training efficiency, even with lower-quality data. Furthermore, in continued pre-training of Qwen3-8B-Base on SciencePedia, OPUS achieves superior performance using only 0.5B tokens compared to full training with 3B tokens, demonstrating significant data efficiency gains in specialized domains.

cs.CL↗

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

Dataset Distillation (DD) seeks to create a compact dataset from a large, real-world dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillation-based dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1\% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18.

cs.LG↗

Efficient Cloud-edge Collaborative Approaches to SPARQL Queries over Large RDF graphs

With the increasing use of RDF graphs, storing and querying such data using SPARQL remains a critical problem. Current mainstream solutions rely on cloud-based data management architectures, but often suffer from performance bottlenecks in environments with limited bandwidth or high system load. To address this issue, this paper explores for the first time the integration of edge computing to move graph data storage and processing to edge environments, thereby improving query performance. This approach requires offloading query processing to edge servers, which involves addressing two challenges: data localization and network scheduling. First, the data localization challenge lies in computing the subgraphs maintained on edge servers to quickly identify the servers that can handle specific queries. To address this challenge, we introduce a new concept of pattern-induced subgraphs. Second, the network scheduling challenge involves efficiently assigning queries to edge and cloud servers to optimize overall system performance. We tackle this by constructing a overall system model that jointly captures data distribution, query characteristics, network communication, and computational resources. Accordingly, we further propose a joint formulation of query assignment and computational resource allocation, modeling it as a Mixed Integer Nonlinear Programming (MINLP) problem and solve this problem using a modified branch-and-bound algorithm. Experimental results on real datasets under a real cloud platform demonstrate that our proposed method outperforms the state-of-the-art baseline methods in terms of efficiency. The codes are available on GitHub

cs.DB↗

Evidence for stellar contamination and water absorption in NGTS-5b's transmission spectra with GTC/OSIRIS

Transmission spectroscopy serves as a valuable tool for probing atmospheric absorption features in the terminator regions of exoplanets. Stellar surface heterogeneity can introduce wavelength-dependent contamination that complicates the interpretation of planetary spectra. We aim to investigate the atmosphere of the warm sub-Saturn NGTS-5b through optical transmission spectroscopy. Two transits were observed with the low-resolution Optical System for Imaging and low-Intermediate-Resolution Integrated Spectroscopy (OSIRIS) on the 10.4 m Gran Telescopio Canarias (GTC). Chromatic transit light curves were modeled to derive optical transmission spectra and multiple Bayesian spectral retrievals were performed to characterize the atmospheric properties. Model comparisons provide strong evidence for contamination from unocculted stellar spots. A joint retrieval of the transmission spectra, assuming equilibrium chemistry, indicates a relatively clear atmosphere with a sub-solar C/O ratio of $<$0.22 (90% upper limit) and a low metallicity of $0.10^{+0.34}_{-0.05} \times$ solar. Retrievals assuming free chemistry yield strong evidence for the presence of $\rm H_2O$, with its abundance constrained to $\log X_{\mathrm{H_2O}} = -0.79^{+0.14}_{-0.17}$. However, the abundances of other species remain unconstrained due to the limited optical wavelength coverage. The discrepancies between the two NGTS-5b transit spectra can be attributed to varying levels of stellar contamination. NGTS-5b thus appears to host a relatively clear, water-rich atmosphere, pending confirmation from additional observations of molecular bands in the infrared.

astro-ph.EP↗

LIR$^3$AG: A Lightweight Rerank Reasoning Strategy Framework for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) effectively enhances Large Language Models (LLMs) by incorporating retrieved external knowledge into the generation process. Reasoning models improve LLM performance in multi-hop QA tasks, which require integrating and reasoning over multiple pieces of evidence across different documents to answer a complex question. However, they often introduce substantial computational costs, including increased token consumption and inference latency. To better understand and mitigate this trade-off, we conduct a comprehensive study of reasoning strategies for reasoning models in RAG multi-hop QA tasks. Our findings reveal that reasoning models adopt structured strategies to integrate retrieved and internal knowledge, primarily following two modes: Context-Grounded Reasoning, which relies directly on retrieved content, and Knowledge-Reconciled Reasoning, which resolves conflicts or gaps using internal knowledge. To this end, we propose a novel Lightweight Rerank Reasoning Strategy Framework for RAG (LiR$^3$AG) to enable non-reasoning models to transfer reasoning strategies by restructuring retrieved evidence into coherent reasoning chains. LiR$^3$AG significantly reduce the average 98% output tokens overhead and 58.6% inferencing time while improving 8B non-reasoning model's F1 performance ranging from 6.2% to 22.5% to surpass the performance of 32B reasoning model in RAG, offering a practical and efficient path forward for RAG systems.

cs.CL↗