SearcharxivSearch

arXiv subjects

Guilin Liu

Publications and source records attributed to Guilin Liu.

At least 19 recordsLinked to original sources

Energy Partition in AGN-driven Bubbles of NGC 4438: From Nuclear Bubbles to a Galaxy-scale Outflow

Jets launched by accreting supermassive black holes represent a major mode of active galactic nucleus (AGN) feedback. However, how their energy is divided among bulk kinetic motion, thermal gas, magnetic fields, cosmic rays (CRs), and radiation - and how this distribution changes with spatial scale - remains poorly constrained. NGC 4438 provides a unique laboratory for probing this evolution, hosting two 200-pc-scale nuclear bubbles and a lopsided ~10 kpc galaxy-scale outflow plausibly associated with the same AGN. We present a multi-wavelength analysis to investigate the morphology, radiation mechanisms, and energetics of these structures. Joint radio-X-ray modeling shows that the non-thermal emission in the nuclear bubbles may require two distinct populations of cosmic-ray electrons, suggesting that in addition to shock acceleration at the bubble rim, the highest-energy particles may be linked to acceleration processes closer to the unresolved central engine. A spatially resolved energy inventory reveals that bulk kinetic energy dominates the current energy budget of the nuclear bubbles, while roughly half of the injected energy has already been transformed into thermal, CR, and magnetic energy, as well as radiative losses. Across all bubble sizes, the thermal and magnetic pressures are consistent within the uncertainties, implying that magnetic fields remain dynamically significant on all examined spatial scales. Furthermore, the empirical correlation between radio luminosity and jet power, established for kiloparsec-scale jet bubbles (MerloniHeinz2007), matches the energetics of the galaxy-scale outflow but substantially overestimates the power of the 200-pc-scale nuclear bubbles, underscoring the scale dependence of jet energy dissipation.

astro-ph.HE

Measuring Outflow Distances in NGC 5548 Using Absorption-Line Variability Diagnostics

AGN-driven outflows serve as a key channel through which the energetic central engine influences host galaxy evolution. Among the physical properties of outflows, their radial distance from the galactic nucleus is particularly important for assessing AGN feedback. In this study, we investigate the UV outflow components in NGC 5548 by analyzing the variability of C IV absorption troughs in optical spectra obtained through multiple HST observations during 2013 and 2014. We construct a set of variability-based diagnostic events, labeled G1 and G2, which are sensitive to the recombination timescale ($t_r$) of ionized gas. By combining these with mock light curves generated using a damped random walk (DRW) model, we numerically establish a mapping between the G1 event probability and $t_r$. This approach allows us to constrain the radial distances of outflow components 1 and 6, whose absorption variability is primarily driven by changes in the incident ionizing continuum, to be $0.77^{+0.10}_{-0.10}$ and $1.72^{+1.74}_{-1.72}$ pc, respectively. These results are consistent with those obtained using a different method in our previous study, as well as with values reported in the literature.

astro-ph.GA

Dissecting the multiple-component outflow in NGC 5548 with absorption-line Variability

AGN-driven outflows are routinely invoked as a key agent of supermassive black holes to regulate the evolution of galaxies. The radial distance from the central engine is a crucial parameter for evaluating the impact of these outflows on the host galaxy. In this work, we estimate the radial distances of ultraviolet (UV) outflow components in NGC 5548 using the most up-to-date absorption-line variability method, combined with multi-epoch HST/COS spectroscopy from the 2014 AGN STORM campaign and archival data observed in 2013. The recombination timescale (tr) of the absorbers are measured by analyzing the detection rate curves of absorption-line variability. In particular, the detection rate curves of the absorption troughs showing blended multiple velocity components are featured by distinct ``multi-step' profiles, allowing for measuring tr for individual components. Among the 6 identified outflow components, four are found to be a few pc from the center and two are 30-40 pc away. Our results agree well with the more reliable results in the literature on components 1 and 4, and show overall consistency with previous works, demonstrating the power of our new methodology especially when it is aided by densely sampled HST spectra.

astro-ph.GA

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently. This token-by-token decoding mismatches the coupled structure of box geometry and creates a practical inference bottleneck due to strictly sequential generation. We introduce LocateAnything, a unified generative grounding and detection framework based on Parallel Box Decoding (PBD). By decoding geometric elements such as bounding boxes and points as atomic units in a single step, LocateAnything preserves intra-box geometric coherence and unlocks substantial parallelism. We show that PBD improves both decoding throughput and localization accuracy. We further develop a scalable data engine and curate LocateAnything-Data, a large-scale dataset with more than 138 million training samples, substantially increasing data diversity for high-precision localization. Extensive evaluations show that LocateAnything advances the speed-accuracy frontier, achieving significantly higher decoding throughput while improving high-IoU localization quality across diverse benchmarks. The results highlight the complementary benefits of Parallel Box Decoding and large-scale training data in enabling efficient and precise unified visual grounding and detection.

cs.CV

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.

cs.LG

CHANG-ES. XXXVIII. A Thin Radio Halo Shaped by Slow Cosmic-Ray Transport in the Quiescent Galaxy NGC 4565

We present the VLA C-array S-band (2--4 GHz) radio continuum observations of the nearby edge-on spiral galaxy NGC 4565, a target from the Continuum Halos in Nearby Galaxies - an EVLA (CHANG-ES) Survey. We conduct rotation measure synthesis to probe the magnetic field structure and analyze the vertical radio continuum intensity profiles using the 1-D cosmic ray transportation models. The radio continuum emission of NGC 4565 is vertically compact, with a vertical-to-radial extent ratio of $\sim 1/6$. Its vertical profile is optimally described by a two-component Gaussian distribution, yielding a mean Gaussian halo scale height of $\sim 3.0$ kpc. The magnetic field is weak, predominantly disk-parallel, with an equipartition strength of $\lesssim 5\ \mu$G and a rotation measure profile indicative of an axisymmetric spiral structure. Nevertheless, we identify a localized, faint vertical magnetic field component in the northeastern region, hinting at an X-shaped structure that spatially coincides with extraplanar structures detected in H I and soft X-ray emission. The CR transport modeling favors a flux-tube advection scenario, with a slow initial velocity of $v_0 \approx 60$ km s$^{-1}$, consistent with a limited energy input from star formation. Therefore, the absence of an extended radio halo can be explained by the low star formation rate, the weak magnetic field, and the inefficient CR transport. The localized X-shaped field may trace a weak, magnetically guided outflow or a tidal perturbation induced by the nearby companion. NGC 4565 is thus a key quiescent benchmark for understanding the physical conditions required to drive large-scale outflows and generate extended radio halos.

astro-ph.GA

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To bridge this gap, we introduce MM-Lifelong, a dataset designed for Multimodal Lifelong Understanding. Comprising 181.1 hours of footage, it is structured across Day, Week, and Month scales to capture varying temporal densities. Extensive evaluations reveal two critical failure modes in current paradigms: end-to-end MLLMs suffer from a Working Memory Bottleneck due to context saturation, while representative agentic baselines experience Global Localization Collapse when navigating sparse, month-long timelines. To address this, we propose the Recursive Multimodal Agent (ReMA), which employs dynamic memory management to iteratively update a recursive belief state, significantly outperforming existing methods. Finally, we establish dataset splits designed to isolate temporal and domain biases, providing a rigorous foundation for future research in supervised learning and out-of-distribution generalization.

cs.CV

Stateful Token Reduction for Long-Video Hybrid VLMs

Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated as token pruning. We study token reduction in hybrid Mamba--Transformer VLMs and find that it is \emph{stateful}: Mamba layers maintain a recurrent state that accumulates information from earlier tokens, allowing discarded tokens to persist, so reduction behaves more like compression than dropping.We support this view with a representation-based probing method measuring how much information from discarded tokens is retained, and analyze layer-wise sparsity and cross-layer importance stability. Our findings show importance is sparse within layers but unstable across layers, making aggressive early pruning unreliable while hybrids remain robust to later reduction.Motivated by this, we propose a hybrid-aware token reduction framework with a low-to-high progressive schedule and a unified query-conditioned importance score for attention and Mamba layers. For Mamba, excluding the position-dependent decay from the recurrence produces a stronger selection signal. Across long-video benchmarks, our method achieves $3.8{\times}$--$4.2{\times}$ prefilling speedups at a 25% token budget while maintaining near-baseline accuracy and improving with light finetuning. Hybrid models benefit from aggressive reduction, improving both efficiency and accuracy, whereas Transformers exhibit the standard trade-off. Our method also outperforms prior baselines on the same hybrid backbone and combines effectively with visual redundancy reduction methods.

cs.CV

PhyCritic: Multimodal Critic Models for Physical AI

With the rapid development of large multimodal models, reliable judge and critic models have become essential for open-ended evaluation and preference alignment, providing pairwise preferences, numerical scores, and explanatory justifications for assessing model-generated responses. However, existing critics are primarily trained in general visual domains such as captioning or image question answering, leaving physical AI tasks involving perception, causal reasoning, and planning largely underexplored. We introduce PhyCritic, a multimodal critic model optimized for physical AI through a two-stage RLVR pipeline: a physical skill warmup stage that enhances physically oriented perception and reasoning, followed by self-referential critic finetuning, where the critic generates its own prediction as an internal reference before judging candidate responses, improving judgment stability and physical correctness. Across both physical and general-purpose multimodal judge benchmarks, PhyCritic achieves strong performance gains over open-source baselines and, when applied as a policy model, further improves perception and reasoning in physically grounded tasks.

cs.CV

ELFO: A Python Package for Emission Line Fitting Optimization in Integral Field Spectroscopy Data

Integral field spectroscopy (IFS) provides spatially resolved spectra, enabling detailed studies that address the physical and kinematic properties of the interstellar medium. A critical step in analyzing IFS data is the decomposition of emission lines, where different velocity components are often modeled with Gaussian profiles. However, conventional fitting methods that treat each spectrum independently often yield spatial discontinuities in the fitting results. Here, we present Emission Line Fitting Optimization (ELFO), a Python package for IFS spectral fitting. ELFO uses the results of neighboring spectra to determine multiple initial guesses and selects the result that exhibits spatial smoothness. We tested ELFO on IFS data of two quasars obtained from the Multi-Unit Spectroscopic Explorer, where it successfully corrected anomalous fits, revealed previously unresolved substructures, and made large-scale kinematic structures more evident. With minor modifications, this method can also be easily adapted to other IFS data and different emission lines.

astro-ph.GA

OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation

This paper presents a family of advanced vision encoder, named OpenVision 3, that learns a single, unified visual representation that can serve both image understanding and image generation. Our core architecture is simple: we feed VAE-compressed image latents to a ViT encoder and train its output to support two complementary roles. First, the encoder output is passed to the ViT-VAE decoder to reconstruct the original image, encouraging the representation to capture generative structure. Second, the same representation is optimized with contrastive learning and image-captioning objectives, strengthening semantic features. By jointly optimizing reconstruction- and semantics-driven signals in a shared latent space, the encoder learns representations that synergize and generalize well across both regimes. We validate this unified design through extensive downstream evaluations with the encoder frozen. For generation, we test it under the RAE framework: ours substantially surpasses the standard CLIP-based encoder (e.g., gFID: 1.87 vs. 2.54 on ImageNet). For multimodal understanding, we plug the encoder into the LLaVA-1.5 and LLaVA-NeXT framework: it performs comparably with a standard CLIP vision encoder (e.g., 63.3 vs. 61.2 on SeedBench, and 59.2 vs. 58.1 on GQA). We provide empirical evidence that generation and understanding are mutually beneficial in our architecture, while further underscoring the critical role of the VAE latent space. We hope this work can spur future research on unified modeling.

eess.IV

NVIDIA Nemotron Nano V2 VL

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant improvements over our previous model, Llama-3.1-Nemotron-Nano-VL-8B, across all vision and text domains through major enhancements in model architecture, datasets, and training recipes. Nemotron Nano V2 VL builds on Nemotron Nano V2, a hybrid Mamba-Transformer LLM, and innovative token reduction techniques to achieve higher inference throughput in long document and video scenarios. We are releasing model checkpoints in BF16, FP8, and FP4 formats and sharing large parts of our datasets, recipes and training code.

cs.LG

STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Recent advances in video-based multimodal large language models (Video-LLMs) have significantly improved video understanding by processing videos as sequences of image frames. However, many existing methods treat frames independently in the vision backbone, lacking explicit temporal modeling, which limits their ability to capture dynamic patterns and efficiently handle long videos. To address these limitations, we introduce STORM (Spatiotemporal TOken Reduction for Multimodal LLMs), a novel architecture incorporating a dedicated temporal encoder between the image encoder and the LLM. Our temporal encoder leverages the Mamba State Space Model to integrate temporal information into image tokens, generating enriched representations that preserve inter-frame dynamics across the entire video sequence. This enriched encoding not only enhances video reasoning capabilities but also enables effective token reduction strategies, including test-time sampling and training-based temporal and spatial pooling, substantially reducing computational demands on the LLM without sacrificing key temporal information. By integrating these techniques, our approach simultaneously reduces training and inference latency while improving performance, enabling efficient and robust video understanding over extended temporal contexts. Extensive evaluations show that STORM achieves state-of-the-art results across various long video understanding benchmarks (more than 5% improvement on MLVU and LongVideoBench) while reducing the computation costs by up to $8\times$ and the decoding latency by 2.4-2.9$\times$ for the fixed numbers of input frames. Project page is available at https://research.nvidia.com/labs/lpr/storm

cs.CV

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer models designed to reduce inference cost for a given accuracy level. To achieve this goal, we replace the majority of self-attention layers in the common Transformer model architecture with Mamba layers that perform constant computation and require constant memory per generated token. We show that Nemotron-H models offer either better or on-par accuracy compared to other similarly-sized state-of-the-art open-sourced Transformer models (e.g., Qwen-2.5-7B/72B and Llama-3.1-8B/70B), while being up to 3$\times$ faster at inference. To further increase inference speed and reduce the memory required at inference time, we created Nemotron-H-47B-Base from the 56B model using a new compression via pruning and distillation technique called MiniPuzzle. Nemotron-H-47B-Base achieves similar accuracy to the 56B model, but is 20% faster to infer. In addition, we introduce an FP8-based training recipe and show that it can achieve on par results with BF16-based training. This recipe is used to train the 56B model. We are releasing Nemotron-H base model checkpoints with support in Hugging Face and NeMo.

cs.CL

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

While Video Large Language Models (Video-LLMs) have shown significant potential in multimodal understanding and reasoning tasks, how to efficiently select the most informative frames from videos remains a critical challenge. Existing methods attempt to optimize frame sampling by reducing inter-frame redundancy or employing unsupervised event localization. However, these approaches often fall short in handling complex instruction-following tasks and scenarios that demand precise temporal modeling, resulting in limited performance in both semantic alignment and temporal reasoning. To address the above challenges, we introduce Instructed Temporal Grounding for Videos (VideoITG), a framework aiming to adaptively customize frame sampling strategies based on user instructions. Specifically, we design the VidThinker pipeline, which automates annotation by generating instruction-conditioned captions, retrieving relevant video segments, and selecting key frames to enable efficient supervision. Using VidThinker, we build the VideoITG-40K dataset with 40K videos and 500K temporal grounding annotations. Our plug-and-play VideoITG model leverages Video-LLMs' visual-language alignment and reasoning for discriminative frame selection. VideoITG consistently boosts the performance on multiple multimodal video understanding benchmarks, demonstrating its effectiveness and potential.

cs.CV

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Recent advancements in reasoning-focused language models such as OpenAI's O1 and DeepSeek-R1 have shown that scaling test-time computation-through chain-of-thought reasoning and iterative exploration-can yield substantial improvements on complex tasks like mathematics and code generation. These breakthroughs have been driven by large-scale reinforcement learning (RL), particularly when combined with verifiable reward signals that provide objective and grounded supervision. In this report, we investigate the effects of prolonged reinforcement learning on a small language model across a diverse set of reasoning domains. Our work identifies several key ingredients for effective training, including the use of verifiable reward tasks, enhancements to Group Relative Policy Optimization (GRPO), and practical techniques to improve training stability and generalization. We introduce controlled KL regularization, clipping ratio, and periodic reference policy resets as critical components for unlocking long-term performance gains. Our model achieves significant improvements over strong baselines, including +14.7% on math, +13.9% on coding, and +54.8% on logic puzzle tasks. To facilitate continued research, we release our model publicly.

cs.LG

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed for accurate reasoning. In this paper, we introduce Argus to address these limitations with a new visual attention grounding mechanism. Our approach employs object-centric grounding as visual chain-of-thought signals, enabling more effective goal-conditioned visual attention during multimodal reasoning tasks. Evaluations on diverse benchmarks demonstrate that Argus excels in both multimodal reasoning tasks and referring object grounding tasks. Extensive analysis further validates various design choices of Argus, and reveals the effectiveness of explicit language-guided visual region-of-interest engagement in MLLMs, highlighting the importance of advancing multimodal intelligence from a visual-centric perspective. Project page: https://yunzeman.github.io/argus/

cs.CV

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Enabling large language models with external tools has become a pivotal strategy for extending their functionality beyond text space. To enhance LLMs' tool-calling abilities, previous approaches primarily rely on supervised fine-tuning (SFT) with trajectories distilled from stronger models, often resulting in imitative reasoning that limits generalization. In this work, we explore rule-based reinforcement learning to enhance tool-calling in LLMs, resulting in Nemotron-Research-Tool-N1, a series of tool-calling reasoning models. Rather than enforcing supervision over intermediate distilled reasoning traces, Tool-N1 is trained with a binary RL reward that assesses only the format validity and functional correctness of tool invocations. This lightweight supervision allows the model to develop reasoning strategies independently, without relying on annotated trajectories. Experiments on several major benchmarks show that Tool-N1-7B/14B clearly outperform GPT-4o. We conduct a systematic study on the design of rule-based reinforcement learning strategies for training tool-calling models. Using 5,518 distilled reasoning trajectories, we compare SFT, RL, and the SFT-then-RL pipeline, finding that the widely adopted SFT-then-RL paradigm does not necessarily outperform pure RL.

cs.CL