SearcharxivSearch

arXiv subjects

Xinyu Dai

Publications and source records attributed to Xinyu Dai.

At least 19 recordsLinked to original sources

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.

cs.AI

The X-Ray Continuum Emission Region in the Lensed Quasar SDSS J133907.23+131038.6 is Much Smaller than the Accretion Disk

We analyze microlensing variability in 15 seasons of optical monitoring data and 4 epochs of new X-ray observations of the doubly-imaged gravitationally lensed quasar SDSS J133907.23+131038.6 to place empirical constraints on the size and structure of that system's X-ray and optical continuum emission regions. Employing a Bayesian Monte Carlo method, we analyzed ground-based optical light curves to constrain the half-light radius of the far-UV source $\log(r_{\rm 1/2, FUV}/{\rm cm})=15.78^{+0.26}_{-0.28}$ at 193 nm, the rest-frame center of the {\it r}-band, assuming a $60^\circ$ inclination angle. This size corresponds to $\sim100\,{\it r}_{\rm g}$ for a $4.0 \times 10^{8} \: {\rm M_{\odot}}$ black hole. We measured the half-light radius of the full band ($0.2-8.0 \: {\rm keV}$) X-ray continuum emission region $\log(r_{\rm 1/2, X_{full}}/{\rm cm})=14.32^{+0.23}_{-0.31}$, a size measurement that is consistent with the radius of the innermost stable circular orbit (ISCO) in the Schwarzschild metric.Two shifted Fe K$\alpha$ lines caused by microlensing are detected in the stacked spectrum of image A at 5.9 and 8.9~keV at $>99\%$ significance.

astro-ph.HE

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.

cs.CV

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.

cs.CL

The Twentieth Data Release of the Sloan Digital Sky Survey: First All-Sky BOSS Spectra, eROSITA-SDSS-V Mapper Coordinated Observations, and a Preview of the Local Volume Mapper

This paper presents the twentieth data release (DR20) from the Sloan Digital Sky Survey, the third data release of its fifth generation (SDSS-V). SDSS-V is a panoptic spectroscopy survey that is mapping the stars, gas, and galaxies through three scientific programs: the Milky Way Mapper (MWM), the Local Volume Mapper (LVM), and the Black Hole Mapper (BHM). DR20 presents the first optical (BOSS) SDSS-V spectra from southern hemisphere for the MWM and BHM surveys; new optical MWM and BHM data from the northern hemisphere are also available, for a total over 3 million spectra of 1.5 million stars and half a million galaxies and quasars, with galactic and extragalactic x-ray targets coordinate with eROSITA DR2. DR20 includes integral field spectroscopy maps from LVM of six targets and 169 tiles, spanning Galactic HII regions, planetary nebulae, and nearby galaxies. Additionally, eighteen value added catalogs are also released with DR20, based on SDSS-V MWM and BHM data, and we present a new LVM visualization tool including an RGB HiPS map as a value added product.

astro-ph.GA

Spectroscopic Analysis of Fermi-detected Blazars using SDSS-V

The automated spectroscopic pipeline of the Sloan Digital Sky Survey (SDSS) systematically assigns Galactic star, galaxy, or typical quasar classifications to jet-dominated blazars, owing to the absence of a non-thermal jet continuum component in its template library. In this study, we present a new, physically motivated, multi-component spectral fitting pipeline that we apply to 746 optical counterparts of Fermi/4FGL-DR4 $\gamma$-ray sources in the SDSS-V Data Release 20 spectroscopic database, yielding 707 well-fitted blazar candidates dominated by Power-law+Galaxy (59.4%), Power-law+Lines (22.6%), and Power-law+QSO (15.0%) model families. Independent WISE infrared (IR) photometry confirms that 96.6% of the 88 sources originally misclassified as Galactic stars by SDSS fall within the canonical blazar region, with 90.9% reclassified as BL Lacertae object (BL Lac) candidates, demonstrating the success of our new pipeline. The new classification scheme naturally recovers the known cosmological separation between BL Lac and Flat Spectrum Radio Quasar (FSRQ) candidates and a separation of approximately one order of magnitude in median $\gamma$-ray luminosity. We compare our redshift estimates to a validation sample of 111 sources in the Third Fermi-LAT Catalogue of High-Energy Sources (3FHL), finding a 10.4% reduction in catastrophic failures ($\eta = 0.387$ vs 0.432) over the SDSS pipeline. The equivalent width analysis validates the traditional |EW| = 5 Ang classification boundary at the population level (BL Lac: 4.47 $\pm$ 0.15 Ang; FSRQ: 22.86 $\pm$ 0.86 Ang), although 49.5% of individual BL Lac candidates exceed this threshold and 51.4% show simultaneous emission and absorption features. This hybrid population challenges the traditional binary blazar classification and points towards a more physically continuous description of blazar properties.

astro-ph.HE

The number of perfect matchings in 3-connected planar graphs

A graph is matchable if it admits a perfect matching. Recently, Goedgebeur et al. asked whether there exists a constant $c<12$ such that infinitely many matchable planar $3$-connected graphs, each with exactly $c$ perfect matchings. We answer this question by proving that every matchable planar $3$-connected graph on at least 40 vertices has at least 12 perfect matchings, and this lower bound is sharp.

math.CO

Bricks in which every vertex is incident with a forcing edge

A matching covered graph is a brick if it is 3-connected and bicritical. An edge of a matching covered graph G is a forcing edge if it lies in precisely one perfect matching of G. We prove that every vertex of a simple brick G is incident with a forcing edge if and only if G is an odd wheel or is isomorphic to one of the following three graphs: the triangular prism \overline{C_6},\overline{C_6}^+ or the bicorn H_8. Here \overline{C_6}^+ is obtained from \overline{C_6} by adding an edge between two nonadjacent vertices.

math.CO

Detection of Variability in Seyfert 2 Galaxies and Measurement of the Optical Scattering Region Size

One main theme of the Unification Model of active galactic nuclei is that there is an obscuring torus structure blocking the direct view of the central engine for Seyfert 2 galaxies. Here, we present the detection of long-term optical variability for a sample of nearby Seyfert 2s. We found that Seyfert 2s exhibit a relatively low but significant level of variability beyond the constant fluxes established by galaxies over month to year time scales. The variability is also detected in the structure functions of Seyfert 2s. Assuming a simple variability suppression model by the scattering region and dilution due to host starlight, where the region smooths the unobscured photon packets from the central engine as the light scatters over the torus, we estimate AGN scattering region sizes by matching the variability amplitudes of Seyfert 1s to 2s. Our measured scattering sizes are largely consistent with the torus size measured using their emission properties, suggesting that the scattering region is of a similar size as the torus. Our results pave the way towards variability as a powerful and independent test for AGN unification models.

astro-ph.GA

Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems

The transition from single-turn models to Multi-Agent Systems (MAS) promises enhanced problem-solving capabilities, yet the centralized orchestration topology remains a critical point of fragility. To analyze this, we propose a Mean-Field Entropy Dynamics framework, modeling the orchestration process as a system governed by the competing forces of task resolution and cumulative context loading. To facilitate validation, we introduce Inverse Workflow Generation (IWG), a multi-agent pipeline that synthesizes process-verifiable, high-complexity benchmarks with dense intermediate checkpoints. We demonstrate that our entropy dynamics model fits empirical trajectories, providing physically interpretable parameters that quantify system stability and performance collapse. Crucially, our analysis uncovers a ``Reasoning Trap": while reasoning-heavy models excel in isolated tasks, they frequently fail as orchestrators due to context squeezing. Elucidating the physical mechanisms underlying the Orchestrator and quantifying systemic uncertainty offers insights for the MASs' architectural design.

cs.AI

PEMark: Watermarking API Responses Based on Proxy Gateways and Position Encoding

Data leakage from API responses has drawn wide attention. APIs are often not fully regulated, making them easy to abuse. One common solution is to embed watermarks into API responses for traceability. However, existing watermarking methods often require modifying database content or API response data. This forces changes to business system code, and may even disrupt normal business operations because data values are altered. In this paper, we propose an original pluggable watermarking scheme based on a watermark proxy gateway and PEMark (Position Encoding-based Watermarking). The key novelty of our approach is exploiting the inherent permutation redundancy in the ordering of JSON/XML key-value pairs -- an overlooked dimension that carries no semantic information yet provides abundant encoding capacity. First, we forward server responses to the watermark proxy gateway, a design that requires zero modification to existing business systems. Then, we embed a watermark into each API response using position encoding, which reorders keys without altering any data values. To the best of our knowledge, this is the first work to achieve distortion-free API response watermarking via position encoding over a proxy gateway. Our method does not modify any data values, so normal business operations continue seamlessly after watermark embedding. Experimental results show that our framework maintains business usability while ensuring that returned API data is traceable. Compared with current mainstream schemes, our method is robust against tampering and insertion attacks (100\% similarity), and can withstand certain levels of deletion attacks.

cs.CR

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

Modality-conflict hallucination occurs when multimodal large language models (MLLMs) prioritize erroneous textual premises over contradictory visual evidence. To understand why visual evidence fails to prevail during generation, we take a mechanistic perspective and examine which internal components drive or resist this failure. We perform head-level causal analysis using path patching across five open-source MLLMs and identify two groups of attention heads with opposing causal roles: hallucination-driving heads and hallucination-resisting heads. We find a consistent asymmetry: driving effects are more broadly distributed and carry greater aggregate weight, whereas resisting effects concentrate in a small number of high-importance heads. Ablation experiments further confirm that these groups exert opposing effects during generation: distributed driving influence and localized resistance together form an imbalanced routing structure that biases generation toward the erroneous premise. Motivated by this finding, we propose MACI (Modality-conflict-Aware Causal Intervention), a conditional intervention that suppresses causally identified hallucination-driving heads only when conflict is detected. Across five MLLMs, MACI achieves the largest hallucination reduction among compared inference-time baselines on the MMMC benchmark with a favorable hallucination-accuracy trade-off, and transfers zero-shot to the SCI-SemanticConflict test.

cs.AI

CogGen: A Cognitively Inspired Recursive Framework for Deep Research Report Generation

The autonomous synthesis of deep research reports represents a critical frontier for Large Language Models (LLMs), demanding sophisticated information orchestration and non-linear narrative logic. Current approaches rely on rigid predefined linear workflows, which cause error accumulation, preclude global restructuring from subsequent insights, and ultimately limit in-depth multimodal fusion and report quality. We propose CogGen, a Cognitively inspired recursive framework for deep research report Generation. Leveraging a Hierarchical Recursive Architecture to simulate cognitive writing, CogGen enables flexible planning and global restructuring. To extend this recursivity to multimodal content, we introduce Abstract Visual Representation (AVR): a concise intent-driven language that iteratively refines visual-text layouts without pixel-level regeneration overhead. We further present CLEF, a Cognitive Load Evaluation Framework, and curate a new benchmark from Our World in Data (OWID). Extensive experiments show CogGen achieves state-of-the-art results among open-source systems, generating reports comparable to professional analysts' outputs and surpassing Gemini Deep Research. Our code and dataset are available at https://github.com/NJUNLP/CogGen.

cs.MA

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process. This isolation leads to a fundamental misalignment: documents identified as topically relevant by information retrieval metrics often fail to provide the actual utility required by the LLM for precise answer generation. To bridge this gap, we introduce ReRanking Preference Optimization (RRPO), a reinforcement learning framework that directly aligns reranking with the LLM's generation quality. By formulating reranking as a sequential decision-making process, RRPO optimizes for context utility using LLM feedback, thereby eliminating the need for expensive human annotations. To ensure training stability, we further introduce a reference-anchored deterministic baseline. Extensive experiments on knowledge-intensive benchmarks demonstrate that RRPO significantly outperforms strong baselines, including the powerful list-wise reranker RankZephyr. Further analysis highlights the versatility of our framework: it generalizes seamlessly to diverse readers (e.g., GPT-4o), integrates orthogonally with query expansion modules like Query2Doc, and remains robust even when trained with noisy supervisors.

cs.CL

The shortest detected intra-day variability of active galactic nuclei in TESS survey

AGNs are known to be variable in almost all wavelengths and timescales. The shortest variability timescale of AGNs can be used to probe the smallest scale structures within AGNs. We aim to measure the shortest detected variability timescale, $t_{min,ul}$, of type 1 radio-quiet Seyfert galaxies and analyse their characteristics. We extracted TESS light curves of 47 Seyfert 1 galaxies. We measured the PSDs of the sample, modelled by a power law model plus a constant noise, and constrained the shortest detected AGN variability timescale as the power law component exceeds the constant noise and systematic uncertainties indicated by the upper limits of non-variable quiescent galaxies' PSDs. We measured the upper limits of the shortest variability timescale to be $\log(t_{min,ul}/hrs)=0.85\pm0.55$. We compared these upper limits to a range of theoretical AGN variability timescales, and the natural interpretation of our measured $t_{min,ul}$ is the light crossing scale from a coherently varying region, where the measured $t_{min,ul}$ corresponds to the range from a few to thousands of gravitational radii. A significant fraction of these light crossing scales is smaller than the accretion disk emission sizes measured by quasar microlensing, reverberation mapping, or theoretical accretion disk models. Since we only measure the upper limits, the true physical shortest variability timescales are even shorter. We measure the power law index to be $2.0\pm0.2$, and find weak anticorrelations with the black hole mass and luminosity. Our analysis suggests that the shortest optical variability is driven by a compact region smaller than the accretion disk size, potentially by X-ray reprocessing. Alternatively, this shortest timescale variability suggests that the accretion disk can be inhomogeneous potentially caused by turbulence from magnetorotational instability or magnetic reconnections. (abridged)

astro-ph.GA

WebNavigator: Global Web Navigation via Interaction Graph Retrieval

Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems from Topological Blindness, where agents are forced to explore via trial-and-error without access to the global topological structure of the environment. To overcome this limitation, we introduce WebNavigator, which reframes web navigation from probabilistic exploration into deterministic retrieval and pathfinding. WebNavigator constructs Interaction Graphs via zero-token cost heuristic exploration offline and implements a Retrieve-Reason-Teleport workflow for global navigation online. WebNavigator achieves state-of-the-art performance on WebArena and OnlineMind2Web. On WebArena multi-site tasks, WebNavigator achieves a 72.9\% success rate, more than doubling the performance of enterprise-level agents. This work reveals that Topological Blindness, rather than model reasoning capabilities alone, is an underestimated bottleneck in autonomous web navigation.

cs.IR

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.

cs.CV

A Decade-Long Increasing Mid-Infrared Luminosity in Galaxy NGC6447: a Turning-On Candidate of Active Galactic Nucleus

It is widely expected that the obscured accretion stage can be the initial turning-on stage of active galactic nuclei from quiescent galaxies. We present mid-infrared light curves of NGC 6447 in 3.5$\mu$m and 4.6 $\mu$m bands observed by WISE/NEOWISE, which show an almost monotonic increasing trend of 1.2 mag over 14 years. The optical light curve from ASAS-SN during the same period is consistent with a constant showing no variability. The mid-infrared color evolution shows that the galaxy transitioned into an active galactic nucleus (AGN) in 2018. The SPHEREx spectrum reveals an increasing continuum resembling warm to hot dust emission from an AGN. NuSTAR detected an X-ray source with a 2-30 keV luminosity of $8.4\times10^{41}$ ergs/s at the lower boundary of AGN X-ray emission range, and a factor of >7 variability in one year compared to the Swift upper limit. NGC 6447 was classified as a quiescent galaxy in the literature. The multi-wavelength timing and spectral properties of NGC 6447 are consistent with the expected AGN turning on event, where the obscuring material around the AGN central engine is gradually dispersed, revealing the central engine. This example shows that long-term infrared variability can be a powerful tool to find similar sources. Based on the sample selection statistics, we estimate the duration of the episodes of AGN accretion (duty cycle) signified by the turning-on event as $10^4$-$10^6$ yr.

astro-ph.GA