SearcharxivSearch

arXiv subjects

Yan Xia

Publications and source records attributed to Yan Xia.

At least 19 recordsLinked to original sources

VibeVoice-ASR-Streaming Technical Report

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.

eess.AS

MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models

Achieving 360{\deg} coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existing multi-view fisheye datasets are limited, and synthesizing rare corner cases typically requires computationally expensive 3D simulations, hindering the training. While generative models have achieved significant success in standard perspective imagery, their application to wide-angle distortion remains unexplored. In this work, we formally introduce the novel problem of multi-view fisheye image generation conditioned on volumetric semantic representations and present two distinct methods. We first propose SyntheOcc-FE, which adapts the SyntheOcc architecture to fisheye data. While effective, this method is constrained by the scarcity of fisheye datasets, which limits its generalization. To overcome these limitations, we propose our second method, MIVIFI (multi-view fisheye), which leverages cross-domain learning with Equirectangular Projections. By bridging the gap between dataset domains using KITTI-360 fisheye images alongside nuScenes multi-view standard images, our approach enables high-fidelity manipulation of scene content. This framework enables the structural modification of semantic occupancy inputs to introduce or eliminate specific actors and facilitates the rendering of diverse meteorological conditions and illumination scenarios absent in the limited fisheye datasets. Quantitative and qualitative experiments demonstrate that our methods achieve robust photorealistic multi-view fisheye image generation and highlight the specific advantages of our cross-domain strategy for handling data scarcity.

cs.CV

Video Diffusion for Satellite-based High-Dynamical-Fidelity Precipitation (HiDFiP) Field Generation

High spatiotemporal fidelity precipitation products that accurately capture storm spatial organization, propagation, and lifecycle evolution, are essential for advancing hydrometeorological research and operations at regional and global scales. Satellite products offer the only near-global precipitation observations, but they still fall short of reproducing the spatiotemporal structure of ground-based references, due largely to dynamic distortions from the inhomogeneity, intermittency, and indirectness of satellite retrievals. Here we propose a video-diffusion framework for satellite-based High-Dynamical-Fidelity Precipitation (HiDFiP) field generation beyond the space-time coverage of ground-radar, using radar-rich CONUS as a testbed. The framework performs explicit spatiotemporal modeling with IMERG as the primary source and leverages four-dimensional storm-environment information from about 40 ERA5/ERA5-Land atmospheric/land fields to compensate for the temporal information deficit inherent to satellite retrievals. We introduce an extensive metric suite to assess HiDFiP dynamical fidelity in temporal-reconstruction and spatial-transfer settings against MRMS ground-radar precipitation over CONUS. Relative to IMERG and image-wise diffusion baselines, HiDFiP accurately reproduces the storm space-time spectral characteristics; storm timing, location, and directional propagation; precipitation-event episodicity and temporal structure; precipitation-system morphology and spatial organization; and storm-track kinematics and lifecycle evolution. Transferability experiments indicate that HiDFiP generalizes reasonably well to an unseen region. This work advances a video-diffusion paradigm for satellite-based high-spatiotemporal-fidelity precipitation generation and provides an algorithmic and diagnostic foundation for long-term global radar-grade precipitation records.

physics.ao-ph

SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse masks, or pseudo-masks. These weak annotations may include texture-similar distractors and background context alongside the target, contaminating class prototypes or visual prompts before query prediction. We introduce SADe, a predictor-agnostic support decontamination layer that estimates the reliability of selected support patches without query information. Central to SADe is sparse autoencoder (SAE) atom evidence: dense similarity may respond to both target and texture-similar context, whereas contrasting atom activations inside and outside the weak-support region provides factor-level reliability cues. A lightweight router combines atom evidence with dense similarity and episode statistics to predict patch reliability and generate a cleaned support mask. Trained once on synthetic weak-support episodes from FSS-1000, the router is frozen for all target evaluations. The resulting mask supports standalone prediction or can be supplied to heterogeneous FSS models through native support interfaces without altering query-side inference. Under a matched weak-support protocol, SADe achieves the highest query mIoU in six of nine standalone prompt-shot combinations. With the same ProMi query head, it is within 0.03 mIoU of SAM3-derived masks under tight boxes and surpasses them by 11.17 and 19.49 points under box-r2 and box-r4, respectively. As a plug-in, SADe improves over raw support in 70 of 72 matched box-family comparisons across four frozen downstream models and two datasets. On point and scribble prompts, its average performance remains close to the corresponding raw-support baseline. Ablations and atom-removal controls show that atom evidence contributes reliability information beyond dense similarity.

cs.CV

Agentic Autoresearch for CT Reconstruction

Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether a ranking measured on ideal data predicts behavior under realistic noise. We built an agentic loop: the agent edits a solver, runs a short cluster job, reads one frozen metric, and revises. The metric is a calibrated headroom score against the FBP baseline, inside the field of view; every method shares the same differentiable fan-beam projector. We benchmarked 26 methods on Mayo low-dose CT (noise-limited) and a 128-view sparse-view breast task from the noiseless DL-Sparse-View Challenge, with validation-selected iterations scored on a held-out test set. Every trained breast model was then re-scored on noisy inputs (I_0 = 10^5 photons) without retraining, and separately retrained on matched noise. The agent independently implemented, tuned, and benchmarked all 26 methods, and recombined them into a compact solver of 969 parameters that ties the top Mayo tier at the 1% level using 0.4% of the champion's parameters. Benchmarking gives a tier of statistically indistinguishable top methods, not one winner. Mild input noise nearly inverts the breast ranking: the noiseless champion (a supervised image denoiser, hr 0.89) collapses to 0.00, while a learned primal-dual method rises to champion (0.72 to 0.93). An ideal-data leaderboard therefore does not predict robustness. The inversion is a transfer effect, not a permanent deficit: retraining on matched noise restores much of the clean ranking (Spearman rho 0.04 to 0.61). Noise is only the easiest confounder in an open-ended set (beam hardening, scatter, anatomy, disease), so no single-factor challenge certifies generality. Benchmarks should model a broad spectrum of realistic factors at once.

physics.med-ph

VibeVoice-ASR-BitNet Technical Report

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition (RTF < 1) on low-thread-count CPUs. VibeVoice-ASR-BitNet is 1.6--2.3x faster than Whisper.cpp at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.

cs.SD

AI-Powered Browsers Are Broadly Accurate News Summarizers That Reduce Political Bias and Negative Affect

Web browsers now provide AI-generated news summaries for millions of users. Despite their popularity and influence, we lack a systematic understanding of how these systems transform news before people read it. Through a large-scale audit, we investigate the factual accuracy of browser-based AI summarizers and how they alter the political bias, negative affect, and journalistic writing quality of news. Drawing on 13,777 articles from 15 U.S. news outlets, we evaluate their 41,331 summaries generated by three leading AI-powered browsers: Google Chrome (Gemini), Microsoft Edge (Copilot), and Perplexity Comet. We find that browser-based AI summarizers are broadly accurate. Furthermore, they consistently transform news by attenuating ideological bias, partisan stances, negativity, anger, and fear, while increasing clarity and reducing personal tone. With some variations, these patterns hold across browsers, outlet ideologies, and topics. Our findings identify AI-powered browsers as a new class of editorial intermediaries that systematically reshape news, with implications for democratic discourse and AI governance.

cs.CY

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents. However, current BEV representations, typically extracted by LiDAR backbones trained from scratch, are geometry-dominated and lack general semantic priors, inherently limiting the efficacy of feature-level collaboration. Meanwhile, vision foundation models (VFMs) pretrained on large-scale image data have demonstrated strong capability in learning general-purpose and informative visual representations for 2D tasks, and have the potential to enhance agent-wise LiDAR BEV representations for collaboration. Despite this potential, adapting VFMs to LiDAR-based 3D detection remains challenging due to the substantial image-point cloud modality gap. To bridge this gap, we propose ViCo3D, a collaborative 3D object detection framework powered by VFMs. Specifically, ViCo3D adapts VFMs to LiDAR-based collaborative perception from three aspects: First, ViCo3D projects point clouds onto the BEV plane as three-channel images, enabling DINOv2 to extract BEV-space visual features from LiDAR inputs. Besides, to effectively integrate these DINOv2-derived features with LiDAR geometric features, ViCo3D introduces a multi-scale BEV fusion module within the single-agent encoder. In addition, ViCo3D adopts an ego-centric cross-agent fusion strategy to aggregate complementary information from multiple agents. Experiments on DAIR-V2X and V2XSet demonstrate that ViCo3D achieves state-of-the-art 3D detection performance. Remarkably, it delivers up to 1.8x greater collaborative gains than prior methods on DAIR-V2X. The code will be made public available for future investigation.

cs.CV

Spectral Chaos Does Not Determine Quantum Mpemba Crossings

In a symmetry-restoration quantum Mpemba effect, an initial state with stronger local symmetry breaking can lose that memory faster than a state that starts closer to the symmetric manifold. We test whether this local ordering reversal is organized by chaotic thermalization in a clean U(1)-conserving spin chain, comparing spectral level statistics with crossings of the entanglement asymmetry for the same Hamiltonians. We find that Gaussian orthogonal ensemble (GOE)-like level statistics alone do not determine whether Mpemba crossings occur. Across field textures, GOE-like spectra can occur with or without entanglement-asymmetry crossings, and crossings can also appear away from the GOE reference. A near-staggered detuned control further shows that even an inversion of the total charge-sector coherence need not produce an entanglement-asymmetry crossing. Thus the crossing response is controlled not by spectral chaos alone, but by how local charge-sector coherence enters the reduced density matrix.

quant-ph

WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis

Generating CT volumes from MRI and CBCT can improve treatment planning in adaptive radiotherapy while avoiding additional radiation exposure. However, direct regression of CT intensities is challenged by the inherently high dynamic range and long-tailed distributions, thereby averaging out sparse yet clinically important structures. To alleviate this issue, we reformulate the regression target into multiple windowed representations, leveraging the inductive prior that CT intensities are structure-deterministic and window-separable. These windowed views exhibit smoother distributions and admit structured fusion back to the full-range CT. Building on this reformulation, we introduce WING, a WINdow-prior-based Generative network comprising: 1) a new Gated Inception Generator to produce multi-window predictions, enabling multi-shape kernel interactions to capture cross-modality correspondence; 2) a Fuse-and-Refine Transformer to aggregate the windowed outputs and learn residuals for detail refinement; and 3) a joint adversarial training objective to enhance window-conditioned realism. Extensive experiments demonstrate that our compact WING achieves state-of-the-art performance on the MRI-to-CT and CBCT-to-CT benchmarks, while supporting multi-anatomy synthesis with a single model.

cs.CV

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native tools and bound into one navigable deliverable for revision and reuse. We present ResearchStudio-Reel, a native-editable dissemination workspace that binds its three artifacts into one interactive deliverable at the experience level, implemented as five skills executable in Claude Code and Codex: one shared extractor, three editable artifact generators, and one interactive convergence layer. A shared asset bundle feeds a PowerPoint poster and video deck, plus a bilingual Word blog; rather than re-rendering the paper into a fourth format, Paper2Reel converges these already-produced artifacts at the experience level, binding poster regions, video segments, and blog passages into one interactive viewer. Artifact-specific release checks make this delivery contract testable, and Paper2Poster additionally uses a measured-fill loop. On the Paper2Poster benchmark, our Claude Code configuration achieves the best scores among automated systems on all three aesthetic sub-criteria and the best or tied-best scores on two of three information sub-criteria. Under two VLMjudges, it exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins on overall quality on 74 and 95 of the 100 papers under the two judges. The full pipeline additionally packages the native-editable source artifacts and their aligned viewer. Project is available at https://aka.ms/ResearchStudio

cs.CV

Qubit Readout via State-Dependent Radiative Linewidths

Fast qubit readout conventionally encodes state information in a dispersive frequency shift. Here we formulate a linewidth-encoded quantum non-demolition measurement channel in which the qubit state enters the external radiative amplitude, equivalently a state-dependent Lindblad jump operator. Starting from an empty cavity, we show analytically that this dissipative channel imprints state information on the output field at $O(t)$, whereas standard dispersive readout starts at $O(t^2)$ because it requires intracavity buildup and conditional phase accumulation. This short-time scaling produces faster matched-filter signal-to-noise ratio accumulation and persists in finite-resource comparisons, including photon-number limits, external-linewidth budgets, cavity depletion, and pulse-optimized dispersive baselines. We further outline an auxiliary-mode route that converts a qubit-state-dependent auxiliary susceptibility into a state-dependent linewidth. These results identify engineered dissipation as an information-carrying resource for fast quantum non-demolition readout.

quant-ph

Spin-Squeezing-Enhanced Charging for Quantum Dicke Batteries

High-power Dicke quantum batteries (QBs) typically exploit collective superradiance, whereas intrinsic matter-matter interactions are conventionally considered detrimental. Here, we propose a counterintuitive paradigm: these interactions can be controlled and repurposed as a synergistic resource to enhance charging power and capacity. In the low-excitation limit, transverse interactions induce collective spin squeezing, causing critical mode softening and an exponential enhancement of effective coupling, which significantly boosts charging power. At higher excitations, these interactions act as a macroscopic nonlinear torque. By appropriately aligning this torque, we effectively lower phase-space dynamical barriers, guiding the system along optimal rapid-charging paths. Importantly, this cooperative enhancement remains highly robust under realistic dissipation, outperforming ideal, dissipationless Dicke QBs in specific regimes. Our results provide a blueprint for exploiting matter interactions to design dissipation-resistant, high-performance many-body QBs.

quant-ph

Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipulations are difficult to standardize, motivating the use of humanoid robots with dexterous hands. Despite this opportunity, no existing benchmark evaluates humanoid manipulation in precision-critical laboratory environments. We present Labimus, to our knowledge, the first benchmark for humanoid dexterous manipulation in organic chemistry laboratories. Labimus reconstructs over 30 functionally faithful assets from real organic chemistry workstations through real-to-sim modeling, collectively covering the core operations of routine organic chemistry experiments. The benchmark integrates articulated laboratory instruments, particle-based powder physics, and closed-loop instrument readouts, enabling a complete manipulation-to-measurement pipeline. It further defines six atomic operations and a seven-step solid-weighing workflow derived from real laboratory standard operating procedures. We introduce a precision-aware evaluation protocol designed to jointly measure task completion, experimental precision, and long-horizon execution. We benchmark three representative policies under procedural layouts and environmental perturbations. Results reveal a precision gap: policies that successfully complete laboratory tasks can still fail to satisfy the quantitative tolerances required by experimental protocols. Our benchmark exposes a fundamental disconnect between task completion and experimental validity, providing a new testbed for developing reliable humanoid robots for scientific laboratories.

cs.RO

SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models

Aerial 6DoF localization typically relies on precise GNSS signals or radiometrically rich 3D reconstructions, limiting scalability and on-board deployment. We propose SemCityLoc, a semantic-geometric alignment system that reframes aerial pose estimation as structured surface registration between foundation-model-derived visual priors and standardized LoD-compliant 3D city models. Instead of matching sparse contours or dense texture, our method aligns semantic surfaces and monocular depth with lightweight semantic 3D building models, increasing pose discriminability in repetitive and occluded urban environments. To enable accurate evaluation, we introduce SemCityLockeD, the first real-world benchmark combining centimeter-accurate UAV poses with standardized LoD1--LoD3 semantic city models and challenging low-altitude imagery. Experiments demonstrate substantial improvements over existing map-based approaches, improving recall by up to 36% and reducing mean positional error from 9.89m to 2.62m in challenging urban canyons. Our results indicate that semantically structured geometry provides sufficient and scalable constraints for high-precision aerial localization without radiometric scene reconstructions. The code and data are available at https://albertchen98.github.io/SemCityLoc.

cs.CV

BitNet Text Embeddings

LLM-based text embedders have substantially improved retrieval and semantic representation quality, but their deployment remains costly: large backbone models slow down embedding inference, while high-dimensional full-precision embeddings impose substantial storage and bandwidth overhead on large-scale indexes. In this paper, we present BITEMBED, an extreme low-bit framework for LLM-based text embedding that jointly targets encoding efficiency and vector storage. BITEMBED converts pretrained LLM backbones into BitNet-style embedding encoders with ternary weights, quantized activations, and lightweight normalization refinement. The converted model is adapted to representation learning through continual contrastive pre-training, followed by supervised contrastive fine-tuning with both similarity-distribution distillation and attention-relation distillation from a full-precision teacher. Beyond quantizing the backbone, BITEMBED further trains output embeddings to support multiple storage precisions meeting different storage needs in various scenarios. Experiments on MMTEB (eng, v2) with Qwen3-0.6B and Gemma3-270M show that BITEMBED is largely comparable to full precision teacher embedders. Moreover, BITEMBED flexibly obtains text embeddings of various precisions, achieving a trade-off between performance and storage cost.

cs.CL

Learning Robust Pair Confidence for Multimodal Emotion-Cause Pair Extraction

Multimodal emotion-cause pair extraction (MECPE) requires reliable pair confidence over candidate pairs. Existing pair scorers commonly use pair-level cross entropy over valid candidates, which treats links mostly independently. This leaves the relative confidence geometry among competing causes under-constrained, allowing gold pairs to stay close to hard negatives or rely on incidental non-gold context. We study this vulnerability as pair-confidence brittleness and propose RPCL (Robust Pair Confidence Learning), a training-only framework for pair-confidence learning. RPCL encourages pair confidence to be both discriminative and stable: gold pairs are separated from row-wise hard negatives through a confidence-difference margin constraint, and clean pair predictions are aligned with predictions from a corrupted view where non-gold contextual utterance representations are partially corrupted. The original clean pair scorer and decoding pipeline are used unchanged at inference time. On ECF, MECAD, and MEC4, RPCL improves the three-seed mean Pair F1 over a matched base model by 2.58 to 2.83 percentage points in the full text-audio-video setting, and improves mean Pair AUPRC on all three datasets. Diagnostic analysis further shows larger gold-negative confidence gaps and lower margin-violation severity. These results suggest that explicitly shaping pair confidence is an effective training strategy for MECPE.

cs.CL

Anisotropic Rabi Model as a Noise Biased Qubit

We present the quantum anisotropic Rabi model as a potential resource for a noise biased qubit. The system-environment coupling can be biased by tuning the relative strengths of the rotating-wave and counter-rotating-wave interactions, characterized by the anisotropy parameter $\eta$. This anisotropy selectively suppresses dominant decoherence pathways, thereby enabling the construction of a protected logical qubit in the ultrastrong and deep-strong coupling regimes. The logical states (formed by the ground and first excited states of the anisotropic Rabi model) possess coherence times that are enhanced compared to the isotropic case. Moreover, we construct a set of universal gate operations within the logical-state subspace and demonstrate that the gate operations associated with different values of $\eta$ exhibit robustness against external noise. These findings are expected to inspire applications and research directions for the anisotropic Rabi model with promising potential impacts.

quant-ph