SearcharxivSearch

arXiv subjects

Shiyu Li

Publications and source records attributed to Shiyu Li.

At least 19 recordsLinked to original sources

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the ``one-to-many'' nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality through the lens of information fidelity: A caption must maximize the coverage of salient visual information while ensuring strict factuality. We introduce CapQuiz, a novel reference-free benchmark that assesses captions based on their utility in answering human-verified, fine-grained, multiple-choice questions derived from the video. CapQuiz features a hierarchical taxonomy of 10 question types (spanning Descriptive and Inferential categories) across 24 diverse video domains. Extensive experiments demonstrate that CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance.

cs.CV

SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processing long visual token sequences. To reduce inference costs, recent studies have explored visual token pruning through vision-centric or text-guided strategies. However, these methods often overlook high-norm outlier tokens, i.e., tokens with abnormally large feature norms, leading to suboptimal pruning decisions. In this work, we show that such high-norm outlier tokens are highly redundant in both feature and spatial dimensions, yet are often mistakenly preserved as informative cues by existing methods. Motivated by this observation, we propose SinkPruner, a training-free visual token pruning framework for efficient MLLM inference. SinkPruner follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains tokens semantically aligned with the text query. Extensive experiments on twelve image-language and four video-language benchmarks demonstrate the effectiveness, efficiency, and generalizability of our framework. Notably, SinkPruner preserves 96.5% (91.8%) of the original performance of LLaVA-1.5 (Qwen2.5-VL) under an 89% token reduction. Experiments further indicate that our visual sanitizer exhibits promising transferability in enhancing the performance of existing pruning methods. Our code is available at https://github.com/LaVi-Lab/SinkPruner.

cs.CV

Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection

The rapid evolution of large language models necessitates robust machine-generated text detection. Existing paradigms typically follow two isolated tracks. Training-free methods rely on global statistical scalars such as perplexity, while training-based methods utilize semantic hidden states. Both approaches exhibit fundamental vulnerabilities in adversarial scenarios. Global scalars act as lossy compressions that obscure local probabilistic burstiness in interleaved texts, whereas pure semantic models overfit to specific fingerprints and remain susceptible to spoofing. To expose these flaws, we introduce MOSAIC, a comprehensive adversarial benchmark comprising 16000 samples across a full-granularity attack spectrum. To address these challenges, we propose NeuroStat, an end-to-end framework bridging the statistical and semantic gap. NeuroStat captures uncompressed token-level probabilistic logits alongside deep semantic hidden states from a single causal language model backbone. We fuse these heterogeneous signals through Macro-State Residual Modulation, which adaptively calibrates local convolutional features using global uncertainty indicators. Orthogonal and contrastive losses further ensure the learning of complementary representations. Extensive experiments demonstrate that NeuroStat maintains exceptional robustness on MOSAIC compared to the severe degradation of state-of-the-art methods, establishing a new standard for adversarial text detection. Code and the MOSAIC benchmark are available at https://github.com/TencentBAC/NeuroStat.

cs.CL

Compound symmetries and double antisymmetry groups in linear time-invariant photonic systems

Symmetry is fundamental to photonic systems. External (spatial) symmetries and internal symmetries---Lorentz reciprocity, energy conservation, and time-reversal symmetry---constrain the electromagnetic response. Photonic systems can also possess compound symmetries that combine external and internal transformations, exemplified by parity-time (PT) symmetry. However, a unified framework for general compound symmetries involving reciprocity, energy conservation, and time reversal remains lacking, leaving their classification and physical implications unexplored. In this paper, we present such a framework for linear photonic systems. We define compound transformations and symmetries, and derive their constraints on electromagnetic fields and scattering matrices. We show that internal, external, and compound symmetries are naturally described by the theory of double antisymmetry groups. This theory classifies linear time-invariant photonic systems into twelve symmetry categories, each imposing characteristic constraints on the electromagnetic response. We illustrate two representative categories with numerical examples of photonic crystal slabs and apply the theory to examine Kirchhoff's law of thermal radiation for a gyrotropic sphere. Our work provides a systematic foundation for analyzing and engineering symmetry in photonic systems.

physics.optics

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduction. While explicit visual generation incurs prohibitive computational costs, existing latent approaches often rely on external experts or lack rigorous cognitive logic. In this paper, we introduce ProLaViT (Progressive Latent Visual Thought), a framework empowering MLLMs to perform structured visual derivation in the continuous latent space. Unlike works dependent on heterogeneous external models, ProLaViT leverages an endogenous self-distillation mechanism, utilizing the model's own visual encoder to supervise latent thoughts. To facilitate this, we construct a scalable programmatic synthesis pipeline enabling the model to internalize algorithmic precision without inference time tools. We design two reasoning paradigms: (1) Coarse-to-Fine Causal Chain for spatial tasks, guiding attention from global context to local targets. (2) Dialectical Reasoning Chain for logical tasks, incorporating counter-factual thinking for verification. Furthermore, we propose a Distance-Weighted Diversity Loss to impose topology-aware constraints, preventing feature degeneration by enforcing semantic distinctiveness. Extensive experiments demonstrate that ProLaViT outperforms baselines on vision-centric benchmarks, achieving superior accuracy and interpretability with high efficiency.

cs.CV

MedSWFlow: An Open-Source LLM Workflow for Drafting Medical Social Work Case Plans

We present MedSWFlow, an open-source, model-agnostic LLM workflow for drafting medical social work case plans. The framework translates professional case-planning tasks into six stages: assessment, problem analysis, goal setting, intervention planning, risk anticipation, and planned effect evaluation. Drawing on established social work and behavioral frameworks, MedSWFlow standardizes case inputs, builds structured case profiles, and generates reviewable assessment forms and service plans through staged prompting. The system is released as an open-source research framework for reproducible case-plan generation across LLM providers. Outputs are intended as practitioner-reviewed drafts rather than final service decisions. Source code: https://github.com/santhiyacw-droid/MedSWFlow/tree/main.

cs.HC

A Non-Negativity Iterative Approach to Image Deconvolution for SKA

We introduce a novel algorithm for image deconvolution applicable to interferometric radio observations, based on the assumption of non-negative source fluxes. The method enables rapid and efficient image reconstruction in an iterative manner, without requiring prior knowledge or training. Its computational cost scales linearly with the number of pixels: for example, a $512\times 512$ image can be processed in about 1-2 seconds on a standard laptop. We validate the algorithm using both point sources and an extended galaxy image, incorporating a realistic SKA-Low PSF with incomplete $uv$-coverage, though tests are conducted in noise-free simulations. Comparison with the CLEAN method demonstrates that our approach yields a good reconstruction, showing particular promise for the SKA and VLBI observations with sparse $uv$-coverage.

astro-ph.IM

DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects

Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies, exome and genome sequencing have moved many cases from variant detection to post-sequencing interpretation: clinicians must rank patient-specific candidate variants under incomplete fetal or infant phenotypes and heterogeneous evidence from population genetics, variant-effect prediction, gene-disease validity, phenotype ontologies, cellular and pathway context, protein structure and clinical literature. We present DeepBD, a grounded agentic workflow for variant prioritization and diagnostic interpretation of genetic birth defects. DeepBD organizes the workflow into LLM-assisted case structuring, a pretrained evidence engine, specialist evidence modules and a grounded diagnostic review layer. The evidence engine learns patient-specific variant scores from structured rule evidence, sequence and variant-effect representations and phenotype-conditioned biological context, whereas specialist modules and the agentic layer provide tool-based refinement, candidate-pool review and diagnosis-oriented synthesis from ranked candidates. Developed using an in-house fetal and infant cohort comprising 18,622 cases, DeepBD achieved Recall@1/3/5/10 of 0.658/0.882/0.912/0.929 on an internal held-out solved-case benchmark, outperforming standalone Exomiser, DeepRare and prompted LLM reranking baselines evaluated on Exomiser-derived top-20 candidate variants. Ablation and overlap analyses show that rule evidence, mechanistic context, and specialist refinement provide complementary signals. These findings support a grounded agentic workflow that separates evidence integration, tool-based refinement, and LLM-assisted diagnostic review for retrospective variant prioritization in genetic birth defects.

q-bio.GN

Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding

While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent training-free methods address this through image scaling and localized cropping. However, applying these manipulations indiscriminately introduces computational redundancy for simple queries and can degrade accuracy by truncating essential global context or introducing irrelevant background noise. To this end, we propose LazyMCoT, a dynamic and training-free framework that adaptively allocates visual grounding efforts based on sample difficulty. The framework features an Adaptive Routing mechanism that evaluates predictive uncertainty using first-token statistics from a single forward pass. This efficiently bypasses confident cases while ensuring the recall of difficult samples via conformal calibration. For these challenging cases, a Collaborative Grounding module integrates the inherent cross-modal attention of the model with an external visual expert through a two-stage refinement process. This refinement process generates a precise localized display to recover small or occluded targets. Extensive experiments across diverse benchmarks demonstrate that LazyMCoT rivals training-based approaches by simultaneously improving reasoning accuracy and reducing average inference latency. Our code is availble at https://github.com/TencentBAC/LazyMCoT.

cs.CV

Intercoupling of bulk acoustic streaming and acoustothermal effect: A high-frequency focused beam example

High-frequency focused acoustic beams are promising for selective trapping of cells in fluids, but the related acoustic absorption may generate large acoustothermal effect which could cause thermal heating on cells or microparticles and bring extra acoustic body force due to the thermal gradient. The theory of the bulk acoustic streaming and acoustic radiation force in a focused-beam for the three-dimensional selective trapping of a cell has been developed [Li and Gong, Phys. Rev. Fluids, 11, 054201 (2026)], however, the acoustothermal effect and its feedback on the acoustic field at high frequency with strong absorption remain weakly understood. To solve this issue, we develop a theoretical and numerical model that couples acoustic propagation, bulk acoustic streaming, and acoustothermal effect in water. The acoustic body force is decomposed into a viscous-attenuation-induced acoustic body force $\mathbf{f}_{\mathrm{E}}$ and a temperature-gradient-induced acoustic body force $\mathbf{f}_{\mathrm{T}}$, while the temperature field is fed back to the frequency-domain acoustic calculation through the temperature-dependent material properties. Taking the single focused beam for example, within the pressure range constrained by the mechanical index, $\mathbf{f}_{\mathrm{T}}$ remains weaker than $\mathbf{f}_{\mathrm{E}}$, whereas streaming-induced convection can markedly reduce the temperature rise when the thermal Peclet number ($Pe_T$) exceeds unity. This work establishes a theoretical basis for predicting and controlling the intercoupling of bulk acoustic streaming and acoustothermal effec of high-frequency focused beams which will be helpful for the design of single-beam acoustical tweezers.

physics.app-ph

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

Omni-modal retrieval promises a single embedding space for text, image, video, document, and audio inputs, but building such a unified retriever is difficult since these modalities differ in data distribution, architecture, and optimization dynamics. In this work, we present Conan-embedding-v3, a decouple--fuse--recover framework for omni-modal retrieval. Conan-embedding-v3 first trains modality specialists independently and fuses their task vectors into a single dense backbone, a strategy we call Decoupled Specialist Fusion. We show that this fusion composes visual, video, and document retrieval capabilities, but also exposes a failure mode for projector-based modalities: when audio is attached through an external encoder and projector, fusing the backbone leaves the projector calibrated to the audio-specialist backbone, causing a large audio retrieval regression despite copying all audio-specific modules unchanged. We call this failure Projector Drift. To repair it, Conan-embedding-v3 applies Projector Recovery (i.e., full-parameter fine-tuning of the projector while keeping the backbone frozen) followed by balanced multi-modal rehearsal. The resulting model supports these retrieval pathways in one backbone, achieving 74.9 scores on MMEB while obtaining 55.61 on the 30-task MAEB audio suite.

cs.MM

Multi-Camera AR Guidance System for Surgical Instrument Handling and Assembly: Investigating Workload and Efficiency

The handling and assembly of instruments during surgery imposes high cognitive demands on scrub nurses, particularly when instruments are unfamiliar. We present a supporting guidance system for surgical instrumentation that combines multi-camera 6D pose estimation with augmented reality in-situ visualization on a head-mounted display without the requirement for additional markers. Pose estimation and consecutive camera calibration are achieved through known objects. The 6D pose estimation network is trained purely on synthetic data, aiming for better generalizability and real-world applicability. The AR guidance displays tooltip localization cues and step-wise assembly animations. Via gaze-based selection and a foot pedal, users can switch between assembly steps in intraoperative use. In a technical evaluation, our approach outperforms state-of-art 6D pose estimation. A user study with 29 scrub nurses was conducted in a surgical simulation of knee arthroplasty, comparing the system against a paper manual. AR guidance significantly reduced the perceived workload compared. Objectively, AR guidance reduced task completion time by 21.3\% (4.76 minutes). Specifically, scrub nurses less experienced with the instrument set benefited when using the system. Error frequencies were comparable between conditions. Qualitative feedback highlighted improved process clarity, reduced information overload, and perceived independence. To summarize, our marker-free multi-camera AR guidance approach for surgical instruments can, subjectively and objectively, improve intraoperative instrumentation performance, particularly for untrained scrub nurses.

cs.CV

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs

The scaling of large language models (LLMs) is currently bottlenecked by the rigidity of distributed programming. While high-performance libraries like CuBLAS and NCCL provide optimized primitives, they lack the flexibility required for rapidly evolving model architectures. Conversely, existing tensor compilers fail to address the complex memory hierarchy of distributed clusters effectively. To bridge this gap, we propose DITRON, a scalable tile-level compiler that democratizes high-performance distributed kernel development. DITRON introduces a novel hierarchical programming abstraction spanning Core, Device, and Task levels to map tensor programs efficiently onto heterogeneous distributed hardware. This abstraction allows DITRON to support diverse parallelism strategies while abstracting away the complexity of inter-node and intra-node communication. Evaluated across large-scale clusters, DITRON achieves performance parity with or exceeding expert-tuned CUDA libraries, delivering speedups of $6\%-30\%$ on isolated kernels and $5\%-30\%$ on end-to-end inference in vLLM. Furthermore, DITRON demonstrates strong portability, achieving significant speedups on both NVIDIA and AMD platforms. \ours{} has been deployed at the enterprise level for both training and inference. It achieves an MFU improvement of over 10\% in training tasks, saving approximately 500,000 GPU hours of training cost per month. For inference tasks, it delivers an end-to-end gain of over 20\% and has been applied to cloud service inference and edge inference scenarios.

cs.PL

Competition between acoustic radiation force and streaming-induced drag force in focused beams for 3D cell trapping

The ability to trap a single cell or microparticle in three dimensions is important for biomedical and microfluidic applications. Single-beam acoustic tweezers based on focused waves provide a compact and biocompatible approach because of their high spatial resolution and strong intensity gradients. However, 3D trapping remains challenging, especially at high frequencies, because the weak axial restoring radiation force may not overcome the pushing drag force caused by acoustic bulk streaming in free space. The combined effect of acoustic radiation force and streaming-induced drag force on a microparticle has not been systematically studied. Although the radiation force scales with the square of the focal pressure amplitude p_foc, the scaling of streaming-induced drag force with p_foc under different flow conditions remains unclear. Here, we establish a unified theoretical and numerical framework to compare these two effects and derive an explicit scaling law, U0 ~ p_foc^n, for the streaming velocity from the viscous to the inertial regime. We show that n = 2 in the viscous limit (Re_lambda << 1), n = 4/3 in the inertial limit (Re_lambda >> 1), and n lies between 4/3 and 2 in the transition regime (Re_lambda ~ 1). We further introduce the Schiller-Naumann model to estimate the drag force more accurately than the Stokes model. On this basis, we find that the ratio of axial radiation force to drag can vary non-monotonically with p_foc, contrary to the conventional expectation of monotonic increase. This work provides a theoretical basis for optimizing single-beam acoustic tweezers for stable 3D trapping of single cells.

physics.app-ph

Experimental demonstration for precisely tuning the focal length of finite-aperture focused beams and vortex

High-frequency focused ultrasound is widely used in biomedical applications such as high-resolution imaging, neuromodulation, particle manipulation, and so on. However, dynamic tuning of the focal plane in conventional systems often relies on mechanically adjustable components or array-based control with complex system and high cost. In this work, an optically transparent, planar compact piezoelectric ultrasonic transducer was designed and fabricated by truncating an ideal spherical wavefront with a plane, enabling high-frequency focused ultrasound generation and convenient integration with microscopic platforms. The acoustic field was characterized experimentally at the focal plane under the design frequency and at propagation planes near the design frequency to evaluate the focal tuning. An approximate linear relation between the focal length and driving frequency near the design one is derived theoretically, and the finite-range tuning behavior is interpreted using the stationary-phase condition. Both theory and experiment show that the focal length varies approximately linearly with excitation frequency near the design frequency. Water-tank measurements agree well with the theoretical prediction, confirming the proposed model. This work provides a simple and cost-effective approach for focal tuning in compact high-frequency ultrasound devices.

physics.app-ph

MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects

Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or objects, across multiple views. While commonly cameras are calibrated in an initial step or updated through the constant use of markers, another option is to leverage information already present in the scene, like known objects. Another downside of marker-based tracking is that markers have to be tracked inside the field-of-view (FoV) of the cameras. To overcome these limitations, we propose a constant dynamic camera pose estimation leveraging spatiotemporal FoV overlaps of known objects on the fly. To achieve that, we enhance the state-of-the-art object pose estimator to update our spatiotemporal scene graph, enabling a relation even among non-overlapping FoV cameras. To evaluate our approach, we introduce a multi-camera, multi-object pose estimation dataset with temporal FoV overlap, including static and dynamic cameras. Furthermore, in FoV overlapping scenarios, we outperform the state-of-the-art on the widely used YCB-V and T-LESS dataset in camera pose accuracy. Our performance on both previous and our proposed datasets validates the effectiveness of our marker-less approach for AR applications. The code and dataset are available on https://github.com/roth-hex-lab/IEEE-VR-2026-MultiCam.

cs.CV

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although CoT prompting enhances reasoning, its verbosity imposes substantial computational overhead. Recent works often focus exclusively on outcome alignment and lack supervision on the intermediate reasoning process. These deficiencies obscure the analyzability of the latent reasoning chain. To address these challenges, we introduce Render-of-Thought (RoT), the first framework to reify the reasoning chain by rendering textual steps into images, making the latent rationale explicit and traceable. Specifically, we leverage the vision encoders of existing Vision Language Models (VLMs) as semantic anchors to align the vision embeddings with the textual space. This design ensures plug-and-play implementation without incurring additional pre-training overhead. Extensive experiments on mathematical and logical reasoning benchmarks demonstrate that our method achieves 3-4x token compression and substantial inference acceleration compared to explicit CoT. Furthermore, it maintains competitive performance against other methods, validating the feasibility of this paradigm. Our code is available at https://github.com/TencentBAC/RoT

cs.CL

Optimal universal bounds for waves with varied coherence based on supremum and infimum coherence spectra

We establish a majorization-based theory for bounding observables of waves with varied coherence. For any measurement, exact bounds are attained by the maximal and minimal elements in the set of input coherence spectra. The set's supremum and infimum, which may lie outside the set, provide optimal universal bounds: any alternative spectrum yielding universal bounds produces weaker constraints. We present an algorithm to compute the supremum and infimum, and prove that they lie either at singular boundary points or strictly outside the set of coherence spectra.

physics.optics