SearcharxivSearch

arXiv subjects

Chaoyang Li

Publications and source records attributed to Chaoyang Li.

12 recordsLinked to original sources

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.

cs.CL

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

Recent multimodal large language models (MLLMs) support Thinking with Images, invoking visual tools such as zooming and cropping to inspect image regions during inference. Yet these systems remain brittle in fine-grained reasoning: to acquire a decisive detail, a model must ground its attention on the correct region, but knowing which region is correct presupposes having already observed that detail. We identify this circular dependency as the grounding paradox, show that grounding errors are rarely self-corrected within a single trajectory---once a misleading region is inspected, all subsequent reasoning conditions on that observation and the error propagates to the final answer---and observe that because each trajectory constructs its own evidence, answer-level aggregation discards the very information that distinguishes trajectories. We propose Test-Time Scaling over Perception (TTSP), a closed-loop framework that treats perception as the unit of scalable inference and allocates compute along two axes: Entropy-Gated Perceptual Exploration samples diverse trajectories and uses critical-token entropy to withhold evidence the model cannot commit to, while Evidence-Guided Iterative Refinement distills validated observations into a correctable Evidence Ledger that steers later rounds to re-inspect unresolved regions. Across high-resolution and general multimodal benchmarks, TTSP consistently outperforms strong test-time scaling baselines, while improving grounding quality with favorable token efficiency.

cs.CV

Decoupling Defense Strategies for Robust Image Watermarking

Deep learning-based image watermarking, while robust against conventional distortions, remains vulnerable to advanced adversarial and regeneration attacks. Conventional countermeasures, which jointly optimize the encoder and decoder via a noise layer, face 2 inevitable challenges: (1) decrease of clean accuracy due to decoder adversarial training and (2) limited robustness due to simultaneous training of all three advanced attacks. To overcome these issues, we propose AdvMark, a novel two-stage fine-tuning framework that decouples the defense strategies. In stage 1, we address adversarial vulnerability via a tailored adversarial training paradigm that primarily fine-tunes the encoder while only conditionally updating the decoder. This approach learns to move the image into a non-attackable region, rather than modifying the decision boundary, thus preserving clean accuracy. In stage 2, we tackle distortion and regeneration attacks via direct image optimization. To preserve the adversarial robustness gained in stage 1, we formulate a principled, constrained image loss with theoretical guarantees, which balances the deviation from cover and previous encoded images. We also propose a quality-aware early-stop to further guarantee the lower bound of visual quality. Extensive experiments demonstrate AdvMark outperforms with the highest image quality and comprehensive robustness, i.e. up to 29\%, 33\% and 46\% accuracy improvement for distortion, regeneration and adversarial attacks, respectively.

cs.CV

Crucible: Quantifying the Potential of Control Algorithms through LLM Agents

Control algorithms in production environments typically require domain experts to tune their parameters and logic for specific scenarios. However, existing research predominantly focuses on algorithmic performance under ideal or default configurations, overlooking the critical aspect of Tuning Potential. To bridge this gap, we introduce Crucible, an agent that employs an LLM-driven, multi-level expert simulation to turn algorithms and defines a formalized metric to quantitatively evaluate their Tuning Potential. We demonstrate Crucible's effectiveness across a wide spectrum of case studies, from classic control tasks to complex computer systems, and validate its findings in a real-world deployment. Our experimental results reveal that Crucible systematically quantifies the tunable space across different algorithms. Furthermore, Crucible provides a new dimension for algorithm analysis and design, which ultimately leads to performance improvements. Our code is available at https://github.com/thu-media/Crucible.

cs.AI

Towards User-level QoE: Large-scale Practice in Personalized Optimization of Adaptive Video Streaming

Traditional optimization methods based on system-wide Quality of Service (QoS) metrics have approached their performance limitations in modern large-scale streaming systems. However, aligning user-level Quality of Experience~(QoE) with algorithmic optimization objectives remains an unresolved challenge. Therefore, we propose \texttt{LingXi}, the first large-scale deployed system for personalized adaptive video streaming based on user-level experience. \texttt{LingXi} dynamically optimizes the objectives of adaptive video streaming algorithms by analyzing user engagement. Utilizing exit rate as a key metric, we investigate the correlation between QoS indicators and exit rates based on production environment logs, subsequently developing a personalized exit rate predictor. Through Monte Carlo sampling and online Bayesian optimization, we iteratively determine optimal parameters. Large-scale A/B testing utilizing 8\% of traffic on Kuaishou, one of the largest short video platforms, demonstrates \texttt{LingXi}'s superior performance. \texttt{LingXi} achieves a 0.15\% increase in total viewing time, a 0.1\% improvement in bitrate, and a 1.3\% reduction in stall time across all users, with particularly significant improvements for low-bandwidth users who experience a 15\% reduction in stall time.

cs.MM

Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models

Over the past decade, adaptive video streaming technology has witnessed significant advancements, particularly driven by the rapid evolution of deep learning techniques. However, the black-box nature of deep learning algorithms presents challenges for developers in understanding decision-making processes and optimizing for specific application scenarios. Although existing research has enhanced algorithm interpretability through decision tree conversion, interpretability does not directly equate to developers' subjective comprehensibility. To address this challenge, we introduce \texttt{ComTree}, the first bitrate adaptation algorithm generation framework that considers comprehensibility. The framework initially generates the complete set of decision trees that meet performance requirements, then leverages large language models to evaluate these trees for developer comprehensibility, ultimately selecting solutions that best facilitate human understanding and enhancement. Experimental results demonstrate that \texttt{ComTree} significantly improves comprehensibility while maintaining competitive performance, showing potential for further advancement. The source code is available at https://github.com/thu-media/ComTree.

cs.MM

A Localization Method of High Energy Transients for All-Sky Gamma-Ray Monitor

Fast and reliable localization of high-energy transients is crucial for characterizing the burst properties and guiding the follow-up observations. Localization based on the relative counts of different detectors has been widely used for all-sky gamma-ray monitors. There are two major methods for this counts distribution localization: $χ^{2}$ minimization method and the Bayesian method. Here we propose a modified Bayesian method that could take advantage of both the accuracy of the Bayesian method and the simplicity of the $χ^{2}$ method. With comprehensive simulations, we find that our Bayesian method with Poisson likelihood is generally more applicable for various bursts than $χ^{2}$ method, especially for weak bursts. We further proposed a location-spectrum iteration approach based on the Bayesian inference, which could alleviate the problems caused by the spectral difference between the burst and location templates. Our method is very suitable for scenarios with limited computation resources or time-sensitive applications, such as in-flight localization software, and low-latency localization for rapid follow-up observations.

astro-ph.HE

Magnet-free nonreciprocal metasurface for on-demand bi-directional phase modulation

Unconstrained by Lorentz reciprocity, nonreciprocal metasurfaces are uniquely capable of encoding distinctive optical functions on forward- and backward-propagating waves. The nonreciprocal metasurfaces reported to date require external electric or magnetic field biasing or rely on nonlinear effects, both of which are challenging to practically implement. Here, we propose and experimentally realize a magnet-free, linear, and passive nonreciprocal metasurface based on self-biased magnetic meta-atoms. Record transmittance up to 77% and operation angle reaching 64 degree are experimentally demonstrated. Moreover, on-demand bidirectional phase modulation in a "LEGO-like" manner is theoretically proposed and experimentally demonstrated, enabling a cohort of nonreciprocal functionalities such as microwave isolation, nonreciprocal beam steering, nonreciprocal focusing, and nonreciprocal holography. The design can also be extended to MHz and optical frequencies, taking advantage of the wide variety of self-biased gyrotropic materials available. We foresee that the nonreciprocal metasurfaces demonstrated in this work will have a significant practical impact for applications ranging from nonreciprocal antennas and radomes to full-duplex wireless communication and radar systems.

physics.optics

Gain stabilization and consistency correction approach for multiple SiPM-based gamma-ray detectors on GECAM

Each satellite of the Gravitational wave high-energy Electromagnetic Counterpart All-sky Monitor (GECAM, mission) consists of 25 SiPM based gamma-ray detectors (GRDs). Although SiPM based GRD has merits of compact size and low bias-voltage, the drift of the SiPM gain with temperature is a severe problem for GRD performance. An adaptive voltage supply source was designed to automatically adjust the SiPM bias voltage to compensate the temperature effects and keep the gain stable. This approach has been proved to be effective during both the on-ground and in-flight tests. The in-flight measured variation of the SiPM gain is within 2%. To reduce the gain non-uniformity of GRDs, an iterative bias voltage adjustment approach is proposed and implemented. The gain non-uniformity is reduced from 17% to 0.6%. In this paper, the gain stabilization and consistency correction approach are presented and discussed in detail.

physics.ins-det

Design and test of a portable Gamma-Ray Burst simulator for GECAM

The main scientific goal of the Gravitational wave high-energy Electromagnetic Counterpart All-sky Monitor (GECAM) is to monitor various types of Gamma-Ray Bursts (GRB) originated from merger of binary compact stars, which could also produce gravitational wave, and collapse of massive stars. In order to study the response of GECAM Gamma-Ray Detectors (GRDs) to high-energy bursts and test the in-flight trigger and localization software of GECAM before the launch, a portable GRB simulator device is designed and implemented based on grid controlled X-ray tube (GCXT) and direct digital synthesis (DDS) technologies. The design of this GRB simulator which modulates X-ray flux powered by high voltage up to 20 kV is demonstrated, and the time jitter (FWHM) of the device is about 0.9 $μ$s. Before the launch in December, 2020, both two GECAM satellites were irradiated by different types of GRBs (including short and long bursts in duration) generated by this GRB simulator. The light curves detected with GECAM/GRDs are consistent with the programmed input functions within statistical uncertainties, indicating the good performance of both the GRDs and the GRB simulator.

physics.ins-det

Switching the Optical Chirality in Magneto-plasmonic Metasurfaces Using Applied Magnetic Fields

Chiral nanophotonic devices are promising candidates for chiral molecules sensing, polarization diverse nanophotonics and display technologies. Active chiral nanophotonic devices, where the optical chirality can be controlled by an external stimulus has triggered great research interest. However, efficient modulation of the optical chirality has been challenging. Here, we demonstrate switching of the extrinsic chirality by applied magnetic fields in a magneto-plasmonic metasurface device based on a magneto-optical oxide material, Ce1Y2Fe5O12 (Ce:YIG). Thanks to the low optical loss and strong magneto-optical effect of Ce:YIG, we experimentally demonstrated a giant and continuous far-field circular dichroism (CD) modulation by applied magnetic fields from -0.65° to +1.9° at 950 nm wavelength under glancing incident conditions. The far field CD modulation is due to both magneto-optical circular dichroism and near-field modulation of the superchiral fields by applied magnetic fields. Finally, we demonstrate magnetic field tunable chiral imaging in millimeter-scale magneto-plasmonic metasurfaces fabricated using self-assembly. Our results provide a new way for achieving planar integrated, large-scale and active chiral metasurfaces for polarization diverse nanophotonics.

physics.optics

Alignment of the photoelectron spectroscopy beamline at NSRL

The photoelectron spectroscopy beamline at National Synchrotron Radiation Laboratory (NSRL) is equipped with a spherical grating monochromator with the included angle of 174 deg. Three gratings with line density of 200, 700 and 1200 lines/mm are used to cover the energy region from 60 eV to 1000 eV. After several years operation, the spectral resolution and flux throughput were deteriorated, realignment is necessary to improve the performance. First, the wavelength scanning mechanism, the optical components position and the exit slit guide direction are aligned according to the design value. Second, the gratings are checked by Atomic Force Microscopy (AFM). And then the gas absorption spectrum is measured to optimize the focusing condition of the monochromator. The spectral resolving power is recovered to the designed value of 1000@244eV. The flux at the end station for the 200 lines/mm grating is about 10^10 photons/sec/200mA, which is in accordance with the design. The photon flux for the 700 lines/mm grating is about 5 X 10^8 photons/sec/200mA, which is lower than expected. This poor flux throughput may be caused by carbon contamination on the optical components. The 1200 lines/mm grating has roughness much higher than expected so the diffraction efficiency is too low to detect any signal. A new grating would be ordered. After the alignment, the beamline has significant performance improvements in both the resolving power and the flux throughput for 200 and 700 lines/mm gratings and is provided to users.

physics.ins-det