SearcharxivSearch

arXiv subjects

Qing Yin

Publications and source records attributed to Qing Yin.

At least 19 recordsLinked to original sources

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. The generator is factorized causally in time, matching the causal structure of physical dynamics, and is aligned with a latent world-model reward for predictive consistency. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers 4K video generation at 24 FPS in real time, using an optimized GPU serving engine. In quantitative evaluations, Visko Orbis 1.0 achieves the best DOVER aesthetic and technical scores and the best VideoAlign visual and motion quality, and leads three physical-plausibility protocols (VideoPhy-2, Physics-IQ, and VBench-2.0 Physics); in long-form Arena comparisons, it obtains the highest overall-preference and temporal-stability ratings among all the state-of-the-art real-time interactive video generation systems.

cs.CV

Central Limit Theorem for a P\'olya-Friedman Mixed Urn Model

This paper considers a two-color, single-draw urn model with two types of balls, denoted type $1$ and type $2$, with initial counts $Y^1_0\in N^+$ and $Y^2_0\in N^+$, respectively. At each discrete time step, a ball is drawn uniformly at random, its type observed, and then it is returned to the urn. The urn is subsequently updated according to a mixed replacement matrix: with fixed probability $p\in(0,1)$, the Friedman replacement matrix is applied, adding $a$ balls of the drawn type and $b$ balls of the opposite type; with fixed probability $1-p\in (0,1)$, the P\'olya replacement matrix is applied, adding $c$ balls of the drawn type. We establish the central limit theorem for the proportion of type $1$ balls after $n$ draws. Furthermore, we provide corollaries that yield large deviation inequalities and the law of the iterated logarithm related to the proportion of type $1$ balls after $n$ draws.

math.PR

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content must adapt to dynamically evolving event conditions. A fundamental challenge lies in balancing reactivity and stability: models must respond promptly to new events while maintaining temporal coherence over long horizons. Existing approaches distill bidirectional models into autoregressive generators and further adapt them via streaming long tuning, yet often exhibit persistent drift after condition changes. We identify the cause as conditional bias, where the teacher may provide condition-aligned but trajectory-agnostic guidance, biasing generation toward locally valid yet globally inconsistent modes. Inspired by Trust Region Policy Optimization, we propose Delta Forcing, a simple yet effective framework that constrains unreliable teacher supervision within an adaptive trust region. Specifically, Delta Forcing estimates transition consistency from the latent delta between teacher and generator trajectories, and uses it to balance teacher supervision with a monotonic continuity objective. This suppress unreliable teacher-induced shifts while preserving responsiveness to new events. Extensive experiments demonstrate that Delta Forcing significantly improves consistency while maintaining event reactivity.

cs.CV

Moderate Deviation Principle for a Stochastic Approximation Process

In this paper, we investigate a stochastic approximation procedure $\left(X_n\right)_{n\ge 0}$ taking values in $R$. The process is adapted to a filtration $(F_n)_{n\ge 0}$ and satisfies the recursion $X_{n+1}=X_n+\frac{b}{n+1}\big[g(X_n)+U_{n+1}\big]$, where $b>0$, $g:R \to R$ is a function and $\left(U_n\right)_{n\ge 1}$ is a sequence of bounded martingale differences adapted to the filtration $(F_n)_{n\ge 1}$. We establish the moderate deviation principle for the stochastic process $(X_n)_{n\ge 0}$. As auxiliary results, we also obtain the exponential inequality for $(X_n)_{n\ge 0}$ and the moderate deviation principle for weighted sums of bounded martingale differences.

math.PR

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale human-annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision-language model judges that are not specialized for editing quality. We introduce VEFX-Dataset, a human-annotated dataset containing 5,049 video editing examples across 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on VEFX-Dataset, we propose VEFX-Reward, a reward model designed specifically for video editing quality assessment. VEFX-Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores via ordinal regression. We further release VEFX-Bench, a benchmark of 300 curated video-prompt pairs for standardized comparison of editing systems. Experiments show that VEFX-Reward aligns more strongly with human judgments than generic VLM judges and prior reward models on both standard IQA/VQA metrics and group-wise preference evaluation. Using VEFX-Reward as an evaluator, we benchmark representative commercial and open-source video editing systems, revealing a persistent gap between visual plausibility, instruction following, and edit locality in current models. Our project page is https://xiangbogaobarry.github.io/VEFX-Bench/.

cs.CV

PISCO: Precise Video Instance Insertion with Sparse Control

The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, it is crucial to perform precise, targeted modifications. A cornerstone of this transition is video instance insertion, which requires inserting a specific instance into existing footage while maintaining scene integrity. Unlike traditional video editing, this task demands several requirements: precise spatial-temporal placement, physically consistent scene interaction, and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To address the severe distribution shift induced by sparse conditioning in pretrained video diffusion models, we introduce Variable-Information Guidance for robust conditioning and Distribution-Preserving Temporal Masking to stabilize temporal generation, together with geometry-aware conditioning for realistic scene adaptation. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms strong inpainting and video editing baselines under sparse control, and exhibits clear, monotonic performance improvements as additional control signals are provided. Project page: xiangbogaobarry.github.io/PISCO.

cs.CV

Radio-Frequency Quantum Rectification in Kagome Superconductor CsV3Sb5

Rectification of electromagnetic fields into direct current (DC) is pivotal for energy harvesting, wireless charging, and next-generation communication technologies. The superconducting diode effect, which exploits the nonreciprocal transport of dissipationless superconducting currents, offers ultra-low power consumption and high rectification ratios. Combining the superconducting diode effect with the AC Josephson effect holds promise for converting radio-frequency (rf) irradiation into a quantized DC output. However, experimental realization has been hindered by challenges in achieving the necessary symmetry breaking and fabricating high-performance Josephson junctions. Here we demonstrate the quantum rectification in kagome superconductor CsV3Sb5, which hosts emergent Josephson effects and a zero-field Josephson diode. Under rf irradiation, a DC voltage emerges without applied bias, scaling linearly with frequency as V = hf/2e, where h is Planck's constant, f is the microwave frequency, and e is the electron charge. Furthermore, the rectified voltage exhibits quantized steps with increasing rf power, consistent with Shapiro step quantization. Our work establishes CsV3Sb5 as a versatile platform for wireless quantum power supplies and charging, and underscores the intertwined order parameters as a promising pathway for precise quantum matter control.

cond-mat.supr-con

Current-induced magnetoresistance hysteresis in the kagome superconductor CsV$_3$Sb$_5$

We report the observation of current-modulated magnetoresistance hysteresis below the superconducting transition temperature in the kagome superconductor CsV$_3$Sb$_5$. This highly tunable hysteresis behavior is confined to the superconducting state and vanishes when superconductivity is fully suppressed, directly linking magnetoresistance hysteresis to the superconducting order in CsV$_3$Sb$_5$. Additionally, the superconducting diode effect driven by a small magnetic field is observed, indicating the enhanced electronic magnetochiral anisotropy by the chiral domain-wall scattering. Our findings position CsV$_3$Sb$_5$ as a promising platform for exploring nontrivial physical phenomena, including unconventional pairing mechanisms and topological superconductivity.

cond-mat.supr-con

Nonlinear Valley and Spin Valves in Bilayer Graphene

Nonlinear transport plays a vital role in probing the quantum geometry of Bloch electrons, valley chirality, and carrier scattering mechanisms. The nonlinear Hall effect, characterized by a nonlinear scaling of Hall voltage with longitudinal current, has been explored to reveal the Berry curvature and quantum metric related physics. In this work, we extend the study of nonlinear transport to spin and valley degrees of freedom. Using bilayer graphene devices with Fe3GeTe2 contacts, we observe a second-order nonlinear spin current exhibiting spin valve-like behaviors. By tracking magnetic moment precession under an in-plane magnetic field, we identify a significantly enhanced critical magnetic field required for in-plane rotation, suggesting out-of-plane valley polarization induced by ferromagnetic proximity. These findings offer deep insights into the interplay of valley and spin in second-order nonlinear transport, opening avenues for promising device applications.

cond-mat.mes-hall

Orbital anomalous Hall effect in the few-layer Weyl semimetal TaIrTe4

We report on the observation of the linear anomalous Hall effect (AHE) in the nonmagnetic Weyl semimetal TaIrTe4. This is achieved by applying a direct current Idc and an alternating current Iac (Iac<<Idc) in TaIrTe4, where the former induces time-reversal symmetry breaking and the latter probes the triggered AHE. The anomalous Hall resistance VacH/Iac shows a linear dependence on Idc and changes sign with the polarity of Idc. In temperature-dependent measurements, VacH/Iac also experiences a sign reversal at 100 K, consistent with the temperature-dependent nonlinear Hall effect (NLHE). Furthermore, in measurements involving only dc transport, the dc Hall voltage exhibits a quadratic relationship with Idc. When the Idc direction is reversed, the Hall resistance changes sign, demonstrating a colossal nonreciprocal Hall effect (NRHE). Our theoretical calculations suggest that the observed linear AHE, NLHE, and NRHE all dominantly originate from the current-induced orbital magnetization compared to the minor spin contribution. This work provides deep insights into the orbital magnetoelectric effect and nonlinear Hall response, promising precise electric control of out-of-plane polarized orbit flow.

cond-mat.mes-hall

Berry-Esseen bound of modularity in network

In this paper, the model is a specific partition of a given network. Berry-Esseen bound and strong law of large numbers of modularity for the partition are proved when the size of the network gets large.

math.PR

Cramer's moderate deviations of modularity in network

Complex networks play a crucial role in understanding physical, biological, social and technological systems. One of the most relevant features of graphs representing real systems is community structure. In this paper, for a specific partition of a given network, we prove the Cramer's moderate deviations of modularity for the partition when the size of the network gets large.

math.PR

A Simple Yet Effective Approach for Diversified Session-Based Recommendation

Session-based recommender systems (SBRSs) have become extremely popular in view of the core capability of capturing short-term and dynamic user preferences. However, most SBRSs primarily maximize recommendation accuracy but ignore user minor preferences, thus leading to filter bubbles in the long run. Only a handful of works, being devoted to improving diversity, depend on unique model designs and calibrated loss functions, which cannot be easily adapted to existing accuracy-oriented SBRSs. It is thus worthwhile to come up with a simple yet effective design that can be used as a plugin to facilitate existing SBRSs on generating a more diversified list in the meantime preserving the recommendation accuracy. In this case, we propose an end-to-end framework applied for every existing representative (accuracy-oriented) SBRS, called diversified category-aware attentive SBRS (DCA-SBRS), to boost the performance on recommendation diversity. It consists of two novel designs: a model-agnostic diversity-oriented loss function, and a non-invasive category-aware attention mechanism. Extensive experiments on three datasets showcase that our framework helps existing SBRSs achieve extraordinary performance in terms of recommendation diversity and comprehensive performance, without significantly deteriorating recommendation accuracy compared to state-of-the-art accuracy-oriented SBRSs.

cs.IR

Mediation Analysis using Semi-parametric Shape-Restricted Regression with Applications

Often linear regression is used to perform mediation analysis. However, in many instances, the underlying relationships may not be linear, as in the case of placental-fetal hormones and fetal development. Although, the exact functional form of the relationship may be unknown, one may hypothesize the general shape of the relationship. For these reasons, we develop a novel shape-restricted inference-based methodology for conducting mediation analysis. This work is motivated by an application in fetal endocrinology where researchers are interested in understanding the effects of pesticide application on birth weight, with human chorionic gonadotropin (hCG) as the mediator. We assume a practically plausible set of nonlinear effects of hCG on the birth weight and a linear relationship between pesticide exposure and hCG, with both exposure-outcome and exposure-mediator models being linear in the confounding factors. Using the proposed methodology on a population-level prenatal screening program data, with hCG as the mediator, we discovered that, while the natural direct effects suggest a positive association between pesticide application and birth weight, the natural indirect effects were negative.

stat.ME

Beam Detection Based on Machine Learning Algorithms

The positions of free electron laser beams on screens are precisely determined by a sequence of machine learning models. Transfer training is conducted in a self-constructed convolutional neural network based on VGG16 model. Output of intermediate layers are passed as features to a support vector regression model. With this sequence, 85.8% correct prediction is achieved on test data.

physics.data-an

Understanding Diversity in Session-Based Recommendation

Current session-based recommender systems (SBRSs) mainly focus on maximizing recommendation accuracy, while few studies have been devoted to improve diversity beyond accuracy. Meanwhile, it is unclear how the accuracy-oriented SBRSs perform in terms of diversity. Besides, the asserted "trade-off" relationship between accuracy and diversity has been increasingly questioned in the literature. Towards the aforementioned issues, we conduct a holistic study to particularly examine the recommendation performance of representative SBRSs w.r.t. both accuracy and diversity, striving for better understanding the diversity-related issues for SBRSs and providing guidance on designing diversified SBRSs. Particularly, for a fair and thorough comparison, we deliberately select state-of-the-art non-neural, deep neural, and diversified SBRSs, by covering more scenarios with appropriate experimental setups, e.g., representative datasets, evaluation metrics, and hyper-parameter optimization technique. Our empirical results unveil that: 1) non-diversified methods can also obtain satisfying performance on diversity, which might even surpass diversified ones; and 2) the relationship between accuracy and diversity is quite complex. Besides the "trade-off" relationship, they might be positively correlated with each other, that is, having a same-trend (win-win or lose-lose) relationship, which varies across different methods and datasets. Additionally, we further identify three possible influential factors on diversity in SBRSs (i.e., granularity of item categorization, session diversity of datasets, and length of recommendation lists).

cs.IR

Label-dependent and event-guided interpretable disease risk prediction using EHRs

Electronic health records (EHRs) contain patients' heterogeneous data that are collected from medical providers involved in the patient's care, including medical notes, clinical events, laboratory test results, symptoms, and diagnoses. In the field of modern healthcare, predicting whether patients would experience any risks based on their EHRs has emerged as a promising research area, in which artificial intelligence (AI) plays a key role. To make AI models practically applicable, it is required that the prediction results should be both accurate and interpretable. To achieve this goal, this paper proposed a label-dependent and event-guided risk prediction model (LERP) to predict the presence of multiple disease risks by mainly extracting information from unstructured medical notes. Our model is featured in the following aspects. First, we adopt a label-dependent mechanism that gives greater attention to words from medical notes that are semantically similar to the names of risk labels. Secondly, as the clinical events (e.g., treatments and drugs) can also indicate the health status of patients, our model utilizes the information from events and uses them to generate an event-guided representation of medical notes. Thirdly, both label-dependent and event-guided representations are integrated to make a robust prediction, in which the interpretability is enabled by the attention weights over words from medical notes. To demonstrate the applicability of the proposed method, we apply it to the MIMIC-III dataset, which contains real-world EHRs collected from hospitals. Our method is evaluated in both quantitative and qualitative ways.

cs.AI

Label Dependent Attention Model for Disease Risk Prediction Using Multimodal Electronic Health Records

Disease risk prediction has attracted increasing attention in the field of modern healthcare, especially with the latest advances in artificial intelligence (AI). Electronic health records (EHRs), which contain heterogeneous patient information, are widely used in disease risk prediction tasks. One challenge of applying AI models for risk prediction lies in generating interpretable evidence to support the prediction results while retaining the prediction ability. In order to address this problem, we propose the method of jointly embedding words and labels whereby attention modules learn the weights of words from medical notes according to their relevance to the names of risk prediction labels. This approach boosts interpretability by employing an attention mechanism and including the names of prediction tasks in the model. However, its application is only limited to the handling of textual inputs such as medical notes. In this paper, we propose a label dependent attention model LDAM to 1) improve the interpretability by exploiting Clinical-BERT (a biomedical language model pre-trained on a large clinical corpus) to encode biomedically meaningful features and labels jointly; 2) extend the idea of joint embedding to the processing of time-series data, and develop a multi-modal learning framework for integrating heterogeneous information from medical notes and time-series health status indicators. To demonstrate our method, we apply LDAM to the MIMIC-III dataset to predict different disease risks. We evaluate our method both quantitatively and qualitatively. Specifically, the predictive power of LDAM will be shown, and case studies will be carried out to illustrate its interpretability.

cs.AI