SearcharxivSearch

arXiv subjects

Xinyue Liu

Publications and source records attributed to Xinyue Liu.

At least 19 recordsLinked to original sources

A Sharp Four-Layer Theorem for Integer-Occupied Slices of a Planar Disk-Slab

Let a rational positive-definite quadratic sublevel set on a full-rank affine lattice coset be intersected with a closed slab. A deterministic two-dimensional delta=3/4 reduction selects a primitive-dual integer functional and an inner ellipse. In the branch where the corresponding Babai point lies outside that ellipse, the lattice points in the disk-slab occupy at most four integer levels of the selected functional. An explicit rational instance attains four, so the bound is sharp. A strict diameter estimate reduces any counterexample to three blocks of five consecutive levels, and a nearest-integer inequality excludes the central block. Symmetry and cell localization reduce the two side blocks to eight modes: four follow from a shared-cap inequality, and four from an exact four-variable Bernstein certificate. The certificate contains 17,640 positive coefficients, with minimum 625/2048.

math.OC

SpecF2M: A Spectral-Aware Multi-task Network Estimating Axial Length and Refractive Error from Pediatric Fundus Photographs

Spherical Equivalent Refraction (SER) and Axial Length (AL) are core indicators for pediatric myopia screening, yet their measurements require dedicated biometry and cycloplegic refraction. Fundus photography offers an accessible imaging modality, as myopia-related posterior-pole changes are visible in 45$^\circ$ fundus images. However, these cues are often low-contrast, spatially diffuse, and multi-scale. Moreover, AL, Sphere (SPH), and Cylinder (CYL) share partially overlapping but non-identical anatomical correlates. We propose SpecF2M, a spectral-aware multi-task network for estimating AL and SER components from pediatric fundus photographs. SpecF2M integrates a deterministic anatomy-guided enhancement module, a hybrid spatial--spectral backbone combining MixCNN and Hybrid Spectral Learning (HSL) blocks, and an expert-routing head for component-level estimation of AL, SPH, and CYL. On a pediatric cohort of 4,359 eligible child visits and 6,966 fundus images, SpecF2M outperforms controlled CNN/ViT baselines for AL and SPH estimation, achieving MAEs of 0.5347 mm and 0.7062 D, respectively. Component-level analysis further reveals asymmetric task coupling, where CYL exhibits weaker association with fundus-derived myopic patterns than AL/SPH. These results support fundus-based, screening-oriented estimation of pediatric myopia indicators, while external validation remains necessary before deployment.

eess.IV

Onsager-variational-principle-based Lattice Boltzmann Model For Three-phase Dielectric Fluid Flows

Multiphase electrohydrodynamic (EHD) flows play a crucial role in various engineering applications. However, existing numerical studies on three-phase electrohydrodynamic systems predominantly rely on phenomenological models, often neglecting thermodynamic consistency and critical surface charge convection mechanisms. To address these fundamental gaps, this paper proposes a thermodynamically consistent three-phase EHD model derived strictly from the Onsager variational principle. This theoretical framework intrinsically guarantees thermodynamic consistency and accurately captures complex multiphysics interactions without requiring a priori assumptions. Furthermore, a mesoscopic lattice Boltzmann method is developed to solve the proposed model, enabling the natural capture of interfacial evolution and charge transport. The accuracy of the numerical framework are rigorously validated against several benchmark cases, including electroosmotic flow in microchannels, the spreading of a three-phase liquid lens, the equilibrium of static compound droplet, and the deformation of compound droplet under uniform electric field. Using this validated framework, we investigate EHD applications, specifically simulating the complex dynamics of double droplet coalescence and separation under electric field, as well as the behavior of droplets subjected to combined EHD and shear flow. Overall, this work provides a robust, thermodynamically reliable numerical tool for exploring the highly nonlinear behaviors of multiphase EHD systems.

physics.flu-dyn

DSETA: A Dual-Stage Continual Learning Framework for Travel Time Prediction in Dynamic Traffic Environments

Estimated Time of Arrival (ETA) prediction is a core component of intelligent transportation systems. As traffic congestion patterns become increasingly dynamic in large cities, maintaining high prediction accuracy poses a major challenge for ride-hailing platforms. Existing methods either fail to adapt to irregular traffic patterns and sudden congestion, or suffer from new distributions without disentangling long-term trends from short-term fluctuations, thereby degrading model performance in real-world scenarios. To address this challenge, we propose DSETA, an incrementally updated Dual-Stage ETA prediction framework. Specifically, the continual learning process is divided into \textit{inter-day} and \textit{intra-day} stages. We first design the \textit{intra-day} learning stage, which relies entirely on real-time data to enable dynamic adaptation to short-term traffic patterns caused by events like holidays or accidents. Next, we develop the \textit{inter-day} learning stage, which leverages aggregated historical data from a short time window to capture knowledge of long-term distribution shifts, such as seasonal trends and traffic network evolution. Subsequently, to prevent catastrophic forgetting and preserve knowledge of regular patterns, we explore a \textit{Historical Traffic Knowledge Consolidation} module. Finally, we validate DSETA's effectiveness and robustness through extensive offline and online experiments conducted on real-world datasets from DiDi's platform. Online A/B tests across three major cities including Beijing, Wuhan, and Xi'an consistently demonstrated performance gains, achieving MAE reductions of 6.62\%, 0.73\%, and 2.40\% respectively. This framework has been successfully deployed in DiDi's production environment, processing hundreds of millions of daily requests and validating its strong performance in industrial applications.

cs.LG

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/

cs.CV

Vera: Identity-Faithful Human Subject-to-Video Generation

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

cs.CV

Generative AI floods and dilutes the market for books

Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-text AI detection across 14,419 self-published genre-fiction books sold on Amazon from 2023 to 2026, matched to daily sales records through June 2026. None of these books disclose whether or not they contain AI-produced content. We find that books for which we detected substantial AI text ($>$ 25\%) make up a large share of the catalog but a smaller share of sales. Even so, they reach commercial scale, winning a growing share of sales over time and taking more of the scarce top-rank positions once held by books with no detected AI text. Over this period, the number of books with observed sales in a quarter grew 19.2-fold, while quarterly revenue grew only 8.9-fold. The market therefore added selling books faster than it added revenue, and revenue per selling book fell across most genres. Books with no AI text lose the most ground in genres with high AI diffusion, and most of all where Kindle Unlimited availability is high. Among top-selling books, those with substantial AI text draw on more distinctive language from existing books than do books with no AI text; for these books overlap rises with revenue, a gradient we do not detect for books with no AI text. Generative AI can thus reshape a creative market through scale rather than quality. Our results bear directly on the market-effect question at the center of the fair use defense to copyright infringement.

cs.CL

Res$^2$CLIP: Few-Shot Generalist Anomaly Detection with Residual-to-Residual Alignment

Few-shot Generalist Anomaly Detection requires models to generalize to novel categories without retraining, posing significant challenges in real-world scenarios with scarce samples and rapidly changing categories. Existing CLIP-based methods face two major challenges: coarse-grained unified text prompts struggle to adapt to fine-grained foreground-background differences, causing cross-granularity mismatch; and fine-tuning on auxiliary datasets disrupts CLIP's inherent open-world generalization due to domain shift, leading to cross-category generalization degradation. To address these, we propose to shift multimodal alignment entirely into a unified residual space, where residual representations naturally eliminate fine-grained normal feature differences across regions and class-specific biases, simultaneously resolving both problems. Based on this insight, Res$^2$CLIP, the first residual-to-residual alignment framework that symmetrically bridges visual and text modalities within CLIP's residual space, is designed. The framework is developed from a residual perspective into three branches: a text prompt-based branch, a visual prompt-based branch, and a novel residual-to-residual alignment branch. All learnable optimizations are constrained within the residual domain, and the residual alignment optimization objectives are designed to force the model to focus on relative anomaly deviations rather than optimizing class-specific features. Experiments on multiple datasets demonstrate the effectiveness of our architecture. The code is available at https://github.com/hito2448/Res2CLIP.

cs.CV

Meta Additive Model: Interpretable Sparse Learning With Auto Weighting

Sparse additive models have attracted much attention in high-dimensional data analysis due to their flexible representation and strong interpretability. However, most existing models are limited to single-level learning under the mean-squared error criterion, whose empirical performance can degrade significantly in the presence of complex noise, such as non-Gaussian perturbations, outliers, noisy labels, and imbalanced categories. The sample reweighting strategy is widely used to reduce the model's sensitivity to atypical data; however, it typically requires prespecifying the weighting functions and manually selecting additional hyperparameters. To address this issue, we propose a new meta additive model (MAM) based on the bilevel optimization framework, which learns data-driven weighting of individual losses by parameterizing the weighting function via an MLP trained on meta data. MAM is capable of a variety of learning tasks, including variable selection, robust regression estimation, and imbalanced classification. Theoretically, MAM provides guarantees on convergence in computation, algorithmic generalization, and variable selection consistency under mild conditions. Empirically, MAM outperforms several state-of-the-art additive models on both synthetic and real-world data under various data corruptions.

cs.LG

MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however, are limited to static images or single videos, overlooking the complex interactions across multiple videos. To address this gap, we introduce the Multi-Video Perception Evaluation Benchmark (MVPBench), a new benchmark featuring 14 subtasks across diverse visual domains designed to evaluate models on extracting relevant information from video sequences to make informed decisions. MVPBench includes 5K question-answering tests involving 2.7K video clips sourced from existing datasets and manually annotated clips. Extensive evaluations reveal that current models struggle to process multi-video inputs effectively, underscoring substantial limitations in their multi-video comprehension. We anticipate MVPBench will drive advancements in multi-video perception.

cs.CV

GenSpan: Generation-Calibrated Motion Span Priors for Multi-Verb Video Corpus Moment Retrieval

Video Corpus Moment Retrieval (VCMR) aims to retrieve both the correct video and its temporal segment corresponding to a natural-language query, a task that is especially challenging for multi-verb queries where temporal action ordering is critical. Existing approaches often rely solely on text or static images and struggle to capture implicit motion dynamics, leading to retrieval errors and temporal misalignment. We propose GenSpan, a generation-calibrated VCMR framework that constructs short auxiliary videos from LLM-selected subtitle cues and decomposed sub-events, using these as temporal priors rather than direct retrieval targets. A token selector filters candidate-video features aligned with generated motion, and a bidirectional state-space model efficiently predicts video-moment tuples. Experiments on TVR and ActivityNet-Captions demonstrate that GenSpan improves corpus-level retrieval and moment localization, particularly for complex multi-action queries, while reducing computational cost compared to state-of-the-art multimodal baselines.

cs.CV

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

Frontier LLM companies have repeatedly assured courts and regulators that their models do not store copies of training data. They further rely on safety alignment strategies via RLHF, system prompts, and output filters to block verbatim regurgitation of copyrighted works, and have cited the efficacy of these measures in their legal defenses against copyright infringement claims. We show that finetuning bypasses these protections: by training models to expand plot summaries into full text, a task naturally suited for commercial writing assistants, we cause GPT-4o, Gemini-2.5-Pro, and DeepSeek-V3.1 to reproduce up to 85-90% of held-out copyrighted books, with single verbatim spans exceeding 460 words, using only semantic descriptions as prompts and no actual book text. This extraction generalizes across authors: finetuning exclusively on Haruki Murakami's novels unlocks verbatim recall of copyrighted books from over 30 unrelated authors. The effect is not specific to any training author or corpus: random author pairs and public-domain finetuning data produce comparable extraction, while finetuning on synthetic text yields near-zero extraction, indicating that finetuning on individual authors' works reactivates latent memorization from pretraining. Three models from different providers memorize the same books in the same regions ($r \ge 0.90$), pointing to an industry-wide vulnerability. Our findings offer compelling evidence that model weights store copies of copyrighted works and that the security failures that manifest after finetuning on individual authors' works undermine a key premise of recent fair use rulings, where courts have conditioned favorable outcomes on the adequacy of measures preventing reproduction of protected expression.

cs.CL

Beyond False Discovery Rate: A Stepdown Group SLOPE Approach for Grouped Variable Selection

High-dimensional feature selection is routinely required to balance statistical power with strict control of multiple-error metrics such as the k-Family-Wise Error Rate (k-FWER) and the False Discovery Proportion (FDP), yet some existing frameworks are confined to the narrower goal of controlling the expected False Discovery Rate (FDR) and can not exploit the group-structure of the covariates, such as Sorted L-One Penalized Estimation (SLOPE). We introduce the Group Stepdown SLOPE, a unified optimization procedure which is capable of embedding the Lehmann-Romano stepdown rules into SLOPE to achieve finite-sample guarantees under k-FWER and FDP thresholds. Specifically, we derive closed-form regularization sequences under orthogonal designs that provably bound k-FWER and FDP at user-specified levels, and extend these results to grouped settings via gk-SLOPE and gF-SLOPE, which control the analogous group-level errors gk-FWER and gFDP. For non-orthogonal general designs, we provide a calibrated data-driven sequence inspired by Gaussian approximation and Monte-Carlo correction, preserving convexity and scalability. Extensive simulations are conducted across sparse, correlated, and group-structured regimes. Empirical results corroborate our theoretical findings that the proposed methods achieve nominal error control, while yielding markedly higher power than competing stepdown procedures, thereby confirming the practical value of the theoretical advances.

stat.ME

Boosting Meta-Learning for Few-Shot Text Classification via Label-guided Distance Scaling

Few-shot text classification aims to recognize unseen classes with limited labeled text samples. Existing approaches focus on boosting meta-learners by developing complex algorithms in the training stage. However, the labeled samples are randomly selected during the testing stage, so they may not provide effective supervision signals, leading to misclassification. To address this issue, we propose a \textbf{L}abel-guided \textbf{D}istance \textbf{S}caling (LDS) strategy. The core of our method is exploiting label semantics as supervision signals in both the training and testing stages. Specifically, in the training stage, we design a label-guided loss to inject label semantic information, pulling closer the sample representations and corresponding label representations. In the testing stage, we propose a Label-guided Scaler which scales sample representations with label semantics to provide additional supervision signals. Thus, even if labeled sample representations are far from class centers, our Label-guided Scaler pulls them closer to their class centers, thereby mitigating the misclassification. We combine two common meta-learners to verify the effectiveness of the method. Extensive experimental results demonstrate that our approach significantly outperforms state-of-the-art models. All datasets and codes are available at https://anonymous.4open.science/r/Label-guided-Text-Classification.

cs.LG

Optical and Hall conductivity of the two dimensional Hubbard model: effective theory description, sign-problem-free Monte Carlo simulation and applications to the cuprate superconductors

Exact formulas for the optical conductivity and the Hall conductivity of the two dimensional Hubbard model are derived in terms of an effective theory description of the local moment fluctuation in the system. In this framework, the quantum Monte Carlo simulation of the electromagnetic response of such a strongly correlated electron system becomes sign-problem-free in many physically relevant cases. In particular, it is sign-problem-free when we assume the widely used Millis-Monien-Pines form for the phenomenological susceptibility in the effective action of the fluctuating local moment, even though these local moments are now subjected to Landau damping as a result of their coupling to the itinerant quasiparticle on the fermi surface. This is true more generally when a $\varphi^{4}$ term is included in the effective action and is thus not restricted to the Gaussian limit. Here we demonstrate the power of this framework by studying the effect of thermal fluctuation of the local moment on the optical conductivity $\sigma^{xx}(\omega)$ and the Hall conductivity $\sigma^{xy}(\omega)$ of the cuprate superconductors. Both $\sigma^{xx}(\omega)$ and $\sigma^{xy}(\omega)$ calculated are found to exhibit a two-component structure, with a Drude component at low energy and a mid-infrared component at higher energy. Depending on the relative importance of the hole pocket and the electron pocket on the reconstructed fermi surface and the coupling strength to the local moment, the Drude component in $\mathrm{Im}\sigma^{xy}(\omega)$ can be either positive or negative.(full-length abstract can be found in the main text.)

cond-mat.str-el

Interplay between antiferromagnetic spin fluctuation and electron-phonon coupling and the origin of the peak-dip-hump structure in the anti-nodal spectrum of high-$T_{c}$ cuprate superconductors

Electron-phonon coupling is believed to be responsible for many spectral anomalies in the cuprate superconductors. In particular, the $B_{1g}$ buckling mode of the oxygen ion in the $CuO_{2}$ plane has been proposed to be responsible for the dramatic peak-dip-hump(PDH) structure in the anti-nodal spectrum. The recent observation of the exceptional flat quasiparticle dispersion in the anti-nodal region and the sudden suppression of the PDH structure around the pseudogap end point cast doubts on such a scenario. Instead, a scenario involving the coupling to the antiferromagnetic spin fluctuation seems to resolve both puzzles naturally. Here we present a systematic study on the interplay between antiferromagnetic spin fluctuation and electron-phonon coupling in the cuprate superconductors. We show that the coupling strength to the $B_{1g}$ buckling mode is strongly suppressed by the vertex correction caused by the antiferromagnetic spin fluctuation in the $\mathbf{q}\rightarrow 0$ limit as a result of the destructive interference between electron-phonon coupling at electron momentum differ by the antiferromagnetic wave vector. Counterintuitively, we find that the same vertex correction enhances the phonon contribution to the PDH structure. We also find that while the coupling to either the antiferromagnetic spin fluctuation or the $B_{1g}$ buckling mode can generate a PDH structure in the anti-nodal spectrum with similar phenomenologies, the sudden suppression of such a structure around the pseudogap end point should be mainly attributed to the dramatic change in the nature of the spin fluctuation at such a critical doping. We suggest to take the PDH structure in the anti-nodal spectrum as a spectral signature for the emergence of fluctuating local moment in the pseudogap phase and the entrance of a doped Mott insulating state.

cond-mat.str-el

From Narrative to Action: A Hierarchical LLM-Agent Framework for Human Mobility Generation

Understanding and replicating human mobility requires not only spatial-temporal accuracy but also an awareness of the cognitive hierarchy underlying real-world travel decisions. Traditional agent-based or deep learning models can reproduce statistical patterns of movement but fail to capture the semantic coherence and causal logic of human behavior. Large language models (LLMs) show potential, but struggle to balance creative reasoning with strict structural compliance. This study proposes a Hierarchical LLM-Agent Framework, termed Narrative-to-Action, that integrates high-level narrative reasoning, mid-level reflective planning, and low-level behavioral execution within a unified cognitive hierarchy. At the macro level, one agent is employed as a "creative writer" to produce diary-style narratives rich in motivation and context, then uses another agent as a "structural parser" to convert narratives into machine-readable plans. A dynamic execution module further grounds agents in geographic environments and enables adaptive behavioral adjustments guided by a novel occupation-aware metric, Mobility Entropy by Occupation (MEO), which captures heterogeneous schedule flexibility across different occupational personalities. At the micro level, the agent executes concrete actions-selecting locations, transportation modes, and time intervals-through interaction with an environmental simulation. By embedding this multi-layer cognitive process, the framework produces not only synthetic trajectories that align closely with real-world patterns but also interpretable representations of human decision logic. This research advances synthetic mobility generation from a data-driven paradigm to a cognition-driven simulation, providing a scalable pathway for understanding, predicting, and synthesizing complex urban mobility behaviors through hierarchical LLM agents.

cs.MA

Hexagonal InOI monolayer: a 2D phase-change material combining topological insulator states and piezoelectricity

Two-dimensional (2D) phase-change materials (PCMs) with moderate transition barriers and distinctly contrasting properties are highly desirable for multifunctional devices, yet such systems remain scarce. Using first-principles calculations, we propose a hexagonal InOI monolayer as a promising 2D PCM. This material exhibits two distinct polymorphs: an energetically favorable T$^{\prime}$ phase and a metastable T phase, differentiated by iodine atom positions. The T$^{\prime}$-to-T structural phase transition features a moderate energy barrier $E_b$ of 72.1 meV per formula unit, facilitating reversible switching. Notably, strain engineering tailors the electronic transition, inducing either a metal-to-topological-insulator or a metal-to-normal-insulator transformation. Additionally, this phase transition modulates the piezoelectric response and shifts optical absorption from the infrared to the visible range. These multifunctional properties make 2D hexagonal InOI highly promising for applications in non-volatile memory, low-contact-resistance spintronics, and optical switching devices.

cond-mat.mtrl-sci