SearcharxivSearch

arXiv subjects

Jia Shi

Publications and source records attributed to Jia Shi.

At least 19 recordsLinked to original sources

Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence of their own: removing the class name from the prompt collapses ImageNet accuracy from 59.5% to 15.5%. The diagnosis is that the descriptors are conditioned on the label rather than on the images, so they describe the concept in general and mislead exactly when the data shifts; an LLM insists that strawberries are red, but every strawberry in ImageNet-Sketch is a colorless line drawing. We therefore select attributes from the target image collection instead: we score a large attribute pool against the images in CLIP's joint embedding space and keep the top-scoring attributes per class. Selected this way, class-name-free attribute prompts reach 23.8% on ImageNet (against 15.5% for LLM descriptors), the gain holds on four shifted ImageNet variants, and reselecting from the LLM's own pool isolates the selection mechanism as the cause. With one image per class, the selected attributes outperform the prompt-tuning method CoOp by 3 points while fitting in under a minute instead of 14 hours, with no learned soft prompt to obscure the decision. Because the attribute set is chosen by the data, it doubles as a readable summary of a dataset, which we use to describe distribution shift in words. Our code and results are available on our project page: https://ggare-cmu.github.io/AttributeSelect/

cs.CV

Abelian surfaces of small conductor from genus 3 double covers

We describe the construction of a database of roughly half a million abelian surfaces over Q of small conductor arising as Pryms associated to a genus 3 double cover of a genus 1 curve. Our construction uses a method that provides a degree of control over the primes of bad reduction.

math.NT

SubFlow: Sub-mode Conditioned Flow Matching for Diverse One-Step Generation

Flow matching has emerged as a powerful generative framework, with recent few-step methods achieving remarkable inference acceleration. However, we identify a critical yet overlooked limitation: these models suffer from severe diversity degradation, concentrating samples on dominant modes while neglecting rare but valid variations of the target distribution. We trace this degradation to averaging distortion: when trained with MSE objectives, class-conditional flows learn a frequency-weighted mean over intra-class sub-modes, causing the model to over-represent high-density modes while systematically neglecting low-density ones. To address this, we propose SubFlow, Sub-mode Conditioned Flow Matching, which eliminates averaging distortion by decomposing each class into fine-grained sub-modes via semantic clustering and conditioning the flow on sub-mode indices. Each conditioned sub-distribution is approximately unimodal, so the learned flow accurately targets individual modes with no averaging distortion, restoring full mode coverage in a single inference step. Crucially, SubFlow is entirely plug-and-play: it integrates seamlessly into existing one-step models such as MeanFlow and Shortcut Models without any architectural modifications. Extensive experiments on ImageNet-256 demonstrate that SubFlow yields substantial gains in generation diversity (Recall) while maintaining competitive image quality (FID), confirming its broad applicability across different one-step generation frameworks. Project page: https://yexionglin.github.io/subflow.

cs.LG

Lifting $L$-polynomials of genus $3$ curves

Let $C$ be a smooth plane quartic curve over $\mathbb{Q}$. Costa, Harvey and Sutherland provide an algorithm with an implementation, improving Harvey's average polynomial-time algorithm, to compute the $\bmod \ p$ reduction of the numerator of the zeta function of $C$ at all $p\leq B$, where $p$ is an odd prime of good reduction, in $O(B\log^{3+o(1)} N)$ time, which is $O(\log^{4+o(1)}p)$ time on average per prime. Alternatively, their algorithm can do this for a single prime $p$ of good reduction in $O(p^{1/2}\log^2p)$ time. While this algorithm can be used to compute the full zeta function, no implementation of this step currently exists. In this article, we provide an algorithm and an implementation for the group operation on the Jacobian of $C$ over $\mathbb{F}_p$, where $p$ is an odd prime of good reduction. We provide a Las Vegas algorithm that takes the $\bmod \ p$ result of Costa, Harvey and Sutherland's algorithm and uses it to compute the full zeta function. The expected running time of the algorithm is bounded by $O(p^{1/2+o(1)})$, and under heuristic assumptions, we prove an $O(p^{1/4+o(1)})$ bound on its average running time (over all inputs). Our lifting algorithm can also be applied to hyperelliptic curves of genus 3.

math.NT

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

End-to-End autonomous driving (E2E-AD) has emerged as a new paradigm, where trajectory planning plays a crucial role. Existing studies mainly follow two directions: trajectory generation oriented, which focuses on producing high-quality trajectories with simple decision mechanisms, and trajectory selection oriented, which performs multi-dimensional evaluation to select the best trajectory yet lacks sufficient generative capability. In this work, we propose MindDrive, a harmonized framework that integrates high-quality trajectory generation with comprehensive decision reasoning. It establishes a structured reasoning paradigm of "context simulation - candidate generation - multi-objective trade-off". In particular, the proposed Future-aware Trajectory Generator (FaTG), based on a World Action Model (WaM), performs ego-conditioned "what-if" simulations to predict potential future scenes and generate foresighted trajectory candidates. Building upon this, the VLM-oriented Evaluator (VLoE) leverages the reasoning capability of a large vision-language model to conduct multi-objective evaluations across safety, comfort, and efficiency dimensions, leading to reasoned and human-aligned decision making. Extensive experiments on the NAVSIM-v1 and NAVSIM-v2 benchmarks demonstrate that MindDrive achieves state-of-the-art performance across multi-dimensional driving metrics, significantly enhancing safety, compliance, and generalization. This work provides a promising path toward interpretable and cognitively guided autonomous driving.

cs.CV

Lifting $L$-polynomials of genus 2 curves

Let $C$ be a genus $2$ curve over $\mathbb{Q}$. Harvey and Sutherland's implementation of Harvey's average polynomial-time algorithm computes the $\bmod \ p$ reduction of the numerator of the zeta function of $C$ at all good primes $p\leq B$ in $O(B\log^{3+o(1)}B)$ time, which is $O(\log^{4+o(1)} p)$ time on average per prime. Alternatively, their algorithm can do this for a single good prime $p$ in $O(p^{1/2}\log^{1+o(1)}p)$ time. While Harvey's algorithm can also be used to compute the full zeta function, no practical implementation of this step currently exists. In this article, we present an $O(\log^{2+o(1)}p)$ Las Vegas algorithm that takes the $\bmod \ p$ output of Harvey and Sutherland's implementation and outputs the full zeta function. We then benchmark our results against the fastest algorithms currently available for computing the full zeta function of a genus~$2$ curve, finding substantial speedups in both the average polynomial-time and single prime settings.

math.NT

Reconfigurable Intelligent Surface-Enabled Green and Secure Offloading for Mobile Edge Computing Networks

This paper investigates a multi-user uplink mobile edge computing (MEC) network, where the users offload partial tasks securely to an access point under the non-orthogonal multiple access policy with the aid of a reconfigurable intelligent surface (RIS) against a multi-antenna eavesdropper. We formulate a non-convex optimization problem of minimizing the total energy consumption subject to secure offloading requirement, and we build an efficient block coordinate descent framework to iteratively optimize the number of local computation bits and transmit power at the users, the RIS phase shifts, and the multi-user detection matrix at the access point. Specifically, we successively adopt successive convex approximation, semi-definite programming, and semidefinite relaxation to solve the problem with perfect eavesdropper's channel state information (CSI), and we then employ S-procedure and penalty convex-concave to achieve robust design for the imperfect CSI case. We provide extensive numerical results to validate the convergence and effectiveness of the proposed algorithms. We demonstrate that RIS plays a significant role in realizing a secure and energy-efficient MEC network, and deploying a well-designed RIS can save energy consumption by up to 60\% compared to that without RIS. We further reveal impacts of various key factors on the secrecy energy efficiency, including RIS element number and deployment position, user number, task scale and duration, and CSI imperfection.

cs.IT

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

Recent breakthroughs in text-to-speech (TTS) voice cloning have raised serious privacy concerns, allowing highly accurate vocal identity replication from just a few seconds of reference audio, while retaining the speaker's vocal authenticity. In this paper, we introduce CloneShield, a universal time-domain adversarial perturbation framework specifically designed to defend against zero-shot voice cloning. Our method provides protection that is robust across speakers and utterances, without requiring any prior knowledge of the synthesized text. We formulate perturbation generation as a multi-objective optimization problem, and propose Multi-Gradient Descent Algorithm (MGDA) to ensure the robust protection across diverse utterances. To preserve natural auditory perception for users, we decompose the adversarial perturbation via Mel-spectrogram representations and fine-tune it for each sample. This design ensures imperceptibility while maintaining strong degradation effects on zero-shot cloned outputs. Experiments on three state-of-the-art zero-shot TTS systems, five benchmark datasets and evaluations from 60 human listeners demonstrate that our method preserves near-original audio quality in protected inputs (PESQ = 3.90, SRS = 0.93) while substantially degrading both speaker similarity and speech quality in cloned samples (PESQ = 1.07, SRS = 0.08).

cs.SD

Nucleation and Antiphase Twin Control in Bi$_2$Se$_3$ via Step-Terminated Al$_2$O$_3$ Substrates

The epitaxial synthesis of high-quality 2D layered materials is an essential driver of both fundamental physics studies and technological applications. Bi$_2$Se$_3$, a prototypical 2D layered topological insulator, is sensitive to defects imparted during the growth, either thermodynamically or due to the film-substrate interaction. In this study, it is shown that step-terminated Al$_2$O$_3$ substrates with a high miscut angle (3{\deg}) can effectively suppress a particular hard-to-mitigate defect, the antiphase twin. Systematic investigations across a range of growth temperatures and substrate miscut angles confirm that atomic step edges act as preferential nucleation sites, stabilizing a single twin domain. First-principles calculations suggest that there is a significant energy barrier for twin boundary formation at step edges, supporting the experimental observations. Detailed structural characterization indicates that this twin-selectivity is lost through the mechanism of the 2D layers overgrowing the step edges, leading to higher twin density as the thickness increases. These findings highlight the complex energy landscape unique to 2D materials that is driven by the interplay between substrate properties, nucleation dynamics, and defect formation, and overcoming and controlling these are critical to improve material quality for quantum and electronic applications.

cond-mat.mtrl-sci

Joint Sparse Graph for Enhanced MIMO-AFDM Receiver Design

Affine frequency division multiplexing (AFDM) is a promising chirp-assisted multicarrier waveform for future high-mobility communications. This paper is devoted to enhanced receiver design for multiple input and multiple output AFDM (MIMO-AFDM) systems. Firstly, we introduce a unified variational inference (VI) approach to approximate the target posterior distribution, under which the belief propagation (BP) and expectation propagation (EP)-based algorithms are derived. As both VI-based detection and low-density parity-check (LDPC) decoding can be expressed by bipartite graphs in MIMO-AFDM systems, we construct a joint sparse graph (JSG) by merging the graphs of these two for low-complexity receiver design. Then, based on this graph model, we present the detailed message propagation of the proposed JSG. Additionally, we propose an enhanced JSG (E-JSG) receiver based on the linear constellation encoding model. The proposed E-JSG eliminates the need for interleavers, de-interleavers, and log-likelihood ratio transformations, thus leading to concurrent detection and decoding over the integrated sparse graph. To further reduce detection complexity, we introduce a sparse channel method by approaximating multiple graph edges with insignificant channel coefficients into a single edge on the VI graph. Simulation results show the superiority of the proposed receivers in terms of computational complexity, detection and decoding latency, and error rate performance compared to the conventional ones.

eess.SP

A note on the existence of self-similar profiles of the hydrodynamic formulation of the focusing nonlinear Schr\"odinger equation

After performing the Madelung transformation, the nonlinear Schr\"odinger equation is transformed into a hydrodynamic equation akin to the compressible Euler equations with a certain dissipation. In this short note, we construct self-similar solutions of such system in the focusing case for any mass supercritical exponent. To the best of our knowledge these solutions are new, and may formally arise as potential blow-up profiles of the focusing NLS equation.

math.AP

Giant Magneto-Exciton Coupling in 2D van der Waals CrSBr

Controlling magnetic order via external fields or heterostructures enables precise manipulation and tracking of spin and exciton information, facilitating the development of high-performance optical spin valves. However, the weak magneto-optical signals and instability of two dimensional (2D) antiferromagnetic (AFM) materials have hindered comprehensive studies on the complex coupling between magnetic order and excitons in bulk-like systems. Here, we leverage magneto-optical spectroscopy to reveal the impact of magnetic order on exciton-phonon coupling and exciton-magnetic order coupling which remains robust even under non-extreme temperature conditions (80 K) in thick layered CrSBr. A 0.425T in-plane magnetic field is sufficient to induce spin flipping and transition from AFM to ferromagnetic (FM) magnetic order in CrSBr, while magnetic circular dichroism (MCD) spectroscopy under an out-of-plane magnetic field provides direct insight into the complex spin canting behavior in thicker layers. Theoretical calculations reveal that the strong coupling between excitons and magnetic order, especially the 32 meV exciton energy shift during magnetic transitions, stems from the hybridization of Cr and S orbitals and the larger exciton wavefunction radius of higher-energy B excitons. These findings offer new opportunities and a solid foundation for future exploration of 2D AFM materials in magneto-optical sensors and quantum communication using excitons as spin carriers.

cond-mat.mtrl-sci

LCA-on-the-Line: Benchmarking Out-of-Distribution Generalization with Class Taxonomies

We tackle the challenge of predicting models' Out-of-Distribution (OOD) performance using in-distribution (ID) measurements without requiring OOD data. Existing evaluations with "Effective Robustness", which use ID accuracy as an indicator of OOD accuracy, encounter limitations when models are trained with diverse supervision and distributions, such as class labels (Vision Models, VMs, on ImageNet) and textual descriptions (Visual-Language Models, VLMs, on LAION). VLMs often generalize better to OOD data than VMs despite having similar or lower ID performance. To improve the prediction of models' OOD performance from ID measurements, we introduce the Lowest Common Ancestor (LCA)-on-the-Line framework. This approach revisits the established concept of LCA distance, which measures the hierarchical distance between labels and predictions within a predefined class hierarchy, such as WordNet. We assess 75 models using ImageNet as the ID dataset and five significantly shifted OOD variants, uncovering a strong linear correlation between ID LCA distance and OOD top-1 accuracy. Our method provides a compelling alternative for understanding why VLMs tend to generalize better. Additionally, we propose a technique to construct a taxonomic hierarchy on any dataset using K-means clustering, demonstrating that LCA distance is robust to the constructed taxonomic hierarchy. Moreover, we demonstrate that aligning model predictions with class taxonomies, through soft labels or prompt engineering, can enhance model generalization. Open source code in our Project Page: https://elvishelvis.github.io/papers/lca/.

cs.LG

The regularity of the solutions to the Muskat equation: the degenerate regularity near the turnover points

In this paper, we prove that if a solution to the Muskat problem with different densities and the same viscosity is sufficiently smooth, the solution is analytic in a region that degenerates at the turnover points, provided some additional conditions are satisfied. This paper studies the analyticity of the solution near turnover points, complementing the result in $[$Jia Shi. Regularity of solutions to the Muskat equation, Arch Rational Mech Anal 247, 36 (2023)$]$.

math.AP

Non-radial implosion for compressible Euler and Navier-Stokes in $\mathbb{T}^3$ and $\mathbb{R}^3$

In this paper we construct smooth, non-radial solutions of the compressible Euler and Navier-Stokes equation that develop an imploding finite time singularity. Our construction is motivated by the works [Merle, Rapha\"{e}l, Rodnianski, and Szeftel, Ann. of Math., 196(2):567-778, 2022, Ann. of Math., 196(2):779-889, 2022], [Buckmaster, Cao-Labora, and G\'{o}mez-Serrano, arXiv:2208.09445, 2022], but is flexible enough to handle both periodic and non-radial initial data.

math.AP

Thickness dependence of superconductivity in FeSe films

The films of FeSe on substrates have attracted attention because of their unusually high-temperature (Tc) superconducting properties whose origins continue to be debated. To disentangle the competing effects of the substrate and interlayer and intralayer processes, we present here results of density functional theory (DFT)-based analysis of the electronic structure of unsupported FeSe films consisting of 1 to 5 layers (1L-5L). Furthermore, by solving the Bardeen-Schrieffer-Cooper (BCS) equation with spin-wave exchange attraction derived from the Hubbard model, we find the superconducting critical temperature Tc for 1L-5L and bulk FeSe systems in reasonable agreement with experimental data. Our results point to the importance of correlation effects in superconducting properties of single- and multi-layer FeSe films, independently of the role of substrate.

cond-mat.supr-con

Dark exciton energy splitting in monolayer WSe2: insights from time-dependent density-functional theory

We present here a formalism based on time-dependent density-functional theory (TDDFT) to describe characteristics of both intra- and inter-valley excitons in semiconductors, the latter of which had remained a challenge. Through the usage of an appropriate exchange-correlation kernel (nanoquanta), we trace the energy difference between the intra- and inter-valley dark excitons in monolayer (1L) WSe2 to the domination of the exchange part in the exchange-correlation energies of these states. Furthermore, our calculated transition contribution maps establish the momentum resolved weights of the electron-hole excitations in both bright and dark excitons thereby providing a comprehensive understanding of excitonic properties of 1L WSe2. We find that the states consist of hybridized excitations around the corresponding valleys which leads to brightening of the dark excitons, i.e., significantly decreasing their lifetime which is reflected in the PL spectrum. Using many-body perturbation theory, we calculate the phonon contribution to the energy bandgap and the linewidths of the excited electrons, holes and (bright) exciton to find that as the temperature increases the bandgap significantly decreases, while the linewidths increase. Our work paves for describing the ultrafast charge dynamics of transition metal dichalcogenide within an ab initio framework.

cond-mat.mtrl-sci