Searcharxiv⌕ Search

arXiv subjects

Xiang Fan

Publications and source records attributed to Xiang Fan.

At least 37 records · Page 2Linked to original sources

Contrastive Flow Matching

Unconditional flow-matching trains diffusion models to transport samples from a source distribution to a target distribution by enforcing that the flows between sample pairs are unique. However, in conditional settings (e.g., class-conditioned models), this uniqueness is no longer guaranteed--flows from different conditions may overlap, leading to more ambiguous generations. We introduce Contrastive Flow Matching, an extension to the flow matching objective that explicitly enforces uniqueness across all conditional flows, enhancing condition separation. Our approach adds a contrastive objective that maximizes dissimilarities between predicted flows from arbitrary sample pairs. We validate Contrastive Flow Matching by conducting extensive experiments across varying model architectures on both class-conditioned (ImageNet-1k) and text-to-image (CC3M) benchmarks. Notably, we find that training models with Contrastive Flow Matching (1) improves training speed by a factor of up to 9x, (2) requires up to 5x fewer de-noising steps and (3) lowers FID by up to 8.9 compared to training the same models with flow matching. We release our code at: https://github.com/gstoica27/DeltaFM.git.

cs.CV↗

Left-right splitting of elliptic flow in heavy ion collisions: TRENTo-3D initialization and CLVisc hydrodynamic simulations

Using the TRENTo-3D initial condition model coupled with (3+1)-dimensional CLVisc hydrodynamic simulations, we systematically investigate the left-right splitting of elliptic flow ($Δv_{2}$) for soft particles in relativistic heavy-ion collisions. Our study reveals that the final distribution characteristics of $Δv_{2}$ are primarily depend on the odd flow harmonics and $v_{2}$ itself. We find that the parton transverse momentum scale $k_\mathrm{T}$ not only determines the geometric tilt of the QGP fireball but also significantly affects the rapidity dependence of both $v_1$ and $Δv_{2}$, providing new insights into the splitting mechanism of $Δv_{2}$. Furthermore, our results demonstrate that $Δv_{2} (p_\mathrm{T})$ exhibits significant sensitivity to influences such as the sub-nucleonic degrees of freedom (or `hotspots'), transverse momentum scale, and fragmentation region profile. By analyzing the $Δv_{2}$ and $Δv_{2}/v_{2}$ ratio, our findings provide new constraints on the uncertainties of the QGP initial state and provide additional constraints for refining model parameters.

nucl-th↗

A Model Stealing Attack Against Multi-Exit Networks

Compared to traditional neural networks with a single output channel, a multi-exit network has multiple exits that allow for early outputs from the model's intermediate layers, thus significantly improving computational efficiency while maintaining similar main task accuracy. Existing model stealing attacks can only steal the model's utility while failing to capture its output strategy, i.e., a set of thresholds used to determine from which exit to output. This leads to a significant decrease in computational efficiency for the extracted model, thereby losing the advantage of multi-exit networks. In this paper, we propose the first model stealing attack against multi-exit networks to extract both the model utility and the output strategy. We employ Kernel Density Estimation to analyze the target model's output strategy and use performance loss and strategy loss to guide the training of the extracted model. Furthermore, we design a novel output strategy search algorithm to maximize the consistency between the victim model and the extracted model's output behaviors. In experiments across multiple multi-exit networks and benchmark datasets, our method always achieves accuracy and efficiency closest to the victim models.

cs.CR↗

Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion

We introduce Videoshop, a training-free video editing algorithm for localized semantic edits. Videoshop allows users to use any editing software, including Photoshop and generative inpainting, to modify the first frame; it automatically propagates those changes, with semantic, spatial, and temporally consistent motion, to the remaining frames. Unlike existing methods that enable edits only through imprecise textual instructions, Videoshop allows users to add or remove objects, semantically change objects, insert stock photos into videos, etc. with fine-grained control over locations and appearance. We achieve this through image-based video editing by inverting latents with noise extrapolation, from which we generate videos conditioned on the edited image. Videoshop produces higher quality edits against 6 baselines on 2 editing benchmarks using 10 evaluation metrics.

cs.CV↗

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interactions that are necessary to solve a multimodal task? Subsequently, what are the most suitable multimodal models to capture these interactions? To answer these questions, we propose an information-theoretic approach to quantify the degree of redundancy, uniqueness, and synergy relating input modalities with an output task. We term these three measures as the PID statistics of a multimodal distribution (or PID for short), and introduce two new estimators for these PID statistics that scale to high-dimensional distributions. To validate PID estimation, we conduct extensive experiments on both synthetic datasets where the PID is known and on large-scale multimodal benchmarks where PID estimations are compared with human annotations. Finally, we demonstrate their usefulness in (1) quantifying interactions within multimodal datasets, (2) quantifying interactions captured by multimodal models, (3) principled approaches for model selection, and (4) three real-world case studies engaging with domain experts in pathology, mood prediction, and robotic perception where our framework helps to recommend strong multimodal models for each application.

cs.LG↗

Nano: Nested Human-in-the-Loop Reward Learning for Few-shot Language Model Control

Pretrained language models have demonstrated extraordinary capabilities in language generation. However, real-world tasks often require controlling the distribution of generated text in order to mitigate bias, promote fairness, and achieve personalization. Existing techniques for controlling the distribution of generated text only work with quantified distributions, which require pre-defined categories, proportions of the distribution, or an existing corpus following the desired distributions. However, many important distributions, such as personal preferences, are unquantified. In this work, we tackle the problem of generating text following arbitrary distributions (quantified and unquantified) by proposing Nano, a few-shot human-in-the-loop training algorithm that continuously learns from human feedback. Nano achieves state-of-the-art results on single topic/attribute as well as quantified distribution control compared to previous works. We also show that Nano is able to learn unquantified distributions, achieves personalization, and captures differences between different individuals' personal preferences with high sample efficiency.

cs.CL↗

High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning

Many real-world problems are inherently multimodal, from spoken language, gestures, and paralinguistics humans use to communicate, to force, proprioception, and visual sensors on robots. While there has been an explosion of interest in multimodal learning, these methods are focused on a small set of modalities primarily in language, vision, and audio. In order to accelerate generalization towards diverse and understudied modalities, this paper studies efficient representation learning for high-modality scenarios involving a large set of diverse modalities. Since adding new models for every new modality becomes prohibitively expensive, a critical technical challenge is heterogeneity quantification: how can we measure which modalities encode similar information and interactions in order to permit parameter sharing with previous modalities? This paper proposes two new information theoretic metrics for heterogeneity quantification: (1) modality heterogeneity studies how similar 2 modalities {X1,X2} are by measuring how much information can be transferred from X1 to X2, while (2) interaction heterogeneity studies how similarly pairs of modalities {X1,X2}, {X3,X4} interact by measuring how much information can be transferred from fusing {X1,X2} to {X3,X4}. We show the importance of these 2 proposed metrics as a way to automatically prioritize the fusion of modalities that contain unique information or interactions. The result is a single model, HighMMT, that scales up to 10 modalities (text, image, audio, video, sensors, proprioception, speech, time-series, sets, and tables) and 15 tasks from 5 research areas. Not only does HighMMT outperform prior methods on the tradeoff between performance and efficiency, it also demonstrates a crucial scaling behavior: performance continues to improve with each modality added, and it transfers to entirely new modalities and tasks during fine-tuning.

cs.LG↗

MultiZoo & MultiBench: A Standardized Toolkit for Multimodal Deep Learning

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robustness, we release MultiZoo, a public toolkit consisting of standardized implementations of > 20 core multimodal algorithms and MultiBench, a large-scale benchmark spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas. Together, these provide an automated end-to-end machine learning pipeline that simplifies and standardizes data loading, experimental setup, and model evaluation. To enable holistic evaluation, we offer a comprehensive methodology to assess (1) generalization, (2) time and space complexity, and (3) modality robustness. MultiBench paves the way towards a better understanding of the capabilities and limitations of multimodal models, while ensuring ease of use, accessibility, and reproducibility. Our toolkits are publicly available, will be regularly updated, and welcome inputs from the community.

cs.LG↗

MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective computing, robotics, finance, human-computer interaction, and healthcare. Unfortunately, multimodal research has seen limited resources to study (1) generalization across domains and modalities, (2) complexity during training and inference, and (3) robustness to noisy and missing modalities. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robustness, we release MultiBench, a systematic and unified large-scale benchmark spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas. MultiBench provides an automated end-to-end machine learning pipeline that simplifies and standardizes data loading, experimental setup, and model evaluation. To enable holistic evaluation, MultiBench offers a comprehensive methodology to assess (1) generalization, (2) time and space complexity, and (3) modality robustness. MultiBench introduces impactful challenges for future research, including scalability to large-scale multimodal datasets and robustness to realistic imperfections. To accompany this benchmark, we also provide a standardized implementation of 20 core approaches in multimodal learning. Simply applying methods proposed in different research areas can improve the state-of-the-art performance on 9/15 datasets. Therefore, MultiBench presents a milestone in unifying disjoint efforts in multimodal research and paves the way towards a better understanding of the capabilities and limitations of multimodal models, all the while ensuring ease of use, accessibility, and reproducibility. MultiBench, our standardized code, and leaderboards are publicly available, will be regularly updated, and welcomes inputs from the community.

cs.LG↗

Probing stops in the coannihilation region at the HL-LHC: a comparative study of different processes

In the minimal supersymmetric model, the coannihilation of the lighter stop $\tilde{t}_1$ and bino-like dark matter $χ$ provides a feasible way to accommodate the correct dark matter relic abundance. In this scenario, due to the compressed masses, $\tilde{t}_1$ merely appears as missing energy at the LHC and thus the pair production of $\tilde{t}_1$ can only be probed by requiring an associated energetic jet. Meanwhile, since $\tilde{t}_2$ and $\tilde{b}_1$ are correlated in mass and mixing with $\tilde{t}_1$, the production of $\tilde{t}_2\tilde{t}_2^*$ or $\tilde{b}_1\tilde{b}_1^*$, each of which dominantly decays into $\tilde{t}_1$ plus $Z$, $h$ or $W$ boson, may serve as a complementary probe. We examine all these processes at the HL-LHC and find that the $2σ$ sensitivity to $χ$ mass can be as large as about 570 GeV, 600 GeV and 1.1 TeV from the production process of $\tilde{t}_1\tilde{t}_1^*+{\rm jet}$, $\tilde{t}_2\tilde{t}_2^*$ and $\tilde{b}_1\tilde{b}_1^*$, respectively.

hep-ph↗

Permutation polynomials of degree 8 over finite fields of odd characteristic

This paper provides an algorithmic generalization of Dickson's method of classifying permutation polynomials (PPs) of a given degree $d$ over finite fields. Dickson's idea is to formulate from Hermite's criterion several polynomial equations satisfied by the coefficients of an arbitrary PP of degree $d$. Previous classifications of PPs of degree at most $6$ were essentially deduced from manual analysis of these polynomial equations. However, these polynomials, needed for that purpose when $d>6$, are too complicated to solve. Our idea is to make them more solvable by calculating some radicals of ideals generated by them, implemented by a computer algebra system (CAS). Our algorithms running in SageMath 8.6 on a personal computer work very fast to determine all PPs of degree $8$ over an arbitrary finite field of odd order $q>8$. The main result is that for an odd prime power $q>8$, a PP $f$ of degree $8$ exists over the finite field of order $q$ if and only if $q\leqslant 31$ and $q\not\equiv 1\ (\mathrm{mod}\ 8)$, and $f$ is explicitly listed up to linear transformations.

math.NT↗

CHNS: A case study of turbulence in elastic media

Recent progress in the study of Cahn-Hilliard Navier-Stokes (CHNS) turbulence is summarized. This is an example of \textit{elastic turbulence}, which can occur in elastic (i.e. self-restoring) media. Such media exhibit memory due freezing-in laws, as does MHD, which in turn constrains the dynamics. We report new results in the theory of CHNS turbulence in 2D, with special emphasis on the role of structure (i.e. `blob') formation and its interaction with the dual cascade. The evolution of a concentration gradient in response to a single eddy -- analogous to flux expulsion in MHD -- is analyzed. Lessons learned are discussed in the context of MHD and other elastic media.

physics.flu-dyn↗

Spontaneous Transport Barriers Quench Turbulent Resistivity in 2D MHD

This Letter identifies the physical mechanism for the quench of turbulent resistivity in 2D MHD. Without an imposed, ordered magnetic field, a multi-scale, blob-and-barrier structure of magnetic potential forms spontaneously. Magnetic energy is concentrated in thin, linear barriers, located at the interstices between blobs. The barriers quench the transport and kinematic decay of magnetic energy. The local transport bifurcation underlying barrier formation is linked to the inverse cascade of $\langle A^2\rangle$ and negative resistivity, which induce local bistability. For small scale forcing, spontaneous layering of the magnetic potential occurs, with barriers located at the interstices between layers. This structure is effectively a magnetic staircase.

physics.flu-dyn↗

Permutation polynomials of degree 8 over finite fields of characteristic 2

Up to linear transformations, we obtain a classification of permutation polynomials (PPs) of degree $8$ over $\mathbb{F}_{2^r}$ with $r>3$. By [J. Number Theory 176 (2017) 466-66], a polynomial $f$ of degree $8$ over $\mathbb{F}_{2^r}$ is exceptional if and only if $f-f(0)$ is a linearized PP. So it suffices to search for non-exceptional PPs of degree $8$ over $\mathbb{F}_{2^r}$, which exist only when $r\leqslant9$ by a previous result. This can be exhausted by the SageMath software running on a personal computer. To facilitate the computation, some requirements after linear transformations and explicit equations by Hermite's criterion are provided for the polynomial coefficients. The main result is that a non-exceptional PP $f$ of degree $8$ over $\mathbb{F}_{2^r}$ (with $r>3$) exists if and only if $r\in\{4,5,6\}$, and such $f$ is explicitly listed up to linear transformations.

math.NT↗

The Weil bound and non-exceptional permutation polynomials over finite fields

A well-known result of von zur Gathen asserts that a non-exceptional permutation polynomial of degree $n$ over $\mathbb{F}_{q}$ exists only if $q<n^{4}$. With the help of the Weil bound for the number of $\mathbb{F}_{q}$-points on an absolutely irreducible (possibly singular) affine plane curve, Chahal and Ghorpade improved von zur Gathen's proof to replace $n^{4}$ by a bound less than $n^{2}(n-2)^{2}$. Also based on the Weil bound, we further refine the upper bound for $q$ with respect to $n$, by a more concise and direct proof following Wan's arguments.

math.NT↗

A minimal $U(1)^\prime$ extension of MSSM in light of the B decay anomaly

Motivated by the $R_K$ and $R_{K^*}$ anomalies from B decays, we extend the minimal supersymmetric model with a non-universal anomaly-free $U(1)^\prime$ gauge symmetry, coupling non-universally to the lepton sector as well as the quark sector. In particular, only the third generation quarks are charged under this $U(1)^\prime$, which can easily evade the dilepton bound from the LHC searches. An extra singlet is introduced to break this $U(1)^\prime$ symmetry allowing for the $μ$-term to be generated dynamically. The relevant constraints of $B_s-\bar{B}_s$ mixing, $D^0-\bar{D}^0$ mixing and the LHC dilepton searches are considered. We find that in the allowed parameter space this $U(1)^\prime$ gauge interaction can accommodate the $R_K$ and $R_{K^*}$ anomalies and weaken considerably the $Z^\prime$ mass limits while remaining perturbative up to the Planck scale.

hep-ph↗

Linear complexity of Ding-Helleseth generalized cyclotomic sequences of order eight

During the last two decades, many kinds of periodic sequences with good pseudo-random properties have been constructed from classical and generalized cyclotomic classes, and used as keystreams for stream ciphers and secure communications. Among them are a family DH-GCS$_{d}$ of generalized cyclotomic sequences on the basis of Ding and Helleseth's generalized cyclotomy, of length $pq$ and order $d=\mathrm{gcd}(p-1,q-1)$ for distinct odd primes $p$ and $q$. The linear complexity (or linear span), as a valuable measure of unpredictability, is precisely determined for DH-GCS$_{8}$ in this paper. Our approach is based on Edemskiy and Antonova's computation method with the help of explicit expressions of Gaussian classical cyclotomic numbers of order $8$. Our result for $d=8$ is compatible with Yan's low bound $(pq-1)/2$ of the linear complexity for any order $d$, which means high enough to resist security attacks of the Berlekamp-Massey algorithm. Finally, we include SageMath codes to illustrate the validity of our result by examples.

math.NT↗