Searcharxiv⌕ Search

arXiv subjects

Fan Chen

Publications and source records attributed to Fan Chen.

At least 73 records · Page 4Linked to original sources

Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy

We study the problem of interactive decision making in which the underlying environment changes over time subject to given constraints. We propose a framework, which we call \textit{hybrid Decision Making with Structured Observations} (hybrid DMSO), that provides an interpolation between the stochastic and adversarial settings of decision making. Within this framework, we can analyze local differentially private (LDP) decision making, query-based learning (in particular, SQ learning), and robust and smooth decision making under the same umbrella, deriving upper and lower bounds based on variants of the Decision-Estimation Coefficient (DEC). We further establish strong connections between the DEC's behavior, the SQ dimension, local minimax complexity, learnability, and joint differential privacy. To showcase the framework's power, we provide new results for contextual bandits under the LDP constraint.

cs.LG↗

Quantum Neural Network Extraction Attack via Split Co-Teaching

Quantum Neural Networks (QNNs), now offered as QNN-as-a-Service (QNNaaS), have become key targets for model extraction attacks. Existing methods use ensemble learning to train substitute QNNs, but our analysis reveals significant limitations in real-world environments, where noise and cost constraints undermine their effectiveness. In this work, we introduce a novel attack, \textit{split co-teaching}, which uses label variations to \textit{split} queried data by noise sensitivity and employs \textit{co-teaching} schemes to enhance extraction accuracy. The experimental results show that our approach outperforms classical extraction attacks by 6.5\%$\sim$9.5\% and existing QNN extraction methods by 0.1\%$\sim$3.7\% across various tasks.

quant-ph↗

LSTM-QGAN: Scalable NISQ Generative Adversarial Network

Current quantum generative adversarial networks (QGANs) still struggle with practical-sized data. First, many QGANs use principal component analysis (PCA) for dimension reduction, which, as our studies reveal, can diminish the QGAN's effectiveness. Second, methods that segment inputs into smaller patches processed by multiple generators face scalability issues. In this work, we propose LSTM-QGAN, a QGAN architecture that eliminates PCA preprocessing and integrates quantum long short-term memory (QLSTM) to ensure scalable performance. Our experiments show that LSTM-QGAN significantly enhances both performance and scalability over state-of-the-art QGAN models, with visual data improvements, reduced Frechet Inception Distance scores, and reductions of 5x in qubit counts, 5x in single-qubit gates, and 12x in two-qubit gates.

quant-ph↗

Deciding Bank Interest Rates -- A Major-Minor Impulse Control Mean-Field Game Perspective

Deciding bank interest rates has been a long-standing challenge in finance. It is crucial to ensure that the selected rates balance market share and profitability. However, traditional approaches typically focus on the interest rate changes of individual banks, often neglecting the interactions with other banks in the market. This work proposes a novel framework that models the interest rate problem as a major-minor mean field game within the context of an interbank game. To incorporate the complex interactions between banks, we utilize mean-field theory and employ impulsive control to model the overhead in rate adjustments. Ultimately, we solve this optimal control problem using a new deep Q-network method, which iterates the parameterized action value functions for major and minor players and updates the networks in a Fictitious Play way. Our proposed algorithm converges, offering a solution that enables the analysis of strategies for major and minor players in the market under the Nash Equilibrium.

math.OC↗

Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond

Modern Reinforcement Learning (RL) is more than just learning the optimal policy; Alternative learning goals such as exploring the environment, estimating the underlying model, and learning from preference feedback are all of practical importance. While provably sample-efficient algorithms for each specific goal have been proposed, these algorithms often depend strongly on the particular learning goal and thus admit different structures correspondingly. It is an urging open question whether these learning goals can rather be tackled by a single unified algorithm. We make progress on this question by developing a unified algorithm framework for a large class of learning goals, building on the Decision-Estimation Coefficient (DEC) framework. Our framework handles many learning goals such as no-regret RL, PAC RL, reward-free learning, model estimation, and preference-based learning, all by simply instantiating the same generic complexity measure called "Generalized DEC", and a corresponding generic algorithm. The generalized DEC also yields a sample complexity lower bound for each specific learning goal. As applications, we propose "decouplable representation" as a natural sufficient condition for bounding generalized DECs, and use it to obtain many new sample-efficient results (and recover existing results) for a wide range of learning goals and problem classes as direct corollaries. Finally, as a connection, we re-analyze two existing optimistic model-based algorithms based on Posterior Sampling and Maximum Likelihood Estimation, showing that they enjoy sample complexity bounds under similar structural conditions as the DEC.

cs.LG↗

Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability

We develop a unifying framework for information-theoretic lower bound in statistical estimation and interactive decision making. Classical lower bound techniques -- such as Fano's method, Le Cam's method, and Assouad's lemma -- are central to the study of minimax risk in statistical estimation, yet are insufficient to provide tight lower bounds for \emph{interactive decision making} algorithms that collect data interactively (e.g., algorithms for bandits and reinforcement learning). Recent work of Foster et al. (2021, 2023) provides minimax lower bounds for interactive decision making using seemingly different analysis techniques from the classical methods. These results -- which are proven using a complexity measure known as the \emph{Decision-Estimation Coefficient} (DEC) -- capture difficulties unique to interactive learning, yet do not recover the tightest known lower bounds for passive estimation. We propose a unified view of these distinct methodologies through a new lower bound approach called \emph{interactive Fano method}. As an application, we introduce a novel complexity measure, the \emph{Fractional Covering Number}, which facilitates the new lower bounds for interactive decision making that extend the DEC methodology by incorporating the complexity of estimation. Using the fractional covering number, we (i) provide a unified characterization of learnability for \emph{any} stochastic bandit problem, (ii) close the remaining gap between the upper and lower bounds in Foster et al. (2021, 2023) (up to polynomial factors) for any interactive decision making problem in which the underlying model class is convex.

cs.LG↗

Brualdi-Hoffman-Turán problem of the gem

A graph is said to be $F$-free if it does not contain $F$ as a subgraph. Brualdi-Hoffman-Turán problem seeks to determine the maximum spectral radius of an $F$-free graph with given size. The gem consists of a path on $4$ vertices, along with an additional vertex that is adjacent to every vertex of the path. Concerning Brualdi-Hoffman-Turán problem of the gem, when the size is odd, Zhang and Wang [Discrete Math. 347 (2024) 114171] and Yu, Li and Peng [arXiv:2404. 03423] solved it. In this paper, we completely solve the Brualdi-Hoffman-Turán problem type problem of the gem.

math.CO↗

LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences

Throughout its lifecycle, a large language model (LLM) generates a substantially larger carbon footprint during inference than training. LLM inference requests vary in batch size, prompt length, and token generation number, while cloud providers employ different GPU types and quantities to meet diverse service-level objectives for accuracy and latency. It is crucial for both users and cloud providers to have a tool that quickly and accurately estimates the carbon impact of LLM inferences based on a combination of inference request and hardware configurations before execution. Estimating the carbon footprint of LLM inferences is more complex than training due to lower and highly variable model FLOPS utilization, rendering previous equation-based models inaccurate. Additionally, existing machine learning (ML) prediction methods either lack accuracy or demand extensive training data, as they inadequately handle the distinct prefill and decode phases, overlook hardware-specific features, and inefficiently sample uncommon inference configurations. We introduce \coo, a graph neural network (GNN)-based model that greatly improves the accuracy of LLM inference carbon footprint predictions compared to previous methods.

cs.LG↗

CausalVE: Face Video Privacy Encryption via Causal Video Prediction

Advanced facial recognition technologies and recommender systems with inadequate privacy technologies and policies for facial interactions increase concerns about bioprivacy violations. With the proliferation of video and live-streaming websites, public-face video distribution and interactions pose greater privacy risks. Existing techniques typically address the risk of sensitive biometric information leakage through various privacy enhancement methods but pose a higher security risk by corrupting the information to be conveyed by the interaction data, or by leaving certain biometric features intact that allow an attacker to infer sensitive biometric information from them. To address these shortcomings, in this paper, we propose a neural network framework, CausalVE. We obtain cover images by adopting a diffusion model to achieve face swapping with face guidance and use the speech sequence features and spatiotemporal sequence features of the secret video for dynamic video inference and prediction to obtain a cover video with the same number of frames as the secret video. In addition, we hide the secret video by using reversible neural networks for video hiding so that the video can also disseminate secret data. Numerous experiments prove that our CausalVE has good security in public video dissemination and outperforms state-of-the-art methods from a qualitative, quantitative, and visual point of view.

cs.CV↗

Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model

The emerging video LMMs (Large Multimodal Models) have achieved significant improvements on generic video understanding in the form of VQA (Visual Question Answering), where the raw videos are captured by cameras. However, a large portion of videos in real-world applications are edited videos, \textit{e.g.}, users usually cut and add effects/modifications to the raw video before publishing it on social media platforms. The edited videos usually have high view counts but they are not covered in existing benchmarks of video LMMs, \textit{i.e.}, ActivityNet-QA, or VideoChatGPT benchmark. In this paper, we leverage the edited videos on a popular short video platform, \textit{i.e.}, TikTok, and build a video VQA benchmark (named EditVid-QA) covering four typical editing categories, i.e., effect, funny, meme, and game. Funny and meme videos benchmark nuanced understanding and high-level reasoning, while effect and game evaluate the understanding capability of artificial design. Most of the open-source video LMMs perform poorly on the EditVid-QA benchmark, indicating a huge domain gap between edited short videos on social media and regular raw videos. To improve the generalization ability of LMMs, we collect a training set for the proposed benchmark based on both Panda-70M/WebVid raw videos and small-scale TikTok/CapCut edited videos, which boosts the performance on the proposed EditVid-QA benchmark, indicating the effectiveness of high-quality training data. We also identified a serious issue in the existing evaluation protocol using the GPT-3.5 judge, namely a "sorry" attack, where a sorry-style naive answer can achieve an extremely high rating from the GPT judge, e.g., over 4.3 for correctness score on VideoChatGPT evaluation protocol. To avoid the "sorry" attacks, we evaluate results with GPT-4 judge and keyword filtering. The dataset is released at https://github.com/XenonLamb/EditVid-QA.

cs.CV↗

A Brualdi-Hoffman-Turán problem for friendship graph

A graph is said to be $H$-free if it does not contain $H$ as a subgraph. Brualdi-Hoffman-Turán type problem is to determine the maximum spectral radius of an $H$-free graph $G$ with give size $m$. The $F_k$ is the graph consisting of $k$ triangles that intersect in exactly one common vertex, which is known as the friendship graph. In this paper, we resolve a conjecture (the Brualdi-Hoffman-Turán-type problem for $F_k$) of Li, Lu and Peng [Discrete Math. 346 (2023) 113680] by using the $k$-core technique presented in Li, Zhai and Shu [European J. Combin, 120 (2024) 103966].

math.CO↗

IoTCO2: Assessing the End-To-End Carbon Footprint of Internet-of-Things-Enabled Deep Learning

To improve privacy and ensure quality-of-service (QoS), deep learning (DL) models are increasingly deployed on Internet of Things (IoT) devices for data processing, significantly increasing the carbon footprint associated with DL on IoT, covering both operational and embodied aspects. Existing operational energy predictors often overlook quantized DL models and emerging neural processing units (NPUs), while embodied carbon footprint modeling tools neglect non-computing hardware components common in IoT devices, creating a gap in accurate carbon footprint modeling tools for IoT-enabled DL. This paper introduces \textit{\carb}, an end-to-end tool for precise carbon footprint estimation in IoT-enabled DL, with deviations as low as 5\% for operational and 3.23\% for embodied carbon footprints compared to actual measurements across various DL models. Additionally, practical applications of \carb~are showcased through multiple user case studies.

cs.LG↗

A novel volume of fluid ghost-cell immersed boundary method for free surface flow interacting with structures

This paper presents a novel volume of fluid ghost-cell immersed boundary (IB) method for two-phase free surface flow interacting with structures. To circumvent the disturbance occurring around the intersection area of the IB and free surface when using the interpolation method for variable reconstruction, the fluid-structure interaction is firstly considered with the orthogonal IB by mimicking the imposition of boundary conditions in the body-conformal grid method. Treatments are subsequently performed to account for the non-orthogonal effect in accurately simulating the FSI, including the newly proposed flux-scaling and IB velocity re-evaluation methods. Further, a variable smoothing process and a flux correction method are adapted to handle moving boundary cases. Based on OpenFOAM, a two-phase flow solver has been developed. Both stationary and moving immersed boundary cases are used for validations. The numerical results reasonably agree with the corresponding laboratory data and other numerical simulation results, demonstrating the disturbance being effectively depressed and the solver's accuracy in capturing fluid-structure interactions involving free surface flow.

physics.flu-dyn↗

AIPO: Improving Training Objective for Iterative Preference Optimization

Preference Optimization (PO), is gaining popularity as an alternative choice of Proximal Policy Optimization (PPO) for aligning Large Language Models (LLMs). Recent research on aligning LLMs iteratively with synthetic or partially synthetic data shows promising results in scaling up PO training for both academic settings and proprietary trained models such as Llama3. Despite its success, our study shows that the length exploitation issue present in PO is even more severe in Iterative Preference Optimization (IPO) due to the iterative nature of the process. In this work, we study iterative preference optimization with synthetic data. We share the findings and analysis along the way of building the iterative preference optimization pipeline. More specifically, we discuss the length exploitation issue during iterative preference optimization and propose our training objective for iterative preference optimization, namely Agreement-aware Iterative Preference Optimization (AIPO). To demonstrate the effectiveness of our method, we conduct comprehensive experiments and achieve state-of-the-art performance on MT-Bench, AlpacaEval 2.0, and Arena-Hard. Our implementation and model checkpoints will be made available at https://github.com/bytedance/AIPO.

cs.CL↗

The maximum index of signed complete graphs whose negative edges induce a bicyclic graph

Let $Γ=(K_n,H)$ be a signed complete graph whose negative edges induce a subgraph $H$. Let $A(Γ)$ be the adjacency matrix of the signed graph $Γ$. The largest eigenvalue of $A(Γ)$ is called the index of $Γ$. In this paper, the index of all the signed complete graphs whose negative edges induce a bicyclic graph $B$ is investigated. Specifically, the structure of the bicyclic graph $B$ such that $Γ=(K_n,B)$ has the maximum index is determined.

math.CO↗

Near-Optimal Learning and Planning in Separated Latent MDPs

We study computational and statistical aspects of learning Latent Markov Decision Processes (LMDPs). In this model, the learner interacts with an MDP drawn at the beginning of each epoch from an unknown mixture of MDPs. To sidestep known impossibility results, we consider several notions of separation of the constituent MDPs. The main thrust of this paper is in establishing a nearly-sharp *statistical threshold* for the horizon length necessary for efficient learning. On the computational side, we show that under a weaker assumption of separability under the optimal policy, there is a quasi-polynomial algorithm with time complexity scaling in terms of the statistical threshold. We further show a near-matching time complexity lower bound under the exponential time hypothesis.

cs.LG↗

Coexistence of ferromagnetism and superconductivity at KTaO$_3$ heterointerfaces

The coexistence of superconductivity and ferromagnetism is a long-standing issue in superconductivity due to the antagonistic nature of these two ordered states. Experimentally identifying and characterizing novel heterointerface superconductors that coexist with magnetism presents significant challenges. Here, we report the experimental observation of two-dimensional long-range ferromagnetic order in the KTaO$_3$ heterointerface superconductor, showing the coexistence of superconductivity and ferromagnetism. Remarkably, our direct current superconducting quantum interference device measurements reveal an in-plane magnetization hysteresis loop persisting above room temperature. Moreover, the first-principles calculations and X-ray magnetic circular dichroism measurements provide decisive insights into the origin of the observed robust ferromagnetism, attributing it to oxygen vacancies that localize electrons in nearby Ta 5$d$ states. Our findings not only suggest KTaO$_3$ heterointerfaces as time-reversal symmetry breaking superconductors, but also inject fresh momentum into the exploration of the intricate interplay between superconductivity and magnetism, enhanced by the strong spin-orbit coupling inherent to the heavy Ta in 5$d$ orbitals of KTaO$_3$ heterointerfaces.

cond-mat.supr-con↗

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of improving model capabilities from the vision side. Inspired by the successful applications of Mixture-of-Experts (MoE) in LLMs, which improves model scalability during training while keeping inference costs similar to those of smaller models, we propose CuMo. CuMo incorporates Co-upcycled Top-K sparsely-gated Mixture-of-experts blocks into both the vision encoder and the MLP connector, thereby enhancing the multimodal LLMs with minimal additional activated parameters during inference. CuMo first pre-trains the MLP blocks and then initializes each expert in the MoE block from the pre-trained MLP block during the visual instruction tuning stage. Auxiliary losses are used to ensure a balanced loading of experts. CuMo outperforms state-of-the-art multimodal LLMs across various VQA and visual-instruction-following benchmarks using models within each model size group, all while training exclusively on open-sourced datasets. The code and model weights for CuMo are open-sourced at https://github.com/SHI-Labs/CuMo.

cs.CV↗