Searcharxiv⌕ Search

arXiv subjects

Yichen Wang

Publications and source records attributed to Yichen Wang.

At least 55 records · Page 3Linked to original sources

The maximum number of triangles in graphs without the square of a path

The generalized Turán number for $H$ of $G$, denoted by $\ex(n,H,G)$, is the maximum number of copies of $H$ in an $n$-vertex $G$-free graph. When $H$ is an edge, $\ex(n,H,G)$ is the classical Turán number $\ex(n,G)$. Let $P_k$ be the path with $k$ vertices. The square of $P_k$, denoted by $P_k^2$, is obtained by joining the pairs of vertices with distance at most two in $P_k$. The Turán number of $P_k^2$, $\ex(n, P_k^2)$, was determined by several researchers. When $k=3$, $P_3^2$ is the triangle and $\ex(n, P_3^2)$ is well-known from Mantel's theorem. When $k=4$, $\ex(n, P_4^2)$ was solved by Dirac in a more general context. When $k=5,6$, the problem was solved by Xiao, Katona, Xiao, and Zamora. For general $k \ge 7$, the problem was solved by Yuan in a more general context. Recently, Mukherjee determined the generalized Turán number $\ex(n, K_3, P_5^2)$. In this paper, we determine the exact value of $\ex(n, K_3, P_6^2)$ and characterize all the extremal graphs for $n \ge 11$.

math.CO↗

Towards Efficient 3D Object Detection for Vehicle-Infrastructure Collaboration via Risk-Intent Selection

Vehicle-Infrastructure Collaborative Perception (VICP) is pivotal for resolving occlusion in autonomous driving, yet the trade-off between communication bandwidth and feature redundancy remains a critical bottleneck. While intermediate fusion mitigates data volume compared to raw sharing, existing frameworks typically rely on spatial compression or static confidence maps, which inefficiently transmit spatially redundant features from non-critical background regions. To address this, we propose Risk-intent Selective detection (RiSe), an interaction-aware framework that shifts the paradigm from identifying visible regions to prioritizing risk-critical ones. Specifically, we introduce a Potential Field-Trajectory Correlation Model (PTCM) grounded in potential field theory to quantitatively assess kinematic risks. Complementing this, an Intention-Driven Area Prediction Module (IDAPM) leverages ego-motion priors to proactively predict and filter key Bird's-Eye-View (BEV) areas essential for decision-making. By integrating these components, RiSe implements a semantic-selective fusion scheme that transmits high-fidelity features only from high-interaction regions, effectively acting as a feature denoiser. Extensive experiments on the DeepAccident dataset demonstrate that our method reduces communication volume to 0.71\% of full feature sharing while maintaining state-of-the-art detection accuracy, establishing a competitive Pareto frontier between bandwidth efficiency and perception performance.

cs.CV↗

Early GVHD Prediction in Liver Transplantation via Multi-Modal Deep Learning on Imbalanced EHR Data

Graft-versus-host disease (GVHD) is a rare but often fatal complication in liver transplantation, with a very high mortality rate. By harnessing multi-modal deep learning methods to integrate heterogeneous and imbalanced electronic health records (EHR), we aim to advance early prediction of GVHD, paving the way for timely intervention and improved patient outcomes. In this study, we analyzed pre-transplant electronic health records (EHR) spanning the period before surgery for 2,100 liver transplantation patients, including 42 cases of graft-versus-host disease (GVHD), from a cohort treated at Mayo Clinic between 1992 and 2025. The dataset comprised four major modalities: patient demographics, laboratory tests, diagnoses, and medications. We developed a multi-modal deep learning framework that dynamically fuses these modalities, handles irregular records with missing values, and addresses extreme class imbalance through AUC-based optimization. The developed framework outperforms all single-modal and multi-modal machine learning baselines, achieving an AUC of 0.836, an AUPRC of 0.157, a recall of 0.768, and a specificity of 0.803. It also demonstrates the effectiveness of our approach in capturing complementary information from different modalities, leading to improved performance. Our multi-modal deep learning framework substantially improves existing approaches for early GVHD prediction. By effectively addressing the challenges of heterogeneity and extreme class imbalance in real-world EHR, it achieves accurate early prediction. Our proposed multi-modal deep learning method demonstrates promising results for early prediction of a GVHD in liver transplantation, despite the challenge of extremely imbalanced EHR data.

cs.LG↗

Multiscale spatiotemporal heterogeneity analysis of bike-sharing system's self-loop phenomenon: Evidence from Shanghai

Bike-sharing is an environmentally friendly shared mobility mode, but its self-loop phenomenon, where bikes are returned to the same station after several time usage, significantly impacts equity in accessing its services. Therefore, this study conducts a multiscale analysis with a spatial autoregressive model and double machine learning framework to assess socioeconomic features and geospatial location's impact on the self-loop phenomenon at metro stations and street scales. The results reveal that bike-sharing self-loop intensity exhibits significant spatial lag effect at street scale and is positively associated with residential land use. Marginal treatment effects of residential land use is higher on streets with middle-aged residents, high fixed employment, and low car ownership. The multimodal public transit condition reveals significant positive marginal treatment effects at both scales. To enhance bike-sharing cooperation, we advocate augmenting bicycle availability in areas with high metro usage and low bus coverage, alongside implementing adaptable redistribution strategies.

cs.LG↗

Linear recoloring diameter of degenerate chordal graphs and bounded treewidth graphs

Let $G$ be a graph on $n$ vertices and $t$ an integer. The reconfiguration graph of $G$, denoted by $R_t(G)$, consists of all $t$-colorings of $G$ and two $t$-colorings are adjacent if they differ on exactly one vertex. The $t$-recoloring diameter of $G$ is the diameter of $R_t(G)$. For a $d$-degenerate graph $G$, $R_t(G)$ is connected when $t \ge d+2$~(Dyer et al., 2006). Furthermore, the $t$-recoloring diameter is $O(n^2)$ when $t \ge 3(d+1)/2$~(Bousquet et al., 2022), and it is $O(n)$ when $t \ge 2d+2$~(Bousquet and Perarnau, 2016). For a $d$-degenerate and chordal graph $G$, the $t$-recoloring diameter of $G$ is $O(n^2)$ when $t \ge d+2$~(Bonamy et al. 2014). If $G$ is a graph of treewidth at most $k$, then $G$ is also $k$-degenerate, and the previous results hold. Moreover, when $t \ge k+2$, the $t$-recoloring diameter is $O(n^2)$~(Bonamy and Bousquet, 2013). When $k=2$, the $t$-recoloring diameter of $G$ is linear when $t \ge 5$~(Bartier, Bousquet and Heinrich, 2021) and the result is tight. In this paper, we prove that if $G$ is $d$-degenerate and chordal, then the $t$-recoloring diameter of $G$ is $O(n)$ when $t \ge 2d+1$. Moreover, if the treewidth of $G$ is at most $k$, then the $t$-recoloring diameter is $O(n)$ when $t \ge 2k+1$. This result is a generalization of the previous results on graphs of treewidth at most two.

math.CO↗

The $φ$ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control

Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learning curves observed for large over-parametrized deep networks. Capacity measures based on parameter count typically fail to account for these empirical observations. To tackle this challenge, we consider norm-based capacity measures and develop our study for random features based estimators, widely used as simplified theoretical models for more complex networks. In this context, we provide a precise characterization of how the estimator's norm concentrates and how it governs the associated test error. Our results show that the predicted learning curve admits a phase transition from under- to over-parameterization, but no double descent behavior. This confirms that more classical U-shaped behavior is recovered considering appropriate capacity measures based on models norms rather than size. From a technical point of view, we leverage deterministic equivalence as the key tool and further develop new deterministic quantities which are of independent interest.

stat.ML↗

Edge version of the inducibility via the entropy method

The inducibility of a graph $H$ is about the maximum number of induced copies of $H$ in a graph on $n$ vertices. We consider its edge version, that is, the maximum number of induced copies of $H$ in a graph with $m$ edges. Let $c(G,H)$ be the number of induced copies of $H$ in $G$ and $ρ(H,m) = \max \{c(G,H) \mid |E(G)| = m\}$. For any graph $H$, we prove that $ρ(H,m) = Θ(m^{α_f(H)})$ where $α_f(H)$ is the fractional independence number of $H$. Therefore, we now focus on the constant factor in front of $m^{α_f(H)}$. In this paper, we give some results of $ρ(H,m)$ when $H$ is a cycle or path. We conjecture that for any cycle $C_k$ with $k \ge 5$, $ρ(C_k,m)= (1+o(1))\left( m/k\right)^{k/2}$ and the bound achieves by the blow up of $C_k$. For even cycles, we establish an upper bound with an extra constant factor. For odd cycles, we can only establish an upper bound with an extra factor depending on $k$. We prove that $ρ(P_{2l},m) \le \frac{m^l}{2(l-1)^{l-1}}$ and $ρ(P_{2l+1},m) \le \frac{m^{l+1}}{4l^l}$, where $l \ge 2$. We also conjecture the asymptotic value of $ρ(P_k, m)$. The entropy method is mainly used to prove our results.

math.CO↗

Unraveling Misinformation Propagation in LLM Reasoning

Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning, positioning them as promising tools for supporting human problem-solving. However, what happens when their performance is affected by misinformation, i.e., incorrect inputs introduced by users due to oversights or gaps in knowledge? Such misinformation is prevalent in real-world interactions with LLMs, yet how it propagates within LLMs' reasoning process remains underexplored. Focusing on mathematical reasoning, we present a comprehensive analysis of how misinformation affects intermediate reasoning steps and final answers. We also examine how effectively LLMs can correct misinformation when explicitly instructed to do so. Even with explicit instructions, LLMs succeed less than half the time in rectifying misinformation, despite possessing correct internal knowledge, leading to significant accuracy drops (10.02% - 72.20%), and the degradation holds with thinking models (4.30% - 19.97%). Further analysis shows that applying factual corrections early in the reasoning process most effectively reduces misinformation propagation, and fine-tuning on synthesized data with early-stage corrections significantly improves reasoning factuality. Our work offers a practical approach to mitigating misinformation propagation.

cs.CL↗

Treewidth of generalized Hamming graph, bipartite Kneser graph and generalized Petersen graph

Let $t,q$ and $n$ be positive integers. Write $[q] = \{1,2,\ldots,q\}$. The generalized Hamming graph $H(t,q,n)$ is the graph whose vertex set is the cartesian product of $n$ copies of $[q]$ ($q\ge 2$), where two vertices are adjacent if their Hamming distance is at most $t$. In particular, $H(1,q,n)$ is the well-known Hamming graph and $H(1,2,n)$ is the hypercube. In 2006, Chandran and Kavitha described the asymptotic value of $tw(H(1,q,n))$, where $tw(G)$ denotes the treewidth of $G$. In this paper, we give the exact pathwidth of $H(t,2,n)$ and show that $tw(H(t,q,n)) = Θ(tq^n/\sqrt{n})$ when $n$ goes to infinity. Based on those results, we show that the treewidth of the bipartite Kneser graph $BK(n,k)$ is $\binom{n}{k} - 1$ when $n$ is sufficiently large relative to $k$ and the bounds of $tw(BK(2k+1,k))$ are given. Moreover, we present the bounds of the treewidth of the generalized Petersen graph.

math.CO↗

ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents

Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their vulnerabilities. However, these attacks either rely on overly strong assumptions, requiring full knowledge of the victim VLM, which is impractical for attacking VLM-based agents, or exhibit limited effectiveness. The latter stems from disrupting most semantic information in the image, which leads to a misalignment between the perception and the task context defined by system prompts. This inconsistency interrupts the VLM's reasoning process, resulting in invalid outputs that fail to affect interactions in the physical world. To this end, we propose a fine-grained adversarial attack framework, ADVEDM, which modifies the VLM's perception of only a few key objects while preserving the semantics of the remaining regions. This attack effectively reduces conflicts with the task context, making VLMs output valid but incorrect decisions and affecting the actions of agents, thus posing a more substantial safety threat in the physical world. We design two variants of based on this framework, ADVEDM-R and ADVEDM-A, which respectively remove the semantics of a specific object from the image and add the semantics of a new object into the image. The experimental results in both general scenarios and EDM tasks demonstrate fine-grained control and excellent attack performance.

cs.CV↗

Counting induced subgraphs with given intersection sizes

Let $F$ be a graph of order $r$. In this paper, we study the maximum number of induced copies of $F$ with restricted intersections, which highlights the motivation from extremal set theory. Let $L=\{\ell_1,\dots,\ell_s\}\subseteq[0,r-1]$ be an integer set with $s\not\in\{1,r\}$. Let $Ψ_r(n,F,L)$ be the maximum number of induced copies of $F$ in an $n$-vertex graph, where the induced copies of $F$ are $L$-intersecting as a family of $r$-subsets, i.e., for any two induced copies of $F$, the size of their intersection is in $L$. Helliar and Liu initiated a study of the function $Ψ_r(n,K_r,L)$. Very recently, Zhao and Zhang improved their result and showed that $Ψ_r(n,K_r,L)=Θ_{r,L}(n^{s})$ if and only if $\ell_1,\dots,\ell_s,r$ form an arithmetic progression. In this paper, we show that $Ψ_r(n,F,L)=o_{r,L}(n^{s})$ when $\ell_1,\dots,\ell_s,r$ do not form an arithmetic progression. We study the asymptotical result of $Ψ_r(n,C_r,L)$, and determined the asymptotically optimal result when $\ell_1,\dots,\ell_s,r$ form an arithmetic progression and take certain values. We also study the generalized Turán problem, determining the maximum number of $H$, where the copies of $H$ are $L$-intersecting as a family of $r$-subsets. The entropy method is used to prove our results.

math.CO↗

Agentic Satellite-Augmented Low-Altitude Economy and Terrestrial Networks: A Survey on Generative Approaches

The development of satellite-augmented low-altitude economy and terrestrial networks (SLAETNs) demands intelligent and autonomous systems that can operate reliably across heterogeneous, dynamic, and mission-critical environments. To address these challenges, this survey focuses on enabling agentic artificial intelligence (AI), that is, artificial agents capable of perceiving, reasoning, and acting, through generative AI (GAI) and large language models (LLMs). We begin by introducing the architecture and characteristics of SLAETNs, and analyzing the challenges that arise in integrating satellite, aerial, and terrestrial components. Then, we present a model-driven foundation by systematically reviewing five major categories of generative models: variational autoencoders (VAEs), generative adversarial networks (GANs), generative diffusion models (GDMs), transformer-based models (TBMs), and LLMs. Moreover, we provide a comparative analysis to highlight their generative mechanisms, capabilities, and deployment trade-offs within SLAETNs. Building on this foundation, we examine how these models empower agentic functions across three domains: communication enhancement, security and privacy protection, and intelligent satellite tasks. Finally, we outline key future directions for building scalable, adaptive, and trustworthy generative agents in SLAETNs. This survey aims to provide a unified understanding and actionable reference for advancing agentic AI in next-generation integrated networks.

cs.NI↗

DPNO: A Dual Path Architecture For Neural Operator

Neural operators have emerged as a powerful tool for solving partial differential equations (PDEs) and other complex scientific computing tasks. However, the performance of single operator block is often limited, thus often requiring composition of basic operator blocks to achieve better per-formance. The traditional way of composition is staking those blocks like feedforward neural networks, which may not be very economic considering parameter-efficiency tradeoff. In this pa-per, we propose a novel dual path architecture that significantly enhances the capabilities of basic neural operators. The basic operator block is organized in parallel two paths which are similar with ResNet and DenseNet. By introducing this parallel processing mechanism, our architecture shows a more powerful feature extraction and solution approximation ability compared with the original model. We demonstrate the effectiveness of our approach through extensive numerical experi-ments on a variety of PDE problems, including the Burgers' equation, Darcy Flow Equation and the 2d Navier-Stokes equation. The experimental results indicate that on certain standard test cas-es, our model achieves a relative improvement of over 30% compared to the basic model. We also apply this structure on two standard neural operators (DeepONet and FNO) selected from different paradigms, which suggests that the proposed architecture has excellent versatility and offering a promising direction for neural operator structure design.

math.NA↗

Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy

Can we simulate a sandbox society with generative agents to model human behavior, thereby reducing the over-reliance on real human trials for assessing public policies? In this work, we investigate the feasibility of simulating health-related decision-making, using vaccine hesitancy, defined as the delay in acceptance or refusal of vaccines despite the availability of vaccination services (MacDonald, 2015), as a case study. To this end, we introduce the VacSim framework with 100 generative agents powered by Large Language Models (LLMs). VacSim simulates vaccine policy outcomes with the following steps: 1) instantiate a population of agents with demographics based on census data; 2) connect the agents via a social network and model vaccine attitudes as a function of social dynamics and disease-related information; 3) design and evaluate various public health interventions aimed at mitigating vaccine hesitancy. To align with real-world results, we also introduce simulation warmup and attitude modulation to adjust agents' attitudes. We propose a series of evaluations to assess the reliability of various LLM simulations. Experiments indicate that models like Llama and Qwen can simulate aspects of human behavior but also highlight real-world alignment challenges, such as inconsistent responses with demographic profiles. This early exploration of LLM-driven simulations is not meant to serve as definitive policy guidance; instead, it serves as a call for action to examine social simulation for policy development.

cs.MA↗

Optimal Differentially Private Ranking from Pairwise Comparisons

Data privacy is a central concern in many applications involving ranking from incomplete and noisy pairwise comparisons, such as recommendation systems, educational assessments, and opinion surveys on sensitive topics. In this work, we propose differentially private algorithms for ranking based on pairwise comparisons. Specifically, we develop and analyze ranking methods under two privacy notions: edge differential privacy, which protects the confidentiality of individual comparison outcomes, and individual differential privacy, which safeguards potentially many comparisons contributed by a single individual. Our algorithms--including a perturbed maximum likelihood estimator and a noisy count-based method--are shown to achieve minimax optimal rates of convergence under the respective privacy constraints. We further demonstrate the practical effectiveness of our methods through experiments on both simulated and real-world data.

math.ST↗

Feature Learning beyond the Lazy-Rich Dichotomy: Insights from Representational Geometry

Integrating task-relevant information into neural representations is a fundamental ability of both biological and artificial intelligence systems. Recent theories have categorized learning into two regimes: the rich regime, where neural networks actively learn task-relevant features, and the lazy regime, where networks behave like random feature models. Yet this simple lazy-rich dichotomy overlooks a diverse underlying taxonomy of feature learning, shaped by differences in learning algorithms, network architectures, and data properties. To address this gap, we introduce an analysis framework to study feature learning via the geometry of neural representations. Rather than inspecting individual learned features, we characterize how task-relevant representational manifolds evolve throughout the learning process. We show, in both theoretical and empirical settings, that as networks learn features, task-relevant manifolds untangle, with changes in manifold geometry revealing distinct learning stages and strategies beyond the lazy-rich dichotomy. This framework provides novel insights into feature learning across neuroscience and machine learning, shedding light on structural inductive biases in neural circuits and the mechanisms underlying out-of-distribution generalization.

cs.LG↗

UAV-Assisted Integrated Communication and Over-the-Air Computation with Interference Awareness

Over the air computation (AirComp) is a promising technique that addresses big data collection and fast wireless data aggregation. However, in a network where wireless communication and AirComp coexist, mutual interference becomes a critical challenge. In this paper, we propose to employ an unmanned aerial vehicle (UAV) to enable integrated communication and AirComp, where we capitalize on UAV mobility with alleviated interference for performance enhancement. Particularly, we aim to maximize the sum of user transmission rate with the guaranteed AirComp accuracy requirement, where we jointly optimize the transmission strategy, signal normalizing factor, scheduling strategy, and UAV trajectory. We decouple the formulated problem into two layers where the outer layer is for UAV trajectory and scheduling, and the inner layer is for transmission and computation. Then, we solve the inner layer problem through alternating optimization, and the outer layer is solved through soft actor critic based deep reinforcement learning. Simulation results show the convergence of the proposed learning process and also demonstrate the performance superiority of our proposal as compared with the baselines in various situations.

eess.SP↗

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

With the significant advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly. Previous studies have highlighted VLMs' vulnerability to jailbreak attacks, where carefully crafted inputs can lead the model to produce content that violates ethical and legal standards. However, existing methods struggle against state-of-the-art VLMs like GPT-4o, due to the over-exposure of harmful content and lack of stealthy malicious guidance. In this work, we propose a novel jailbreak attack framework: Multi-Modal Linkage (MML) Attack. Drawing inspiration from cryptography, MML utilizes an encryption-decryption process across text and image modalities to mitigate over-exposure of malicious information. To align the model's output with malicious intent covertly, MML employs a technique called "evil alignment", framing the attack within a video game production scenario. Comprehensive experiments demonstrate MML's effectiveness. Specifically, MML jailbreaks GPT-4o with attack success rates of 97.80% on SafeBench, 98.81% on MM-SafeBench and 99.07% on HADES-Dataset. Our code is available at https://github.com/wangyu-ovo/MML.

cs.CV↗