SearcharxivSearch

arXiv subjects

Yannan Chen

Publications and source records attributed to Yannan Chen.

At least 19 recordsLinked to original sources

Recursive-Line Zarankiewicz Numbers with Four Columns

The recursive-line Zarankiewicz number maximizes the number of squares in a structured irreducible sum-of-squares representation encoded by an augmentation of an extremal $C_4$-free bipartite graph. We determine its four-column behavior under the strengthened recursive definition in the manuscript of L\"ofberg and Qi dated 9 September 2026. Combining AI-assisted discovery with exact certificate verification and finite exclusion computations, we determine eighteen of the nineteen values for $2\le m\le20$ and isolate the only unresolved case to $37\le\zr(14,4)\le38$. More significantly, we prove the first eventual exact formula in the four-column setting: \[ \zr(m,4)=\floor{\frac{5m+6}{2}}\qquad(m\ge15). \] The upper bound follows from the classical identity $z(m,4)=m+6$ and a sharp cell count. For the matching lower bound, we construct a two-row extension chain from an explicit $20\times4$ seed and derive the odd orders by a fixed deletion. Analytic propagation, together with two independently audited symbolic certificate tables, proves the construction for arbitrary chain length rather than merely for a finite computational range. Thus every extremal configuration has no holes when $m$ is even and exactly one hole when $m$ is odd, and the same exact formula holds for the second-order number $z_2(m,4)$.

math.CO

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, making it difficult for MLLMs to align their responses with the reasoning path required by the given multimodal input. This limitation becomes more pronounced in complex multimodal tasks, thereby restricting further improvements in MLLM performance. To address this issue, we propose a new multimodal ICL framework that combines contrastive demonstration modeling with the self-refinement capability of MLLMs. Specifically, our framework reformulates each demonstration by explicitly contrasting a suboptimal response with a better response under the same input, together with a reasoning path that reveals how the response should be refined. This contrastive formulation makes the reasoning path toward the desired response more explicit and guides the MLLM beyond superficial imitation. Furthermore, because effective refinement depends on the current response, we introduce a response-conditioned retrieval mechanism to select demonstrations whose reasoning paths are more relevant to the current response. In addition, we use a lightweight alignment controller to predict response quality and determine whether further refinement is needed. Experiments on three types of multimodal tasks show that the proposed framework consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).

cs.AI

Exact and Asymptotic Values for Weak Limited Augmented Zarankiewicz Numbers in the $m\times 3$ Case

We determine the exact weak limited augmented Zarankiewicz numbers $z_{wL}(m,3)$ for all $m\ge 3$: \[ z_{wL}(m,3)= \begin{cases} m+3+\left\lceil \dfrac{m}{2}\right\rceil+1, & 9\le m\le 15,\\[2mm] m+3+\left\lfloor \dfrac{2m-4}{3}\right\rfloor, & m\ge 16, \end{cases} \] with $z_{wL}(3,3)=6$, $z_{wL}(4,3)=8$, and $z_{wL}(m,3)=2m$ for $5\le m\le 9$. In particular, \[ \lim_{m\to\infty} \frac{z_{wL}(m,3)}{m} = \frac{5}{3}. \] The proof is fully analytic, relying on a uniform base classification, two constructive lower-bound families (staircase and $5m/3$), and a sharp upper-bound argument based on a peeling lemma and the analysis of two W2-sensitive boundary cases. Numerical MILP computations were used only as proof-mining tools to identify the structural lemmas; the final theorem is unconditional. We also extend the known range of the original limited numbers $z_L(m,3)$ through $m=13$, where the gap to $z_{wL}(m,3)$ is only 2 or 3.

math.OC

Joint Low-Dimensional Modeling and Sampling Design for Sparse On-Orbit Antenna Pattern Reconstruction

Accurately reconstructing satellite transmit-antenna patterns on orbit is difficult because only sparse directional measurements are available during normal mission operations. This paper develops a cooperative on-orbit pattern-reconstruction framework that converts received calibration power into normalized directional samples and represents the antenna power pattern using a truncated discrete cosine transform (DCT) basis. The resulting low-dimensional model transforms high-dimensional pattern recovery into coefficient estimation, for which a closed-form maximum-likelihood estimator and error characterization are derived. The analysis shows how DCT truncation error, measurement noise, and sampled-basis conditioning jointly affect reconstruction accuracy. For regularly accessible angular sectors, midpoint-uniform sampling provides an information-balanced baseline for the retained DCT modes. For constrained feasible opportunities, D-optimal sampling is used to select informative measurement directions. Simulations verify the accuracy of the angular discretization, the sample efficiency of the truncated-DCT model, and the reconstruction gain of D-optimal sampling under irregular orbit-generated opportunities.

eess.SP

Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization

Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabeled target data by training with target pseudo-labels. CST introduces a target classifier and employs an inner-outer loop updating strategy, addressing the issue of unreliable pseudo-labels and enabling pseudo-labels to generalize across domains. Despite its success in image classification, extending CST to object detection faces three main challenges. First, the upper bound of CST in object detection is constrained by three types of unreliable pseudo-labels, such as classification error alone, localization error alone, and their combination. Second, since object detection involves detecting multiple target objects, directly applying CST leads to training insta bility. Third, a wider numerical range of regression coordinates leads to exploding losses. To this end, we apply CST to both classification and regression and propose the Dual-Stream Bilevel-Cycle Optimization framework. Specifically, we construct CST upon Mean Teacher to prevent training instability and use extra normalization to map the regression bounding box into a standardized space, effectively addressing exploding losses. Also, we provide a theoretical derivation of the regression bound. Extensive experiments across four cross domain standard scenarios demonstrate that our framework achieves considerable results.

cs.CV

The SOS Rank of a $5 \times 4$ Biquadratic Form via Orthogonality

Biquadratic forms arise naturally in polynomial optimization, tensor analysis, and quantum information theory. A key problem is determining the minimal number of squares needed in a sum-of-squares (SOS) representation of such a form, known as its SOS rank. For fixed dimensions $(m,n)$, the maximum possible SOS rank over all biquadratic forms in $m$ and $n$ variables is denoted $\operatorname{BSR}(m,n)$. Recent advances have established lower bounds on $\operatorname{BSR}(m,n)$ via combinatorial constructions involving bipartite graphs and the orthogonality method. In particular, for the case $(m,n)=(5,4)$, it was shown that $\operatorname{BSR}(5,4)\ge 11$ using only nondegenerate 2-edges. In this paper, we extend this framework by incorporating a degenerate $2$-edge, which introduces a cross term where the two $y$-indices coincide. We construct an explicit $5\times 4$ biquadratic form and apply the orthogonality method to prove that its SOS rank is $12$, thereby improving the lower bound to $\operatorname{BSR}(5,4)\ge12$. This result demonstrates that degenerate $2$-edges yield additional algebraic flexibility beyond purely combinatorial bounds and extends the applicability of the orthogonality method to forms with cross terms involving identical $y$-indices.

math.CO

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when these cues are missing. Faithful attribution methods can reveal such problematic reliance through localized explanations, but they are typically used post hoc and are not fed back into the model. To address this limitation, we propose Subset-Selected Counterfactual Augmentation (SS-CA), a training strategy that masks decision-relevant regions to construct counterfactual samples and guide the model toward more robust decision boundaries. Specifically, we extend LIMA, a subset-selection-based faithful attribution method, to Counterfactual LIMA to identify regions whose removal shifts the model toward a competing class. SS-CA then selects near-boundary masks that reduce the logit gap while preserving the original semantics, and applies an adaptive counterfactual filling strategy to replace the masked regions without introducing external semantics. Feeding these counterfactual samples back into training encourages the model to exploit the remaining informative evidence and shifts the decision boundary toward a more robust one. Extensive experiments across five ImageNet variants show that SS-CA effectively improves ID accuracy, OOD generalization, and perturbation robustness, achieving gains of 5.70%/18.04% on ImageNet-1k/ImageNet-R with CLIP ViT/32b, 9.52%/11.33% on ImageNet-R/ImageNet-S on TinyImageNet-200 with ResNet-101, and about 4% under Gaussian Noise corruption. The code will be released soon.

cs.CV

Fast Fractional Programming for Multi-Cell Integrated Sensing and Communications

This paper concerns the coordinate multi-cell beamforming design for integrated sensing and communications (ISAC). In particular, we assume that each base station (BS) has massive antennas. The optimization objective is to maximize a weighted sum of the data rates (for communications) and the Fisher information (for sensing). We first show that the conventional beamforming method for the multiple-input multiple-output (MIMO) transmission, i.e., the weighted minimum mean square error (WMMSE) algorithm, works for the ISAC problem case from a fractional programming (FP) perspective. However, the WMMSE algorithm frequently requires computing the $N\times N$ matrix inverse, where $N$ is the number of transmit or receive antennas, so the algorithm becomes quite costly when antennas are massively deployed. To address this issue, we develop a nonhomogeneous bound and use it in conjunction with the FP technique to solve the ISAC beamforming problem without the need to invert any large matrices. It is further shown that the resulting new FP algorithm has an intimate connection with gradient projection, based on which we can accelerate the convergence via Nesterov's gradient extrapolation.

cs.IT

Even Order Pascal Tensors are Positive Definite

In this paper, we show that even order Pascal tensors are positive definite, and odd order Pascal tensors are strongly completely positive. The significance of these is that our induction proof method also holds for some other families of completely positives tensors, whose construction satisfies certain rules, such an inherence property holds. We show that for all tensors in such a family, even order tensors would be positive definite, and odd order tensors would be strongly completely positive, as long as the matrices in this family are positive definite. In particular, we show that even order generalized Pascal tensors would be positive definite, and odd order generalized Pascal tensors would be strongly completely positive, as long as generalized Pascal matrices are positive definite. We also investigate even order positive definiteness and odd order strongly completely positivity for fractional Hadamard power tensors. Furthermore, we study determinants of Pascal tensors. We prove that the determinant of the $m$th order two dimensional symmetric Pascal tensor is equal to the $m$th power of the factorial of $m-1$.

math.RA

Accelerating Quadratic Transform and WMMSE

Fractional programming (FP) arises in various communications and signal processing problems because several key quantities in the field are fractionally structured, e.g., the Cramér-Rao bound, the Fisher information, and the signal-to-interference-plus-noise ratio (SINR). A recently proposed method called the quadratic transform has been applied to the FP problems extensively. The main contributions of the present paper are two-fold. First, we investigate how fast the quadratic transform converges. To the best of our knowledge, this is the first work that analyzes the convergence rate for the quadratic transform as well as its special case the weighted minimum mean square error (WMMSE) algorithm. Second, we accelerate the existing quadratic transform via a novel use of Nesterov's extrapolation scheme [1]. Specifically, by generalizing the minorization-maximization (MM) approach in [2], we establish a nontrivial connection between the quadratic transform and the gradient projection, thereby further incorporating the gradient extrapolation into the quadratic transform to make it converge more rapidly. Moreover, the paper showcases the practical use of the accelerated quadratic transform with two frontier wireless applications: integrated sensing and communications (ISAC) and massive multiple-input multiple-output (MIMO).

cs.IT

Mixed Max-and-Min Fractional Programming for Wireless Networks

Fractional programming (FP) plays a crucial role in wireless network design because many relevant problems involve maximizing or minimizing ratio terms. Notice that the maximization case and the minimization case of FP cannot be converted to each other in general, so they have to be dealt with separately in most of the previous studies. Thus, an existing FP method for maximizing ratios typically does not work for the minimization case, and vice versa. However, the FP objective can be mixed max-and-min, e.g., one may wish to maximize the signal-to-interference-plus-noise ratio (SINR) of the legitimate receiver while minimizing that of the eavesdropper. We aim to fill the gap between max-FP and min-FP by devising a unified optimization framework. The main results are three-fold. First, we extend the existing max-FP technique called quadratic transform to the min-FP, and further develop a full generalization for the mixed case. Second. we provide a minorization-maximization (MM) interpretation of the proposed unified approach, thereby establishing its convergence and also obtaining a matrix extension; another result we obtain is a generalized Lagrangian dual transform which facilitates the solving of the logarithmic FP. Finally, we present three typical applications: the age-of-information (AoI) minimization, the Cramer-Rao bound minimization for sensing, and the secure data rate maximization, none of which can be efficiently addressed by the previous FP methods.

cs.IT

Multi-Linear Pseudo-PageRank for Hypergraph Partitioning

Motivated by the PageRank model for graph partitioning, we develop an extension of PageRank for partitioning uniform hypergraphs. Starting from adjacency tensors of uniform hypergraphs, we establish the multi-linear pseudo-PageRank (MLPPR) model, which is formulated as a multi-linear system with nonnegative constraints. The coefficient tensor of MLPPR is a kind of Laplacian tensors of uniform hypergraphs, which are almost as sparse as adjacency tensors since no dangling corrections are incorporated. Furthermore, all frontal slices of the coefficient tensor of MLPPR are M-matrices. Theoretically, MLPPR has a solution, which is unique under mild conditions. An error bound of the MLPPR solution is analyzed when the Laplacian tensor is slightly perturbed. Computationally, by exploiting the structural Laplacian tensor, we propose a tensor splitting algorithm, which converges linearly to a solution of MLPPR. Finally, numerical experiments illustrate that MLPPR is powerful and effective for hypergraph partitioning problems.

math.OC

A Low Rank Quaternion Decomposition Algorithm and Its Application in Color Image Inpainting

In this paper, we propose a lower rank quaternion decomposition algorithm and apply it to color image inpainting. We introduce a concise form for the gradient of a real function in quaternion matrix variables. The optimality conditions of our quaternion least squares problem have a simple expression with this form. The convergence and convergence rate of our algorithm are established with this tool.

math.OC

A Tensor Rank Theory and Maximum Full Rank Subtensors

A matrix always has a full rank submatrix such that the rank of this matrix is equal to the rank of that submatrix. This property is one of the corner stones of the matrix rank theory. We call this property the max-full-rank-submatrix property. Tensor ranks play a crucial role in low rank tensor approximation, tensor completion and tensor recovery. However, their theory is still not matured yet. Can we set an axiom system for tensor ranks? Can we extend the max-full-rank-submatrix property to tensors? We explore these in this paper. We first propose some axioms for tensor rank functions. Then we introduce proper tensor rank functions. The CP rank is a tensor rank function, but is not proper. There are two proper tensor rank functions, the max-Tucker rank and the submax-Tucker rank, which are associated with the Tucker decomposition. We define a partial order among tensor rank functions and show that there exists a unique smallest tensor rank function. We introduce the full rank tensor concept, and define the max-full-rank-subtensor property. We show the max-Tucker tensor rank function and the smallest tensor rank function have this property. We define the closure for an arbitrary proper tensor rank function, and show that it is still a proper tensor rank function and has the max-full-rank-subtensor property. An application of the submax-Tucker rank is also presented.

math.RA

A Polynomially Irreducible Functional Basis of Elasticity Tensors

Tensor function representation theory is an essential topic in both theoretical and applied mechanics. For the elasticity tensor, Olive, Kolev and Auffray (2017) proposed a minimal integrity basis of 297 isotropic invariants, which is also a functional basis. Inspired by Smith's and Zheng's works, we use a novel method in this article to seek a functional basis of the elasticity tensor, that contains less number of isotropic invariants. We achieve this goal by constructing 22 intermediate tensors consisting of 11 second order symmetrical tensors and 11 scalars via the irreducible decomposition of the elasticity tensor. Based on such intermediate tensors, we further generate 429 isotropic invariants which form a functional basis of the elasticity tensor. After eliminating all the invariants that are zeros or polynomials in the others, we finally obtain a functional basis of 251 isotropic invariants for the elasticity tensor.

math-ph

Triple Decomposition and Tensor Recovery of Third Order Tensors

In this paper, we introduce a new tensor decomposition for third order tensors, which decomposes a third order tensor to three third order low rank tensors in a balanced way. We call such a decomposition the triple decomposition, and the corresponding rank the triple rank. For a third order tensor, its CP decomposition can be regarded as a special case of its triple decomposition. The triple rank of a third order tensor is not greater than the middle value of the Tucker rank, and is strictly less than the middle value of the Tucker rank for an essential class of examples. These indicate that practical data can be approximated by low rank triple decomposition as long as it can be approximated by low rank CP or Tucker decomposition. This theoretical discovery is confirmed numerically. Numerical tests show that third order tensor data from practical applications such as internet traffic and video image are of low triple ranks. A tensor recovery method based on low rank triple decomposition is proposed. Its convergence and convergence rate are established. Numerical experiments confirm the efficiency of this method.

math.NA

Tensor Norm, Cubic Power and Gelfand Limit

We establish two inequalities for the nuclear norm and the spectral norm of tensor products. The first inequality indicates that the nuclear norm of the square matrix is a matrix norm. We extend the concept of matrix norm to tensor norm. We show that the $1$-norm, the Frobenius norm and the nuclear norm of tensors are tensor norms, but the infinity norm and the spectral norm of tensors are not tensor norms. We introduce the cubic power for a general third order tensor, and show that a Gelfand formula holds for a general third order tensor. In that formula, for any norm, a common spectral radius-like limit exists for that third order tensor. We call such a limit the Gelfand limit. The Gelfand limit is zero if the third order tensor is nilpotent, and is one or zero if the third order tensor is idempotent. The Gelfand limit is not greater than any tensor norm of that third order tensor, and the cubic power of that third order tensor tends to zero as the power increases to infinity if and only if the Gelfand limit is less than one. The cubic power and the Gelfand limit can be extended to any higher odd order tensors.

math.NA

An Irreducible Polynomial Functional Basis of Two-dimensional Eshelby Tensors

Representation theorems for both isotropic and anisotropic functions are of prime importance in both theoretical and applied mechanics. The Eshelby inclusion problem is very fundamental, and is of particular importance in the design of advanced functional composite materials. In this paper, we discuss about two-dimensional Eshelby tensors (denoted as $M^{(2)}$). Eshelby tensors satisfy the minor index symmetry $M_{ijkl}^{(2)}=M_{jikl}^{(2)}=M_{ijlk}^{(2)}$ and have wide applications in many fields of mechanics. In view of the representation of two-dimensional irreducible tensors in complex field, we obtain a minimal integrity basis of ten isotropic invariants of $M^{(2)}$. Remarkably, note that an integrity basis is always a functional basis, we further confirm that the minimal integrity basis is also an irreducible function basis of isotropic invariants of $M^{(2)}$.

math-ph