SearcharxivSearch

arXiv subjects

Zihao Zhang

Publications and source records attributed to Zihao Zhang.

At least 19 recordsLinked to original sources

Unimodality for IDP Lattice Simplices of Prime Normalized Volume

Recently, Ferroni constructed a family of counterexamples to the well-known conjecture in Ehrhart theory stating that the $h^*$-polynomial of a lattice polytope with the integer decomposition property is unimodal. This raises the question of whether the $h^*$-polynomial of a lattice simplex with the integer decomposition property remains unimodal. In this note, we prove that every lattice simplex with the integer decomposition property and prime normalized volume has a unimodal $h^*$-polynomial. Furthermore, we establish several sufficient conditions for the unimodality of the $h^*$-polynomial of such simplices.

math.CO

An Empirical Study of Harness Design for Coding Agents

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.

cs.AI

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model's world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.

cs.CV

Real-Rootedness and Gamma-Positivity for a Variation of the Morris Constant Term

Beck and Pixton expressed the Ehrhart polynomial of the Birkhoff polytope as a weighted sum of constant terms of several multivariate rational functions. Xin and Zhang studied a class of constant terms $h_n(t)$, which can be regarded as a variation of the Morris constant term. They proved that $h_n(t)$ is a polynomial of degree $(n-1)^2$ and obtained many nice properties involving the Morris constant term identity. Let $h_n^*(y)=(1-y)^{(n-1)^2+1}\sum_{t\geq0}h_n(t)y^t$. For fixed $n\geq 3$, we obtain the following three main results: (i): $h_n^*(y)$ is a polynomial with positive integer coefficients. (ii): $h_n^*(y)$ is real-rooted. In particular, all its roots are non-positive real numbers. (iii): $h_n^*(y)$ is Gamma-positive. Furthermore, $h_n^*(y)$ is palindromic, unimodal, and ultra log-concave. This confirms Xin and Zhang's conjecture regarding $h_n^*(y)$. As a byproduct, we prove that every root of a Gamma polynomial associated with $h_n^*(y)$ is a negative real number.

math.CO

MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

Training-free zero-shot composed image retrieval finds a target image in a gallery from a reference image and a text edit without learning from task-specific image triplets. Existing methods typically describe the target as a whole and match this description with a global image representation. This global matching can mix different semantic cues and lose fine- grained details. We propose MULVEC, a role-aware method whose compiler produces a structured query record that is mapped to four retrieval roles: Global describes the full target, Desired states what should appear, Preserve states what should remain, and Forbidden states what should disappear. Frozen encoders map the query to one target description vector and role-specific probe vectors, while each candidate is represented by one global visual vector and a bank of local visual vectors. The retrieval roles then use this shared evidence for their respective purposes, and a fixed weighted sum of their scores ranks the entire gallery in a single retrieval pass. Across CIRCO, CIRR, and FashionIQ and three backbone scales, MULVEC improves CIRCO mAP@5 by up to 23.0% over the strongest compared method and gives the best CIRR and FashionIQ results in our comparison.

cs.CV

Taylor Positivity of Ehrhart Polynomials

Let $P$ be a $d$-dimensional lattice polytope with Ehrhart polynomial $L_P(t)$. Motivated by the study of Ehrhart positivity and magic positivity, we investigate the Taylor coefficients $\mathsf{A}_j(P;k)$ in the shifted expansion $L_P(t)=\sum_{j=0}^{d}\mathsf{A}_j(P;k)(t-k)^j$ about a real center $k$. In this paper, we obtain the following four main results. (i) We give exact formulas for these coefficients in terms of the ordinary Ehrhart coefficients, the $h^*$-vector, elementary symmetric functions, and Stirling numbers. (ii) We denote by $τ(P)$ and $τ^+(P)$ the smallest nonnegative integral centers at which all Taylor coefficients are nonnegative and positive, respectively. If $s$ is the degree of the $h^*$-polynomial, then $0\leqτ(P)\leqτ^+(P)\leq\min\{\max\{0,s-1\},\lfloor\frac{d-1}{2}\rfloor\}$. As an application, we slightly improve an upper bound due to Beck, De Loera, Develin, Pfeifle, and Stanley. That is, every real root of $L_P(t)$ lies in $[-d,\lfloor\frac{d-1}{2}\rfloor)$. (iii) Let $ρ(P)$ be the smallest nonnegative real center such that the Taylor coefficients are nonnegative. If $λ_{\mathbb{R}}(f)$ denotes the largest real zero of $f(t)$, with value $-\infty$ when no such zero exists, then $ρ(P)=\max\{0,\max_{0\leq j<d}λ_{\mathbb{R}}\!(L_P^{(j)})\}$. (iv) We establish structural properties of the Taylor coefficients $\mathsf{A}_j(P;k)$, including derivative interlacing, palindromic reflection symmetries, and Laguerre and Newton inequalities. As a final note, these results provide a systematic partial answer to an open problem listed on the website of the American Institute of Mathematics.

math.CO

Adaptive Beam Hopping and Power Control for Dual-Layer Over-the-Air Online Federated Learning in LEO Satellite Networks

This paper investigates over-the-air (OTA) computation enabled online federated learning (FL) in low-Earth orbit (LEO) satellite networks. Specifically, we consider a dual-layer OTA aggregation architecture, where ground devices upload analog model updates to serving satellites via uplink OTA aggregation, and satellites forward the aggregated signals to a data processing center through the second round OTA aggregation. Then, we formulate a long-term data-utilization maximization problem in which devices continuously collect new data and untrained samples gradually lose freshness. The problem is subject to the satellite beam budget, transmit-power limit, and global mean squared error (MSE) constraint that governs end-to-end aggregation distortion. This yields a coupled mixed-integer nonlinear programming (MINLP) problem, involving tightly coupled discrete beam-hopping decisions and continuous power control. Due to the combinatorial action space and nonconvex constraints, the problem is NP-hard and computationally intractable. Furthermore, the time-varying satellite topology and dynamic data generation render it a sequential decision-making problem, necessitating adaptive online scheduling. To address these issues, we cast the problem as a Markov decision process and develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that jointly optimizes adaptive beam hopping and power control, using an MSE-aware reward to balance data utilization and aggregation accuracy. Numerical simulation results verify that the proposed algorithm consistently outperforms other benchmark schemes, achieving superior long-term data utilization and faster FL convergence while satisfying the MSE requirement.

cs.IT

Unequal urban capacities for mobility adaptation under fuel-price shocks

What a city makes reachable depends less on what it contains than on who can still afford to move when travel costs rise. We leverage the 2026 US-Iran oil shock as a natural experiment, applying a hierarchical panel regression discontinuity design to 1.7 trillion point-of-interest visits across 122,000 neighbourhoods in China and the United States. Mobility range declined in nearly three-quarters of neighbourhoods, but responses varied systematically with pre-shock urban conditions. Exposure to energy-intensive travel explained the largest share of modelled heterogeneity in both countries, while adaptive capacity and activity composition further shaped how travel was reorganized. Longer baseline travel intensified contraction, whereas greater car dependence constrained adjustment. Crucially, similar mobility outcomes arose from different processes: some neighbourhoods maintained travel by absorbing higher costs, whereas others appeared structurally locked into travel they could not reorganize. Fuel-price shocks, therefore, act as urban stress tests, revealing which neighbourhoods a city keeps connected.

stat.AP

AREAs-Lab: An Interactive Environment for AI-driven Requirement Elicitation for AI Systems

Building effective AI systems increasingly depends on writing high-quality task requirements, yet users often struggle to articulate the constraints, preferences, and edge cases that determine success. This problem is especially acute in AI development, where behavior is shaped not only by human expectations but also by data characteristics. We present AREAs-Lab, an interactive environment for AI-driven Requirement Elicitation for AI systems. In AREAs-Lab, an assistant iteratively refines an initially incomplete requirement by analyzing the underlying dataset and asking targeted clarification questions to uncover the user's latent intent. To study this setting systematically, we construct a synthetic benchmark grounded in 16 public datasets spanning diverse domains and task types. Each benchmark instance includes a user profile, a complete reference requirement, and an intentionally underspecified version that serves as the assistant's starting point. We further introduce an automated evaluation pipeline based on an AI-simulated user that reveals hidden information only when appropriately prompted, enabling scalable and reproducible assessment of interactive elicitation quality. AREAs-Lab provides a controlled testbed for studying how AI assistants can transform vague user goals into actionable requirements for AI systems.

cs.HC

Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise convolutions, we identify a critical oversight: pointwise convolutions dominate parameter volume (>87% in models like RepLKNet-31B) and constitute the primary deployment bottleneck on resource-constrained edge devices. This results in prohibitive storage costs and severe memory-loading constraints on resource-limited devices (e.g., smartphones with 4-12 GB Random Access Memory (RAM)). To overcome this, we propose Channel Group-Shared (CGS) low-rank approximation, a novel Singular Value Decomposition (SVD)-based parameter-sharing strategy. CGS constructs a structured low-rank paradigm isomorphic to SVD decomposition, comprising shared (high-parameter-cost) down/up-projection matrices across channel groups within a layer and channel-group-specific (low-parameter-cost) scalable diagonal matrices. This group-sharing design achieves significant parameter reduction. Extensive experiments demonstrate that large-kernel CNNs (RepLKNet, ConvNeXt, SLaK) enhanced with CGS strike an empirically favorable balance between competitive performance and substantially reduced storage costs. Crucially, by alleviating storage constraints, reducing memory bandwidth pressure during loading, and minimizing model loading latency, CGS enables the feasible deployment of pre-trained large-kernel CNN models on edge devices, thereby bridging the gap between high-performance vision models and practical edge deployment.

cs.LG

AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

Multi-vector representations have emerged as an effective paradigm for multimodal retrieval, representing each sample with multiple complementary embeddings to capture fine-grained cross-modal information. However, existing approaches typically employ a fixed representation capacity, assigning the same number of vectors to all samples regardless of their individual retrieval demands. Such a fixed-capacity formulation overlooks the fact that different samples may require different amounts of representation capacity for effective retrieval. In this work, we introduce \emph{Sample-Adaptive Multi-Vector Representation} (SAMVR), a new problem setting for multimodal retrieval that studies how multi-vector representation capacity can be allocated at the sample level. Under SAMVR, each sample is represented by a \emph{content-adaptive embedding set} (CAES), whose capacity is determined according to the sample-specific retrieval utility of additional representation vectors. To instantiate SAMVR, we propose \emph{AdaptiveEmbed}, a unified framework for learning sample-adaptive multi-vector representations. AdaptiveEmbed learns structured multi-vector representations through \emph{Multi-Group Contrastive Learning} (MGCL) with the symmetric \emph{set-to-set similarity} (SetSim), and further employs \emph{Utility Policy Optimization} (UPO) to determine sample-specific representation capacity via \emph{Marginal Utility Allocation} (MUA). Experiments across multimodal retrieval benchmarks involving image, text, video, and audio show that sample-adaptive capacity allocation achieves overall better retrieval performance than fixed-capacity multi-vector representations, validating the effectiveness of SAMVR for multimodal retrieval. These results establish SAMVR as a viable formulation for adaptive capacity allocation in multi-vector multimodal retrieval.

cs.CV

Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing

Supervised proxy-based deep cross-modal hashing has become the dominant paradigm for large-scale retrieval. However, prevalent methods model class proxies as deterministic points in the embedding space. This rigid assumption causes severe gradient conflicts in multi-label scenarios, where gradient conflicts arising from label co-occurrence lead to severe gradient contention and optimization collapse. To resolve this, we propose Kent-based Distributional Proxy Hashing (KDPH), a novel framework that shifts proxy representation from static points to flexible anisotropic Kent distributions on the hypersphere. Unlike point proxies that must shift their positions to accommodate conflicting gradients, KDPH absorbs these conflicts by dynamically adjusting its directional variance. This allows the proxy to maintain a stable semantic mean direction while stretching to cover diverse label correlations. Furthermore, to ensure stable training of these geometric parameters, we derive a tailored loss function incorporating the Cayley transform to enforce strict orthogonality. To the best of our knowledge, KDPH is the first framework to successfully introduce the Kent distributions into cross-modal hashing. Experiments on three benchmark datasets demonstrate that KDPH mitigates proxy collapse and chaotic oscillation, significantly outperforms state-of-the-art methods. Code is available at https://github.com/Senmo996/KDPH-official-code.

cs.CV

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically condition the drafter on a fixed visual interface, such as a predefined visual-token budget or a static compressed representation. However, our controlled visual-budget analysis shows that visual demand varies substantially across tasks and decoding stages, which means more visual input is not always beneficial. Actually, insufficient evidence may weaken visual grounding, while excessive context adds overhead and may disrupt drafting. We propose FOVEA (Focused On-demand Visual Evidence Adaptation), a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state. A cumulative-mass rule determines both how many and which entries are selected. The selected entries are aggregated into a visual readout and fused with the current draft hidden state through a lightweight gated residual correction. Rather than inserting visual tokens into the autoregressive context, the correction modifies only the representation passed to the language-model head. Experiments across multiple vision-language backbones and multimodal benchmarks show that FOVEA improves draft acceptance and end-to-end decoding speed, achieving up to $2.13\times$ speedup over autoregressive decoding. These results demonstrate that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout multimodal generation.

cs.CV

The Ehrhart series of magic squares of orders seven and eight

Let $\mathrm{IMS}_n(m)$ count the $n\times n$ nonnegative integer matrices whose row sums, column sums, main-diagonal sum, and antidiagonal sum are all $m$. We determine the Ehrhart series $F_n(q)=\sum_{m\geq0}\mathrm{IMS}_n(m)q^m$ as reduced rational functions for $n=7$ and $n=8$. Their numerator--denominator degrees are respectively $(366,373)$ and $(540,548)$. Both numerators have positive integer coefficients, are palindromic and strictly unimodal. The proofs share one finite architecture: a signed SimpCone decomposition is evaluated by quotient characters over finite fields, a certified common denominator and Ehrhart reciprocity reduce the rational identity to finitely many coefficients, an explicit counting bound lifts modular congruences to integer equalities, and exact gcd computations prove reducedness. For order eight, a face-index pole certificate gives a degree-$598$ common denominator without enumerating the full face lattice, leaving $296$ independent coefficients in degrees $0$ through $295$. Once the candidate rational function is known, the first six production primes certify this finite prefix by the same bounded-coefficient argument used for order seven.

math.CO

A polynomial time algorithm for almost bounded denumerant

Sylvester's denumerant $d(t; \boldsymbol{A})$ counts the number of nonnegative integer solutions to $\sum_{i=1}^{N} a_i x_i = t$, where $\boldsymbol{A} = (a_1, \dots, a_N)$ is a sequence of positive integers with $\gcd(\boldsymbol{A}) = 1$. In 2025, Xin and Zhang gave a polynomial time algorithm in $N$ for computing $d(t; \boldsymbol{A})$ when the entries of $\boldsymbol{A}$ are bounded by a constant. In this paper, we extend this algorithm by incorporating Barvinok's algorithm, enabling it to handle the case where a fixed number of entries of $\boldsymbol{A}$ are allowed to be unbounded.

math.CO

Polynomial-Time Lattice-Point Counting without Barvinok Decomposition

By using constant term manipulations, we present the first polynomial-time algorithm for lattice-point counting in fixed dimension that does not rely on Barvinok's unimodular decomposition. The algorithm instead operates directly on a rational generating function in the form of a nested root average, as produced by the \texttt{SimpCone[S]} framework. By means of a residue-lattice argument based on Minkowski's theorem, we construct a short multiplier that induces an exact non-coprime split of the outermost average. The resulting child terms are encoded as joint root averages, and Smith normal form is used to restore the recursive structure. Two structural invariants---the generation condition and full-column independence---ensure that the recursion is well defined and that all required pole exchanges are valid. For a fixed-dimensional simplicial cone, the algorithm achieves recursion depth \(O_d(1+\log\log(2+\ind(\mathcal K^*)))\) and produces a signed sum of at most \((1+\log \ind(\mathcal K^*))^{O_d(1)}\) unimodular cone generating functions. The framework uniformly handles numerators that are Laurent polynomials, not merely monomials, thereby giving a polynomial-time algorithm for MacMahon's partition analysis when the dimension is fixed.

math.CO

Hankel Transform and $(α,β)$ Somos-4 Sequences

An $(α,β)$ Somos-$4$ sequence $S_n$ is defined by the recurrence $S_nS_{n-4}=αS_{n-1}S_{n-3}+βS_{n-2}^2$ ($n\geq 4$), with suitable initial values, where $α$ and $β$ are constant parameters. A widely studied question is the following: When does the Hankel transform of a generating function become an $(α,β)$ Somos-4 sequence? In particular, how can $α$ and $β$ be derived for such a function? A sufficient condition for this problem has been established by Wang and Zhang. In this paper, we obtain the following three main results. (i): We extend the Wang--Zhang sufficient condition by working over the rational function field. Then we combine this result with the Sulanke--Xin quadratic transformation to resolve all of Barry's currently unsolved $(α,β)$ Somos-4 conjectures, which arise in diverse contexts, including generalized Catalan recurrences, Riordan arrays, generalized Bernstein arrays, and elliptic curves. (ii): We show that the odd and even subsequences of an $(α,β)$ Somos-4 sequence are again $(α,β)$ Somos-4 sequences with transformed parameters. This is employed to establish Barry's Hurwitz transform conjecture. (iii): Using the theory of orthogonal polynomials, we prove a Hankel determinant formula and thereby prove a conjecture related to the $(α,β)$ Somos-4 sequence. In addition, we prove some conjectures on formulas for periodic Hankel determinants.

math.CO

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still lack reliable object-level grounding diagnostics and therefore tend to apply coarse-grained interventions, which can impair visual understanding, shorten responses, and reduce coverage of genuinely grounded objects. The key challenge is thus to detect, during generation, whether each emerging object mention is supported by reliable visual evidence, so that hallucination can be mitigated selectively. Yet output confidence reflects next-token plausibility rather than visual support, allowing language priors to make absent objects appear certain. We show that the missing diagnostic evidence is encoded in an Intrinsic Grounding Signature (IGS), a distributed signed attention pattern that remains informative for such confident hallucinations. Based on IGS, we propose Verifier-Guided Decoding (VGD), a decoding framework in which a lightweight verifier examines each emerging object mention, rolls back the KV cache when the mention is identified as high risk, suppresses the object and its synonyms, and regenerates the affected continuation. Because VGD intervenes only on object mentions identified as high risk, it reduces object hallucination while preserving the model's original visual understanding and grounded object coverage. Experiments on CHAIR and AMBER-G show that VGD achieves state-of-the-art object hallucination reduction: at @rec90, it cuts AMBER-G CHAIR by 43.6\% while retaining 99.6\% of grounded-object coverage, and reduces CHAIR-MSCOCO CHAIR$_i$/CHAIR$_s$ by 37.0\%/30.4\% without shortening captions.

cs.CV