SearcharxivSearch

arXiv subjects

Ke Ye

Publications and source records attributed to Ke Ye.

At least 19 recordsLinked to original sources

Kempe factorizations for rational curves on $\operatorname{SO}_4(\mathbb{R})$

We study constructive Kempe factorizations for rational curves on $\operatorname{SO}_4(\mathbb{R})$. Motivated by motion-polynomial factorization and rational matrix curves on real classical groups, we prove that every rational curve of degree $2d$ with $d\ge1$ first factors into $d$ quadratic rational curves and then into a product of at most $2d$ planar rotation curves, where each planar rotation curve fixes a two-dimensional plane pointwise. The construction proceeds by extracting left and right isoclinic polynomial parts via Cayley's factorization, factoring the corresponding quaternion polynomials, and pairing linear factors with equal norm polynomials. The resulting algorithms are explicit and are illustrated by examples.

math.AG

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, organization strategies, and procedural knowledge through continued use. We present MindMemOS, a portable and self-evolving memory operating layer that organizes open-world information using a unified entity property timestructure. MindMemOS supports scenario-adaptive memory modeling, higher-order pattern discovery, autonomous memory refinement, and continuous skill evolution. Its MindMemEvolve algorithm employs validation-driven evolutionary search to optimize memory schemas for target scenarios, whiledreaming consolidates accumulated memories by merging redundant records and resolving conflicts. In addition, implicit corrective feedback serves as a human-in-the-loop signal for identifying and revising potentially inaccurate or misaligned memories. Its MindSkillEvolve algorithm further transforms agent execution trajectories into reusable and progressively refined skills. MindMemOS achieves 94.03% accuracy on LOCOMO and 70.63% on PersonaMem. MindSkillEvolve improves SpreadsheetBench success by 9.2 percentage points over the initial-skill baseline.

cs.AI

Complexity of Low-Degree Skew Polynomial Multiplication over Finite Fields

In this note, we study the complexity of multiplication in skew polynomial rings over finite fields. We prove that the product of two elements in $\mathbb{F}_{q^n}[x;\sigma]$ of degree at most $d < n$ can be computed using $\widetilde O(d^{\omega_K-1}n)$ arithmetic operations over $\mathbb{F}_q$, where $\sigma$ is the $q$-Frobenius automorphism. This matches the conjectural upper bound of Caruso--Le Borgne~[ISSAC'17] and is quasi-optimal in view of the lower bound of Chen--Ye [ISSAC'24]. The proof reduces the finite-field case to the split algebra case using the equivariant multiplication theory of Couveignes--Ezome~[J.~Algebra, 2023], and then applies existing fast algorithms.

cs.SC

An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing

In recent years, there have been notable advancements in the area of instruction-based image editing (IIE), which focuses on the automatic alteration of input images using a model. Nevertheless, assessing the effectiveness of these editing models poses a considerable challenge due to the intricate nature of instructions and the wide variety of edits. To tackle this problem, one urgent task in this domain is the development of a robust evaluation framework that can precisely gauge the quality of editing outcomes and offer valuable benchmarks to guide future improvements. To address this challenge, we present a comprehensive evaluation benchmark named I2EBench2.0, designed for single-round and multi-round assessment of IIE models. I2EBench2.0 has four key features: 1) Evaluation Across Single and Multi-rounds: I2EBench2.0 simultaneously evaluates both single-round and multi-round instruction-based edits, assessing the precision and consistency of the edits. 2) Extensive Evaluation Criteria: I2EBench2.0 encompasses a broad range of criteria, evaluating both high-level and low-level aspects of each IIE model. Specifically, it incorporates 16 dimensions for single-round evaluations and 7 for multi-round evaluations. 3) Alignment with Human Judgment: To ensure our benchmark aligns with human evaluation, we conducted a comprehensive user study for each criterion. 4) Research-driven Insights: By analyzing the strengths and weaknesses of current IIE models across all 16 single-round and 7 multi-round dimensions, we provide critical insights aimed at directing future research in this area. We tested eight recently developed IIE models using I2EBench2.0 and derived academic insights through meticulous comparison and analysis. The related code, dataset, and images generated by all IIE models are available on GitHub: https://github.com/cocoshe/I2EBench.

cs.CV

Linear representations of manifolds

A finite-dimensional linear representation of a group or an algebra may be regarded as a map into a space of matrices, endowing abstract elements with coordinates, and encoding algebraic operations as matrix products. With this in mind, we define a linear representation of a $\mathsf{G}$-manifold $\mathcal{M} $ as a map into a space of matrices, representing points as matrices and the $\mathsf{G}$-action as matrix products. We show that this generalizes group representations to any $\mathsf{G}$-manifold that may not have a group structure, with homogeneous spaces $\mathsf{G}/\mathsf{H}$ an important special case; and in this case it also generalizes Cartan embeddings of symmetric spaces to more general $\mathsf{G}/\mathsf{H}$. To demonstrate the utility of such manifold representations, we use them to provide effective bounds for Mostow-Palais $\mathsf{G}$-equivariant embeddings of $\mathsf{G}$-manifolds into $\mathsf{G}$-modules $\mathbb{V}$. Unlike Whitney and Nash embeddings, Mostow-Palais embeddings have no known effective bounds; before our work, it was only known that $\dim \mathbb{V} < \infty$ if $\mathsf{G}$ is compact. We will give explicit values for $\dim \mathbb{V}$ and show that our bounds are sharp. Furthermore, our method is constructive, giving explicit expressions for these minimal-dimensional Mostow-Palais embeddings.

math.DG

Geometry of multilinear varieties over infinite fields and its applications

Multilinear varieties, defined as the sets of rational points of varieties cut out by multilinear functions, were first introduced and studied by Gowers and Mili\'{c}evi\'{c}[Proc. Edinb. Math. Soc., 2021] for finite $\mathbb{K}$. In this paper, we investigate multilinear varieties over infinite fields from a geometric perspective. We establish two fundamental results: a codimension formula for the Zariski closure of a multilinear variety, and the existence of a high-dimensional irreducible subvariety passing through any given $\mathbb{K}$-rational point. These results serve as a geometric foundation for analyzing various ranks of tensors and homogeneous polynomials, including partition rank, analytic rank, geometric rank, (collective) strength and (collective) Birch rank. As applications, we resolve the Adiprasito-Kazhdan-Ziegler conjecture [arXiv:2102.03659, 2021] on the stability of partition rank for perfect infinite fields. We thereby settle the stability conjecture for collective strength [Selecta Math., 2024], as well as the conjecture on the linear equivalence between strength and Birch rank [arXiv:2410.00248, 2024] for such fields. Moreover, our results immediately yield a strengthening of the theorems of Bik-Draisma-Snowden [arXiv:2401.02067, 2024] and Lampert-Snowden [arXiv:2406.18498, 2024], for multilinear varieties over infinite fields.

math.AG

Tur\'{a}n problems for multilinear maps

We study Tur\'{a}n-type extremal problems for alternating and unrestricted multilinear maps. For alternating order-$d$ multilinear maps $T: (\mathbb F^n)^d\to \mathbb{F}^m$, we determine, over algebraically closed fields of arbitrary characteristic, the largest $k$ such that every $T$ vanishes identically on $\mathbb{V}^d$ for some $k$-dimensional subspace $\mathbb{V}$. This extends the bilinear formula of Buhler, Gupta, and Harris [J. Algebra, 1987] to arbitrary order and resolves a question of Qiao [Discrete Anal., 2023]. We also solve the analogous problem for arbitrary, not necessarily alternating, multilinear maps by determining the largest $k$ such that every $T$ vanishes on $\mathbb{V}_1\times\cdots\times \mathbb{V}_d$ for some $k$-dimensional subspaces $\mathbb{V}_1,\dots,\mathbb{V}_d$. These results yield exact values, over algebraically closed fields, of the Feldman--Propp number [Adv. Math., 1992], the Tur\'{a}n number [Discrete Anal., 2023], and the Gow--Quinlan number [Linear Multilinear Algebra, 2006] associated with alternating multilinear maps. Finally, motivated by the Erd\H{o}s box problem, we give a purely algebraic derivation of the Conlon--Pohoata--Zakharov lower bound [Discrete Anal., 2021] by combining analytic and partition rank estimates with an incidence count. In the relevant parameter range, we further show that every multilinear map defined over a finite field has many isotropic tuples of $2$-dimensional subspaces over extensions of sufficiently divisible degree. This rules out the natural route to improving the Conlon--Pohoata--Zakharov exponent by selecting multilinear maps with substantially fewer bad isotropic configurations.

math.CO

ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control

Expert-level scientific reasoning remains challenging for large language models, particularly on benchmarks such as Humanity's Last Exam (HLE), where rigid tool pipelines, brittle multi-agent coordination, and inefficient test-time scaling often limit performance. We introduce ReThinker, a confidence-aware agentic framework that orchestrates retrieval, tool use, and multi-agent reasoning through a stage-wise Solver-Critic-Selector architecture. Rather than following a fixed pipeline, ReThinker dynamically allocates computation based on model confidence, enabling adaptive tool invocation, guided multi-dimensional reflection, and robust confidence-weighted selection. To support scalable training without human annotation, we further propose a reverse data synthesis pipeline and an adaptive trajectory recycling strategy that transform successful reasoning traces into high-quality supervision. Experiments on HLE, GAIA, and XBench demonstrate that ReThinker consistently outperforms state-of-the-art foundation models with tools and existing deep research systems, achieving state-of-the-art results on expert-level reasoning tasks.

cs.AI

From Text to Simulation: A Multi-Agent LLM Workflow for Automated Chemical Process Design

Process simulation is a critical cornerstone of chemical engineering design. Current automated chemical design methodologies focus mainly on various representations of process flow diagrams. However, transforming these diagrams into executable simulation flowsheets remains a time-consuming and labor-intensive endeavor, requiring extensive manual parameter configuration within simulation software. In this work, we propose a novel multi-agent workflow that leverages the semantic understanding capabilities of large language models(LLMs) and enables iterative interactions with chemical process simulation software, achieving end-to-end automated simulation from textual process specifications to computationally validated software configurations for design enhancement. Our approach integrates four specialized agents responsible for task understanding, topology generation, parameter configuration, and evaluation analysis, respectively, coupled with Enhanced Monte Carlo Tree Search to accurately interpret semantics and robustly generate configurations. Evaluated on Simona, a large-scale process description dataset, our method achieves a 31.1% improvement in the simulation convergence rate compared to state-of-the-art baselines and reduces the design time by 89. 0% compared to the expert manual design. This work demonstrates the potential of AI-assisted chemical process design, which bridges the gap between conceptual design and practical implementation. Our workflow is applicable to diverse process-oriented industries, including pharmaceuticals, petrochemicals, food processing, and manufacturing, offering a generalizable solution for automated process design.

cs.AI

RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs

Reinforcement learning (RL) has emerged as the de-facto paradigm for improving the reasoning capabilities of large language models (LLMs). We have developed RLAX, a scalable RL framework on TPUs. RLAX employs a parameter-server architecture. A master trainer periodically pushes updated model weights to the parameter server while a fleet of inference workers pull the latest weights and generates new rollouts. We introduce a suite of system techniques to enable scalable and preemptible RL for a diverse set of state-of-art RL algorithms. To accelerate convergence and improve model quality, we have devised new dataset curation and alignment techniques. Large-scale evaluations show that RLAX improves QwQ-32B's pass@8 accuracy by 12.8% in just 12 hours 48 minutes on 1024 v5p TPUs, while remaining robust to preemptions during training.

cs.LG

Atlas Gaussian processes on restricted domains and point clouds

In real-world applications, data often reside in restricted domains with unknown boundaries, or as high-dimensional point clouds lying on a lower-dimensional, nontrivial, unknown manifold. Traditional Gaussian Processes (GPs) struggle to capture the underlying geometry in such settings. Some existing methods assume a flat space embedded in a point cloud, which can be represented by a single latent chart (latent space), while others exhibit weak performance when the point cloud is sparse or irregularly sampled. The goal of this work is to address these challenges. The main contributions are twofold: (1) We establish the Atlas Brownian Motion (BM) framework for estimating the heat kernel on point clouds with unknown geometries and nontrivial topological structures; (2) Instead of directly using the heat kernel estimates, we construct a Riemannian corrected kernel by combining the global heat kernel with local RBF kernel and leading to the formulation of Riemannian-corrected Atlas Gaussian Processes (RC-AGPs). The resulting RC-AGPs are applied to regression tasks across synthetic and real-world datasets. These examples demonstrate that our method outperforms existing approaches in both heat kernel estimation and regression accuracy. It improves statistical inference by effectively bridging the gap between complex, high-dimensional observations and manifold-based inferences.

stat.ML

Upper Bounds for $s$-Distance Subspaces

As a generalization of equiangular lines, equiangular subspaces were first systematically studied by Balla, Dr\"{a}xler, Keevash and Sudakov in 2017. In this paper, we extend their work to $s$-distance subspaces, i.e., to sets of $k$-dimensional subspaces in $\mathbb{R}^n$ whose pairwise distances take $s$ distinct values. We establish upper bounds on the maximum cardinality of such sets. In particular, our bounds generalize and improve results of Balla and Sudakov.

math.MG

Isotropy and completeness indices of multilinear maps

Structures of multilinear maps are characterized by invariants. In this paper we introduce two invariants, named the isotropy index and the completeness index. These invariants capture the tensorial structure of the kernel of a multilinear map. We establish bounds on both indices in terms of the partition rank, geometric rank, analytic rank and height, and present three applications: 1) Using the completeness index as an interpolator, we establish upper bounds on the aforementioned tensor ranks in terms of the subrank. This settles an open problem raised by Kopparty, Moshkovitz and Zuiddam, and consequently answers a question of Derksen, Makam and Zuiddam. 2) We prove a Ramsey-type theorem for the two indices, generalizing a recent result of Qiao and confirming a conjecture of his. 3) By computing the completeness index, we obtain a polynomial-time probabilistic algorithm to estimate the height of a polynomial ideal.

math.CO

Extremal constructions for apex partite hypergraphs

We establish new lower bounds for the Tur\'an and Zarankiewicz numbers of certain apex partite hypergraphs. Given a $(d-1)$-partite $(d-1)$-uniform hypergraph $\mathcal{H}$, let $\mathcal{H}(k)$ be the $d$-partite $d$-uniform hypergraph whose $d$th part has $k$ vertices that share $\mathcal{ H}$ as a common link. We show that $ex(n,\mathcal{H}(k))=\Omega_{\mathcal{ H}}(n^{d-\frac{1}{e(\mathcal{H})}})$ if $k$ is at least exponentially large in $e(\mathcal{H})$. Our bound is optimal for all Sidorenko hypergraphs $\mathcal{H}$ and verifies a conjecture of Lee for such hypergraphs. In particular, for the complete $d$-partite $d$-uniform hypergraphs $\mathcal{K}^{(d)}_{s_1,\dots,s_d}$, our result implies that $ex(n,\mathcal{K}^{(d)}_{s_{1},\cdots,s_{d}})=\Theta(n^{d-\frac{1}{s_{1}\cdots s_{d-1}}})$ if $s_{d}$ is at least exponentially large in terms of $s_{1}\cdots s_{d-1}$, improving the factorial condition of Pohoata and Zakharov and answering a question of Mubayi. Our method is a generalization of Bukh's random algebraic method [Duke Math.J. 2024] to hypergraphs, and extends to the sided Zarankiewicz problem.

math.CO

EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations

Reliable trajectory prediction from an ego-centric perspective is crucial for robotic navigation in human-centric environments. However, existing methods typically assume noiseless observation histories, failing to account for the perceptual artifacts inherent in first-person vision, such as occlusions, ID switches, and tracking drift. This discrepancy between training assumptions and deployment reality severely limits model robustness. To bridge this gap, we introduce EgoTraj-Bench, built upon TBD dataset, which is the first real-world benchmark that aligns noisy, first-person visual histories with clean, bird's-eye-view future trajectories, enabling robust learning under realistic perceptual constraints. Building on this benchmark, we propose BiFlow, a dual-stream flow matching model that concurrently denoises historical observations and forecasts future motion. To better model agent intent, BiFlow incorporates our EgoAnchor mechanism, which conditions the prediction decoder on distilled historical features via feature modulation. Extensive experiments show that BiFlow achieves state-of-the-art performance, reducing minADE and minFDE by 10-15% on average and demonstrating superior robustness. We anticipate that our benchmark and model will provide a critical foundation for robust real-world ego-centric trajectory prediction. The benchmark library is available at: https://github.com/zoeyliu1999/EgoTraj-Bench.

cs.CV

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabilities, typically fail to decompose high-level commands into executable action sequences from static visual input alone. To address this challenge, we introduce Super-Mimic, a hierarchical framework that enables zero-shot robotic imitation by directly inferring procedural intent from unscripted human demonstration videos. Our framework is composed of two sequential modules. First, a Human Intent Translator (HIT) parses the input video using multimodal reasoning to produce a sequence of language-grounded subtasks. These subtasks then condition a Future Dynamics Predictor (FDP), which employs a generative model that synthesizes a physically plausible video rollout for each step. The resulting visual trajectories are dynamics-aware, explicitly modeling crucial object interactions and contact points to guide the low-level controller. We validate this approach through extensive experiments on a suite of long-horizon manipulation tasks, where Super-Mimic significantly outperforms state-of-the-art zero-shot methods by over 20%. These results establish that coupling video-driven intent parsing with prospective dynamics modeling is a highly effective strategy for developing general-purpose robotic systems.

cs.RO

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

Large Language Models (LLMs) demonstrate impressive mathematical reasoning abilities, but their solutions frequently contain errors that cannot be automatically checked. Formal theorem proving systems such as Lean 4 offer automated verification with complete accuracy, motivating recent efforts to build specialized prover LLMs that generate verifiable proofs in formal languages. However, a significant gap remains: current prover LLMs solve substantially fewer problems than general-purpose LLMs operating in natural language. We introduce Hilbert, an agentic framework that bridges this gap by combining the complementary strengths of informal reasoning and formal verification. Our system orchestrates four components: an informal LLM that excels at mathematical reasoning, a specialized prover LLM optimized for Lean 4 tactics, a formal verifier, and a semantic theorem retriever. Given a problem that the prover is unable to solve, Hilbert employs recursive decomposition to split the problem into subgoals that it solves with the prover or reasoner LLM. It leverages verifier feedback to refine incorrect proofs as necessary. Experimental results demonstrate that Hilbert substantially outperforms existing approaches on key benchmarks, achieving 99.2\% on miniF2F, 6.6\% points above the best publicly available method. Hilbert achieves the \textbf{strongest known result} from a publicly available model on PutnamBench. It solves 462/660 problems (70.0\%), outperforming proprietary approaches like SeedProver (50.4\%) and achieving a 422\% improvement over the best publicly available baseline. Thus, Hilbert effectively narrows the gap between informal reasoning and formal proof generation. Code is available at https://github.com/Rose-STL-Lab/ml-hilbert.

cs.AI

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?

Pairwise preferences over model responses are widely collected to evaluate and provide feedback to large language models (LLMs). Given two alternative model responses to the same input, a human or AI annotator selects the "better" response. This approach can provide feedback for domains where other hard-coded metrics are difficult to obtain (e.g., chat response quality), thereby helping model evaluation or training. However, for some domains high-quality pairwise comparisons can be tricky to obtain - from AI and humans. For example, for responses with many factual statements, annotators may disproportionately weigh writing quality rather than underlying facts. In this work, we explore augmenting standard AI annotator systems with additional tools to improve performance on three challenging response domains: long-form factual, math and code tasks. We propose a tool-using agentic system to provide higher quality feedback on these domains. Our system uses web-search and code execution to ground itself based on external validation, independent of the LLM's internal knowledge and biases. We provide extensive experimental results evaluating our method across the three targeted response domains as well as general annotation tasks, using RewardBench (incl. AlpacaEval and LLMBar), RewardMath, as well as three new datasets for domains with saturated pre-existing datasets. Our results indicate that external tools can indeed improve performance in many, but not all, cases. More generally, our experiments highlight the sensitivity of performance to simple parameters (e.g., prompt) and the need for improved (non-saturated) annotator benchmarks. We share our code at https://github.com/apple/ml-agent-evaluator.

cs.CL