SearcharxivSearch

arXiv subjects

Ruixiang Zhang

Publications and source records attributed to Ruixiang Zhang.

At least 19 recordsLinked to original sources

Mixed-norm Brascamp-Lieb inequalities

Christ [Chr01] looked at a certain boundedness problem for trilinear operators. In this paper we set up a framework of mixed-norm Brascamp-Lieb inequalities and identify a new interesting regime not covered by the study of classical Brascamp-Lieb. Positive results in [Chr01] can be viewed as the first nontrivial progress in this new regime. We will also prove another mixed-norm Brascamp-Lieb inequality, showcasing how one can use a tensor product trick to slightly sharpen Christ's argument and also obtain the endpoint case. We then discuss the connections between mixed-norm Brascamp-Lieb and the Kakeya Conjectures, as well as unique new difficulties for these mixed-norm Brascamp-Lieb inequalities compared to the classical setting.

math.CA

Curved Kakeya problems and the projective geometry of paths

We introduce a general framework for curved Kakeya problems in $\mathbb{R}^n$, encompassing those arising from H\"ormander-type oscillatory integrals. Every family of curves determines a spray geometry, which allows us to use the projective geometry of paths in the study of curved Kakeya problems. We focus on the two extremes of the "best" and "worst" possible behaviors of curved Kakeya sets. We characterize when the incidence structure underlying Wolff's hairbrush argument persists. In particular, we prove that the existence of many totally geodesic surfaces, as required by Wolff's hairbrush argument, is equivalent to projective flatness of the associated spray. Within this projectively flat class, Bourgain's condition provides a clean dichotomy: when it holds, the family is direction-equivalent to a Bochner--Riesz type family of lines and satisfies the Katz--Wolff condition, and thus the Wang--Zahl result is applicable; when it fails, every totally geodesic surface supports a two-dimensional Kakeya set. We also show that under an extra semi-algebraic assumption, a family of curves in $\mathbb{R}^3$ admits a curved Kakeya set of Hausdorff dimension $2$ if and only if it admits a curved Kakeya set contained in a surface. Equivalently, if this compression is absent, every associated curved Kakeya set has dimension strictly greater than $2$.

math.CA

CSPMNet: Pareto-Efficient Automatic Modulation Classification With Learnable Complex Subband Phase Motion

Automatic modulation classification (AMC) is an essential technique for noncooperative spectrum monitoring and intelligent wireless receivers. However, practical AMC models must identify modulation formats from short and noisy I/Q observations while maintaining low computational and storage overhead. Existing deep-learning approaches often improve recognition accuracy by expanding generic neural backbones, which increases deployment cost and weakens their suitability for resource-constrained receivers. To bridge the gap between recognition performance and model efficiency, this letter proposes a Complex Subband Phase-Motion Network, designated as CSPMNet, for lightweight AMC from raw I/Q samples. Specifically, learnable complex subband filters are introduced to adaptively extract frequency-selective baseband responses while preserving the algebraic coupling between in-phase and quadrature components. Then, an amplitude-preserving phase-motion module captures multi-lag temporal rotation dynamics within each subband, and a lightweight temporal classifier performs efficient sequence aggregation. Rigorous experimental evaluations on public RadioML benchmark datasets demonstrate that CSPMNet achieves highly competitive recognition accuracy while requiring substantially lower model complexity than many existing AMC models.

eess.SP

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling

Code generation is typically trained in the primal space of programs: a model produces a candidate solution and receives sparse execution feedback, often a single pass/fail bit. Test-time scaling enriches the inference procedure by sampling multiple candidates and judging among them, but the comparative information this process reveals is discarded after inference. We argue that this information defines a dual judgment space that provides a far richer training signal: the model learns not from an isolated success or failure, but from the relative correctness structure across its own plausible attempts, identifying which succeed, which fail, and what distinguishes them. We introduce DuST (Dual Self-Training), a framework for self-training from the dual judgment space. DuST samples candidate programs from the model's own distribution, labels them through sandbox execution, retains groups containing both successes and failures, and trains the model to rank candidates by execution correctness using GRPO. The objective is purely discriminative: the model is never directly rewarded for generating correct programs. Dual self-training improves both judgment and generation. Across five models spanning two families and three scales (4B to 30B), DuST consistently improves Best-of-4 test-time scaling on LiveCodeBench. For Qwen3-30B-Thinking on LiveCodeBench v6, judgment quality improves by +6.2 NDCG, single-sample pass@1 improves by +3.1, and Best-of-4 accuracy improves by +4.1. The trained model's single rollout matches the base model's Best-of-4 performance. SFT on the same ranking data improves judgment without improving generation, confirming that on-policy RL is the mechanism that transfers dual-space learning back into primal generation.

cs.LG

SEI-SHIELD: Robust Specific Emitter Identification Under Label Noise Via Self-Supervised Filtering and Iterative Rescue

Specific Emitter Identification (SEI) provides physical-layer device authentication for wireless communications and Internet of Things (IoT) systems. While deep learning (DL) has significantly advanced SEI performance, label noise severely degrades system reliability in non-cooperative environments. Label noise originates from channel-induced ambiguities, annotation errors, and deliberate data poisoning by intelligent jammers injecting misleading signals. While recent SEI methods attempt to mitigate label noise, they fundamentally rely on corrupted supervised signals to guide sample selection, inevitably leading to confirmation bias and suboptimal feature spaces. To address this challenge, we propose SEI-SHIELD, a robust SEI framework that integrates self-supervised contrastive pre-training with iterative sample selection. Specifically, SEI-SHIELD employs Momentum Contrast (MoCo) with RF-tailored augmentations to extract intrinsically robust, label-independent representations directly from complex-valued I/Q signals. In addition, K-nearest neighbors (KNN)-based noise filtering identifies corrupted samples through neighborhood label consistency analysis in the learned feature space. Furthermore, an iterative rescue mechanism using prediction confidence and prototype cosine similarity progressively recovers correctly labeled hard samples inadvertently discarded during filtering. Comprehensive experiments on the POWDER and ORACLE datasets demonstrate that SEI-SHIELD achieves state-of-the-art (SOTA) accuracy under various noise rates, substantially outperforming existing noise-robust paradigms, including advanced regularization techniques and sample selection frameworks.

eess.SP

Generalized Small Object Detection:A Point-Prompted Paradigm and Benchmark

Small object detection (SOD) remains challenging due to extremely limited pixels and ambiguous object boundaries. These characteristics lead to challenging annotation, limited availability of large-scale high-quality datasets, and inherently weak semantic representations for small objects. In this work, we first address the data limitation by introducing TinySet-9M, the first large-scale, multi-domain dataset for small object detection. Beyond filling the gap in large-scale datasets, we establish a benchmark to evaluate the effectiveness of existing label-efficient detection methods for small objects. Our evaluation reveals that weak visual cues further exacerbate the performance degradation of label-efficient methods in small object detection, highlighting a critical challenge in label-efficient SOD. Secondly, to tackle the limitation of insufficient semantic representation, we move beyond training-time feature enhancement and propose a new paradigm termed Point-Prompt Small Object Detection (P2SOD). This paradigm introduces sparse point prompts at inference time as an efficient information bridge for category-level localization, enabling semantic augmentation. Building upon the P2SOD paradigm and the large-scale TinySet-9M dataset, we further develop DEAL (DEtect Any smalL object), a scalable and transferable point-prompted detection framework that learns robust, prompt-conditioned representations from large-scale data. With only a single click at inference time, DEAL achieves a 31.4% relative improvement over fully supervised baselines under strict localization metrics (e.g., AP75) on TinySet-9M, while generalizing effectively to unseen categories and unseen datasets. Our project is available at https://zhuhaoraneis.github.io/TinySet-9M/.

cs.CV

Embarrassingly Simple Self-Distillation Improves Code Generation

Can a large language model (LLM) improve at code generation using only its own raw outputs, without a verifier, a teacher model, or reinforcement learning? We answer in the affirmative with simple self-distillation (SSD): sample solutions from the model with certain temperature and truncation configurations, then fine-tune on those samples with standard supervised fine-tuning. SSD improves Qwen3-30B-Instruct from 42.4% to 55.3% pass@1 on LiveCodeBench v6, with gains concentrating on harder problems, and it generalizes across Qwen and Llama models at 4B, 8B, and 30B scale, including both instruct and thinking variants. To understand why such a simple method can work, we trace these gains to a precision-exploration conflict in LLM decoding and show that SSD reshapes token distributions in a context-dependent way, suppressing distractor tails where precision matters while preserving useful diversity where exploration matters. Taken together, SSD offers a complementary post-training direction for improving LLM code generation. Our code is available at https://github.com/apple/ml-ssd

cs.CL

Power loss for the Mizohata-Takeuchi conjecture on $C^k$ convex hypersurfaces

We find a family of compact $C^k$ hypersurfaces where the local Mizohata-Takeuchi Conjecture fails with a power loss of $R^{\alpha}$ for any $\alpha<\frac{n-1}{n-1+k}$. Moreover, this family is dense in the $C^k$ topology, and so the local Mizohata-Takeuchi conjecture fails for many convex hypersurfaces. In particular, the local Mizohata-Takeuchi Conjecture fails with a power loss of $R^\alpha$ for any $\alpha<\frac{n-1}{n+1}$ for many $C^2$ convex hypersurfaces. This power matches the best known upper bound in a paper by Tony Carbery, Marina Iliopoulou and Hong Wang up to the endpoint. For the proof, our weight is positive definite as in the first author's recent $\log(R)$-loss counterexample, and our construction is based on a projection of a higher rank lattice. As a by-product, we also construct compact convex $C^2$ hypersurfaces whose rescaling contains many lattice points in any dimension.

math.CA

Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling

Standard discrete diffusion models treat all unobserved states identically by mapping them to an absorbing [MASK] token. This creates an 'information void' where semantic information that could be inferred from unmasked tokens is lost between denoising steps. We introduce Continuously Augmented Discrete Diffusion (CADD), a framework that augments the discrete state space with a paired diffusion in a continuous latent space. This yields graded, gradually corrupted states in which masked tokens are represented by noisy yet informative latent vectors rather than collapsed 'information voids'. At each reverse step, CADD may leverage the continuous latent as a semantic hint to guide discrete denoising. The design is clean and compatible with existing discrete diffusion training. At sampling time, the strength and choice of estimator for the continuous latent vector enables a controlled trade-off between mode-coverage (generating diverse outputs) and mode-seeking (generating contextually precise outputs) behaviors. Empirically, we demonstrate CADD improves generative quality over mask-based diffusion across text generation, image synthesis, and code modeling, with consistent gains on both qualitative and quantitative metrics against strong discrete baselines.

stat.ML

ChipChat: Low-Latency Cascaded Conversational Agent in MLX

The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.

eess.AS

Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows

Autoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility. In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space. We propose a novel framework TarFlowLM, that employs transformer-based autoregressive normalizing flows to model these continuous representations. This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process. We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models. Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework.

cs.LG

DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation

Diffusion large language models (dLLMs) are compelling alternatives to autoregressive (AR) models because their denoising models operate over the entire sequence. The global planning and iterative refinement features of dLLMs are particularly useful for code generation. However, current training and inference mechanisms for dLLMs in coding are still under-explored. To demystify the decoding behavior of dLLMs and unlock their potential for coding, we systematically investigate their denoising processes and reinforcement learning (RL) methods. We train a 7B dLLM, \textbf{DiffuCoder}, on 130B tokens of code. Using this model as a testbed, we analyze its decoding behavior, revealing how it differs from that of AR models: (1) dLLMs can decide how causal their generation should be without relying on semi-AR decoding, and (2) increasing the sampling temperature diversifies not only token choices but also their generation order. This diversity creates a rich search space for RL rollouts. For RL training, to reduce the variance of token log-likelihood estimates and maintain training efficiency, we propose \textbf{coupled-GRPO}, a novel sampling scheme that constructs complementary mask noise for completions used in training. In our experiments, coupled-GRPO significantly improves DiffuCoder's performance on code generation benchmarks (+4.4\% on EvalPlus) and reduces reliance on AR bias during decoding. Our work provides deeper insight into the machinery of dLLM generation and offers an effective, diffusion-native RL training framework. https://github.com/apple/ml-diffucoder.

cs.CL

STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis

We present STARFlow, a scalable generative model based on normalizing flows that achieves strong performance in high-resolution image synthesis. The core of STARFlow is Transformer Autoregressive Flow (TARFlow), which combines the expressive power of normalizing flows with the structured modeling capabilities of Autoregressive Transformers. We first establish the theoretical universality of TARFlow for modeling continuous distributions. Building on this foundation, we introduce several key architectural and algorithmic innovations to significantly enhance scalability: (1) a deep-shallow design, wherein a deep Transformer block captures most of the model representational capacity, complemented by a few shallow Transformer blocks that are computationally efficient yet substantially beneficial; (2) modeling in the latent space of pretrained autoencoders, which proves more effective than direct pixel-level modeling; and (3) a novel guidance algorithm that significantly boosts sample quality. Crucially, our model remains an end-to-end normalizing flow, enabling exact maximum likelihood training in continuous spaces without discretization. STARFlow achieves competitive performance in both class-conditional and text-conditional image generation tasks, approaching state-of-the-art diffusion models in sample quality. To our knowledge, this work is the first successful demonstration of normalizing flows operating effectively at this scale and resolution.

cs.CV

Discrete Neural Flow Samplers with Locally Equivariant Transformer

Sampling from unnormalised discrete distributions is a fundamental problem across various domains. While Markov chain Monte Carlo offers a principled approach, it often suffers from slow mixing and poor convergence. In this paper, we propose Discrete Neural Flow Samplers (DNFS), a trainable and efficient framework for discrete sampling. DNFS learns the rate matrix of a continuous-time Markov chain such that the resulting dynamics satisfy the Kolmogorov equation. As this objective involves the intractable partition function, we then employ control variates to reduce the variance of its Monte Carlo estimation, leading to a coordinate descent learning algorithm. To further facilitate computational efficiency, we propose locally equivaraint Transformer, a novel parameterisation of the rate matrix that significantly improves training efficiency while preserving powerful network expressiveness. Empirically, we demonstrate the efficacy of DNFS in a wide range of applications, including sampling from unnormalised distributions, training discrete energy-based models, and solving combinatorial optimisation problems.

cs.LG

Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion

Discrete diffusion is a promising framework for modeling and generating discrete data. In this work, we present Target Concrete Score Matching (TCSM), a novel and versatile objective for training and fine-tuning discrete diffusion models. TCSM provides a general framework with broad applicability. It supports pre-training discrete diffusion models directly from data samples, and many existing discrete diffusion approaches naturally emerge as special cases of our more general TCSM framework. Furthermore, the same TCSM objective extends to post-training of discrete diffusion models, including fine-tuning using reward functions or preference data, and distillation of knowledge from pre-trained autoregressive models. These new capabilities stem from the core idea of TCSM, estimating the concrete score of the target distribution, which resides in the original (clean) data space. This allows seamless integration with reward functions and pre-trained models, which inherently only operate in the clean data space rather than the noisy intermediate spaces of diffusion processes. Our experiments on language modeling tasks demonstrate that TCSM matches or surpasses current methods. Additionally, TCSM is versatile, applicable to both pre-training and post-training scenarios, offering greater flexibility and sample efficiency.

cs.LG

What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?

Language models usually use left-to-right (L2R) autoregressive factorization. However, L2R factorization may not always be the best inductive bias. Therefore, we investigate whether alternative factorizations of the text distribution could be beneficial in some tasks. We investigate right-to-left (R2L) training as a compelling alternative, focusing on multiple-choice questions (MCQs) as a test bed for knowledge extraction and reasoning. Through extensive experiments across various model sizes (2B-8B parameters) and training datasets, we find that R2L models can significantly outperform L2R models on several MCQ benchmarks, including logical reasoning, commonsense understanding, and truthfulness assessment tasks. Our analysis reveals that this performance difference may be fundamentally linked to multiple factors including calibration, computability, and directional conditional entropy. We analyze the impact of these factors through controlled simulation studies using arithmetic tasks, where the impacting factors can be better disentangled. Our work demonstrates that exploring alternative factorizations of the text distribution can lead to improvements in LLM capabilities and provides theoretical insights into optimal factorization towards approximating human language distribution, and when each reasoning order might be more advantageous. Our code and checkpoints are released at https://github.com/apple/ml-reversal-blessing.

cs.CL

Composition and Control with Distilled Energy Diffusion Models and Sequential Monte Carlo

Diffusion models may be formulated as a time-indexed sequence of energy-based models, where the score corresponds to the negative gradient of an energy function. As opposed to learning the score directly, an energy parameterization is attractive as the energy itself can be used to control generation via Monte Carlo samplers. Architectural constraints and training instability in energy parameterized models have so far yielded inferior performance compared to directly approximating the score or denoiser. We address these deficiencies by introducing a novel training regime for the energy function through distillation of pre-trained diffusion models, resembling a Helmholtz decomposition of the score vector field. We further showcase the synergies between energy and score by casting the diffusion sampling procedure as a Feynman Kac model where sampling is controlled using potentials from the learnt energy functions. The Feynman Kac model formalism enables composition and low temperature sampling through sequential Monte Carlo.

stat.ML