Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.

cs.CL↗

Pulsed Accretion onto Eccentric Binaries in Highly Misaligned Circumbinary Disks

We present three-dimensional smoothed particle hydrodynamics simulations of highly misaligned circumbinary disks (CBDs) around moderately eccentric equal-mass binaries ($e_\mathrm{b}=0.5$). We show that the binary accretion is modulated on the binary orbital period and exhibits two pulses near periastron. The dominant pulse peaks before periastron for an initial binary-disk misalignment of $60^\circ$, shifts to after periastron at $90^\circ$, and occurs at an even later post-periastron phase at $120^\circ$. We further show that the two pulses are accompanied by a time-dependent response of the circumstellar disks (CSDs) and by different distributions of accreting material within the cavity and around the CSDs. The qualitative pre- versus post-periastron distinction is also present in individual binary orbits despite variations in pulse amplitude. Our results motivate future tests of whether pulse timing is related to binary-disk orientation.

astro-ph.EP↗

Emergence of Nanoscale Modulation in Liquid Crystals as a Result of Short-Range Order Parameter Condensation

The discovery of the twist-bend nematic phase in liquid crystals composed of bent-core and dimeric molecules has revealed an unexpected mechanism for the spontaneous formation of nanoscale periodic structures in soft condensed matter. Unlike conventional liquid-crystalline phases, the twist-bend phase exhibits a nanoscale heliconical modulation despite the absence of molecular chirality. In this note, dedicated to the late R. Meyer, we discuss a Landau phenomenological interpretation of this phenomenon. The central idea is that a short-range orientational order parameter undergoes condensation, giving rise to a heliconical structure characterized by a finite wave vector. We emphasize that the conventional nematic order is already long-ranged, whereas the additional order parameter describes local orientational correlations hidden within the nematic state. The resulting phase transition bears a close analogy to the de Gennes theory of the nematic--smectic-A transition. Fluctuation effects are expected to drive the transition weakly first order. The theory also predicts a new Goldstone mode associated with the spontaneously broken continuous symmetry of the heliconical state

cond-mat.soft↗

Representation-Aware Transport-Information Measure for Non-inclusive Discrete Supports

Information-theoretic measures for comparing probability distributions are widely used across physics and other fields. When two discrete distributions have non-inclusive supports, however, the Kullback-Leibler (KL) divergence is in general not directly applicable, and various alternative divergences and distances have been introduced. These measures compare the resulting distributions themselves, but do not generally retain information about the representation transformations by which the discrete distributions are generated from underlying continuous ones. Here we introduce a representation-aware transport-information measure for discrete distributions with non-inclusive supports, formulated based on the standard KL divergence. We consider two continuous reference distributions, each transformed into a discrete representation through its own discretization scheme. Rather than comparing only the resulting discrete distributions or their continuous references, we additionally retain local information associated with the representation-change schemes. The resulting measure can therefore distinguish discrete representations that may have identical discrete probability landscapes but originate from different continuous references or discretization schemes. The construction is based on the transport-information cost of continuous-to-discrete representation in the framework of unavoidable canonical nonlinearity (UCN). UCN provides a non-arbitrary correspondence between the transport cost of discretization as an extrinsic geometric operation and the information-theoretic indistinguishability of nearby continuous distributions, thereby allowing a discrete representation to be associated with a local family of underlying continuous distributions on the statistical manifold.

cond-mat.stat-mech↗

Deterministic Linear-Time Modular Subset Sum

We give a deterministic $O(m)$-time algorithm for exact modular subset sum over every modulus $m$ on compact input: distinct residues with multiplicities. It reports all reachable residues and answers one target query, returning a witness when the target is reachable. The algorithm uses $O(m)$ auxiliary words on an arithmetic word-RAM. This improves Potępa's deterministic $O(m\log mα(m))$-time bound under the same input convention. The running time matches the cost of explicitly reporting all $m$ reachability bits. We represent reachable residues as runs along cycles of repeated addition. Newly reached residues pay for scans of partial runs, and processing prime factors in increasing order makes cycle rebuilding linear. A theorem on subset sums of distinct units limits the number of adaptive boundary batches to $O(m^{3/4})$; radix sorting their $O(m)$ total keys also takes $O(m)$ time.

cs.DS↗

PipeDRAM: A Data-Transposition-Free Processing-Using-DRAM Architecture with Hardware/Software Pipelining

Processing-using-DRAM (PUD) architectures exploit the analog operational properties of DRAM to perform bulk bitwise Boolean and arithmetic operations inside memory arrays by organizing data in a vertical layout, where operand bits are stacked along DRAM columns. However, modern computing systems natively employ a horizontal data layout that preserves the cache line abstraction, leverages spatial locality in row buffers, and enables high memory throughput. This fundamental mismatch forces existing PUD architectures to frequently perform data layout transformations between horizontal and vertical formats, incurring significant performance, energy, and system integration overheads. Our goal is to eliminate data transposition overheads in PUD systems at low cost. To this end, we propose PipeDRAM, a PUD architecture that eliminates the need for runtime data layout transformation, enabling PUD operations directly over horizontally laid-out data. PipeDRAM's key ideas are to (i) deterministically reorganize bits inside each memory request to enable a PUD-friendly data placement within a DRAM array in a horizontal data layout, and (ii) employ a pipeline-based execution model that overlaps bit-dependent and bit-independent in-DRAM operations to exploit bit-level parallelism across the memory array. We compare PipeDRAM to different computing platforms. PipeDRAM provides (i) 11.8x, 11.8x, and 80.4x higher performance and (ii) 25.4x, 3.0x, and 38.0x lower energy consumption than three state-of-the-art PUD systems. PipeDRAM incurs low area cost on top of a DRAM chip (1.86%) and CPU die (0.05%). To enable further research on PUD systems, we open-source PipeDRAM at https://github.com/CMU-SAFARI/PipeDRAM.

cs.AR↗

DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation

Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.

cs.RO↗

TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback

Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.

cs.CV↗

An ETH-Tight, Constructive FPT Algorithm for the Cone and Polytope Intersection Problem

In a landmark paper, Goemans and Rothvoss (2020) established an XP algorithm running in time $\text{enc}(P)^{2^{O(d)}} \cdot \text{enc}(Q)^{O(1)}$ for the Cone and Polytope Intersection problem: finding a vector $y \in \operatorname{int{.}cone}(P \cap \mathbb{Z}^d) \cap Q$ together with a sparse certificate $λ\in \mathbb{Z}_{\ge 0}^{P \cap \mathbb{Z}^d}$ supported on at most $2^{2d+1}$ generators, where $P \subseteq \mathbb{R}^d$ is a bounded rational polyhedron and $Q \subseteq \mathbb{R}^d$ is an arbitrary rational polyhedron. For high-multiplicity bin packing, this gives a running time of ${|I|}^{2^{O(d)}}$, where $|I|$ denotes the encoding length of the input. Recently, Koana and Kumabe (2026) proved that the decision variant of this problem is fixed-parameter tractable (FPT) parameterized by the number of item types $d$ with running time $2^{d^{O(d)}} \cdot {|I|}^{O(1)} = 2^{2^{O(d \log d)}} \cdot {|I|}^{O(1)}$. In this work, we generalize the framework of Koana and Kumabe from standard bin packing to the full Cone and Polytope Intersection Problem of Goemans and Rothvoss, directly encompassing high-multiplicity bin packing, point-in-cone, and scheduling. Secondly, by combining Carathéodory-type integer cone bounds (Eisenbrand and Shmonin, 2006) with active support enumeration, we reduce the running time to: $$2^{2^{O(d)}} \cdot (\text{enc}(P) + \text{enc}(Q))^{O(1)}.$$ Under the Exponential Time Hypothesis (ETH), the double-exponential lower bound of Kowalik, Lassota, Majewski, Pilipczuk, and Sokołowski (2024) for point-in-cone and Jansen, Ohnesorge, and Pirotton (2026) for high-multiplicity bin packing implies that this parameter dependence is asymptotically optimal. Finally, we provide an explicit decompression algorithm that extracts a solution with sparse support $|\text{supp}(λ)| \le 2^{2d+1}$ in single-exponential FPT time.

cs.DS↗

Toward provably private learning from federated data

Federated Learning (FL) allows devices with private data to collaborate in training a shared model. We present a next-generation FL system based on Trusted Execution Environments (TEEs) that addresses operational challenges associated with earlier systems and provides externally verifiable central Differential Privacy (DP) guarantees for the first time while offering a better privacy-utility tradeoff. In our system, devices upload data encrypted with keys managed by a TEE-hosted Key Management Service (KMS). The uploaded data is cryptographically tied to a policy limiting the set of Python programs that may later process the data in server-side TEEs. External parties may inspect public transparency logs to observe the set of workloads allowed by these policies. Our experimental results show that the new system improves device coverage and favorably shifts privacy-utility curves by enabling collected data to be integrated into the server-side workload at a schedule that optimizes DP guarantees and is unaffected by device availability. Our new system has been productionized, enabling models for the Android Keyboard (Gboard) to be trained faster and achieve better accuracy under smaller, now externally verifiable privacy budgets in comparison to models trained using the prior system.

cs.CR↗

Riesz kernels of hyperbolic polynomials: positivity, admissible exponents and Jordan rigidity

Scott and Sokal asked whether every homogeneous polynomial with the half-plane property has a completely monotone negative power. We answer this question affirmatively and prove the Riesz-positivity conjecture of Michałek, Sturmfels, Uhler and Zwiernik, restated by Kozhasov, Michałek and Sturmfels: in $n$ variables, every exponent $α\ge4096n^2$ is admissible, independently of the degree and coefficients. For complete hyperbolic polynomials the Riesz density is strictly log-concave, with relative Gaussian error at most $512n^2/α$ and explicit curvature bounds. We characterize admissible exponents by a common spectral Dirichlet law, prove $n\le m+αm(m-1)$ with its equality case, and obtain the sharp degree-dependent gap $0<α<1/(2(m-1))$ whenever the degree-$m$ polynomial has a nonlinear irreducible factor. A nonnegative fourth-order defect of $-\log p$ vanishes at one point precisely for products of positive integer powers of Euclidean Jordan determinants. We classify the corresponding logarithmic Monge-Amp{è}re equation and answer the question of Etingof, Kazhdan and Polishchuk about polynomial multiplicative Legendre transforms within irreducible complete hyperbolic polynomials; the general question has counterexamples, the Clifford quartics of Kogiso and Sato. The classification extends to reducible polynomials under a boundary-visibility hypothesis, and asymptotic common-power formulas for the Riesz densities force exact Jordan formulas.

math.CO↗

Gap-free Differentially Private PCA for Gaussian Data

We give a gap-free $(ε,δ)$-differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data. The algorithm is based on a private variant of the power iteration method, and it is computationally efficient.

cs.DS↗

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.

cs.LG↗

Temporal-Attention Head Specialization During Video Diffusion Training

Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mechanism comes from analyzing trained models, so when and where temporal-attention structure forms during training remains poorly characterized. Population averages can also hide it, since a few specializing heads and a diffusing majority cancel in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters), scoring each head with an entropy-normalized measure of cross-frame attention concentration (CFAC) under a preregistered change-point and effect-size selection rule. The census reveals the sparse picture that averages obscure. Aggregate CFAC is flat or decreasing in every run, while a small minority of heads, roughly 4--13% in full-grid runs, develops pronounced concentration. Across seeds, the reproducible signal is positional but block-level. Selected heads repeatedly arise in the first temporal block, whereas individual head coordinates do not reproduce once block membership is accounted for. Among the analyzed 760M selected heads, attention maps converge to a small repertoire of local frame-routing motifs, self-frame diagonals and adjacent-frame bands, even when the responsible coordinates differ across runs. Correlation and ablation analyses do not establish a causal link to generated video quality, and we bound our claims accordingly. Beyond this STDiT family, the study contributes a transferable methodology. Checkpoint-resolved, per-head analysis under fixed selection rules can expose sparse temporal organization in other factorized video diffusion transformers and, with adapted routing metrics, in joint spatio-temporal architectures.

cs.CV↗

Video-to-Music Generation for Gameplay Videos

Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of clean soundtracks, free of sound effects and voice-overs, matched to gameplay audio via audio fingerprinting. With this dataset, we train a simple encoder-decoder transformer that passes video features directly to a MusicGen decoder, comparing different encoding strategies: textual descriptions (T5), independent frames (ViT), or spatiotemporal patches (ViViT). Each encoder is tested both frozen and fine-tuned, while the decoder is always fine-tuned. Frozen encoders match or outperform their fine-tuned counterparts on every metric, and the frozen ViViT achieves the best overall results. We compare this model with state-of-the-art baselines using both objective metrics and a listening study (N = 96). Despite having up to 18% fewer parameters, our model outperforms all baselines on objective metrics, surpasses GVMGen in the listening study, and performs comparably to OSSL.

cs.SD↗

Read-Rezayi fractional Chern insulators in modulated Bernal graphene

Fibonacci anyons provide a universal platform for topological quantum computation, and emerge as low-energy excitations in the $\mathbb{Z}_3$ Read-Rezayi phase in the fractional quantum Hall effect. However, realistic microscopic realizations of this phase in the absence of a magnetic field have remained elusive. We study a model of periodically modulated Bernal bilayer graphene with gate-screened Coulomb interactions. Using the recently developed target-phase optimization method in conjunction with band-projected exact diagonalization, we identify at filling $ν=3/5$ a region of parameter space whose ground state is consistent with a Read-Rezayi fractional Chern insulator. The partially filled band from which it arises is a part of a two-band complex which mimics geometric aspects of the lowest and first Landau levels, with the ground state at $ν=1/2$ consistent with the Moore-Read state. Our results suggest that modulated Bernal graphene can realize delicate non-Abelian fractional quantum Hall states at zero magnetic field, while demonstrating target-phase optimization as a practical route to discovering such phases in realistic, high-dimensional microscopic models.

cond-mat.str-el↗

Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems

The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of distance for heterogeneous datasets which may admit isomorphic representations. In order to accelerate computation of GW distances, many practitioners employ entropic regularization to obtain an Entropic GW (EGW) problem. The most popular EGW solver is the Mirror Descent (MD) algorithm, which reduces EGW computations to an iterative process where an entropic optimal transport (EOT) problem is solved at each iteration. Despite its widespread use, the convergence of MD for this problem has only been established for restricted classes of costs. On the other hand, a recently proposed dual gradient method is available for general costs, but requires a choice of step size which depends on the regularization parameter. To address these two issues, we introduce Averaged Mirror Descent (AMD), which averages consecutive MD steps, and prove its convergence for arbitrary costs. Then, we establish that the dual gradient method with a fixed step size also converges for arbitrary costs at the cost of a more complicated iteration. In both cases, we also account for inexact iterations which are inescapable in practice. We compare the empirical performance of these methods across various settings and, in particular, show that AMD and the dual gradient method both converge on an example where classical MD fails.

cs.LG↗

LLM Judge Validation Under Sparse Overlap: From Inference to Design

Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $ρ\geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

cs.AI↗