Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,531 records · Page 85Linked to original sources

RAZOR: Pruning Replaceable Experts in LLMs

Mixture-of-experts (MoE) models activate only a few experts per token yet store the entire expert pool. Whole-expert pruning shrinks that pool, but for reasoning models it must remove experts without eroding reasoning ability. Common scores rank experts by routing frequency or output magnitude, which measures isolated contribution rather than deletion damage. What decides the damage is functional replaceability, whether the surviving computation can reproduce what is removed. A large contribution may be replaceable by the remaining mixture, whereas a small one may carry a direction the survivors cannot recover. We introduce RAZOR, a training-free method that scores replaceability from consensus residuals, the deviations of individual expert outputs from their original weighted mixture. Holding the layer input fixed, these residuals yield the exact output change from deleting one expert, including survivor reweighting and the replacement expert promoted by router refill. RAZOR aggregates this change over calibration tokens and prunes to a layerwise budget using forward passes alone, without gradients, subset search, or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, RAZOR attains the highest macro average over nine reasoning-centered tasks among the evaluated pruning methods in all eight model-budget settings. Against REAP on GLM-4.7-Flash and Qwen3.6-35B-A3B, it gains 2.12-5.59 points on this average and lowers reverse KL in all four comparisons. Retained accuracy is not the whole picture, as pruned Qwen3.6-35B-A3B still shifts in response diversity, formatting, and termination.

cs.LG↗

Energy-selective control of noise-assisted multipulsing by weak optical seeding in the dissipative-soliton-resonance regime

We study pulse-number selection in a stochastic cubic--quintic complex Ginzburg--Landau model of a weakly seeded, normal-dispersion laser in the dissipative-soliton-resonance (DSR) regime. Without noise, a single pulse and pulse pairs persist at the same control parameters, and the total energy of an $N$-pulse state follows a ladder constructed from the single-pulse branch. The final energies of noisy trajectories lie close to the same ladder. Multipulsing therefore does not necessarily lose the single-pulse solution. It can reflect which coexisting state the noisy dynamics reaches. Optical seed injection suppresses energy-dependent multipulsing and, at larger seed power, produces a nonmonotonic DSR energy window. An energy--noise scan shows that energy dominates the multipulse probability, whereas additive noise produces only a modest trend common to all energies, without a noise optimum. The noise-induced formation statistics therefore do not establish canonical stochastic resonance or escape from a pre-existing soliton. Coherent control by a weak monochromatic seed extends predominantly single-pulse operation over a noticeably broader energy window, enhancing dissipative-soliton energy scalability. Within an adiabatic approximation, equal energy sharing among coexisting pulses is stable wherever the single-pulse energy grows less than proportionally with the control energy, as it does over the sampled DSR range. Energy exchange between the pulses then relaxes up to about 200 times more slowly than their total energy.

physics.optics↗

HuGo: LLMs as Whole-Body Policy Code Designers for Humanoid Loco-Manipulation

For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is https://iconlab.negarmehr.com/HuGo/

cs.RO↗

Fully 3GPP-Compatible Long-Range Sensing for LEO-ISAC: A Window-Grid Processing Framework

This paper addresses the long-range sensing problem in bistatic low-Earth-orbit integrated sensing and communication (LEO-ISAC) systems. Conventional OFDM-based sensing schemes fail in LEO-ISAC due to excessive propagation delays, which cause cross-symbol misalignment and unequal signal durations. We propose a window-grid processing framework that leverages the deterministic target geometry to resolve both challenges as a pure receiver-side processing: the transmitted sensing signal is 3GPP-compatible. Simulations demonstrate meter-level ranging accuracy for two targets at bistatic ranges beyond 640 km with 100.8 MHz bandwidth.

cs.IT↗

SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.

cs.CL↗

Pulsed Accretion onto Eccentric Binaries in Highly Misaligned Circumbinary Disks

We present three-dimensional smoothed particle hydrodynamics simulations of highly misaligned circumbinary disks (CBDs) around moderately eccentric equal-mass binaries ($e_\mathrm{b}=0.5$). We show that the binary accretion is modulated on the binary orbital period and exhibits two pulses near periastron. The dominant pulse peaks before periastron for an initial binary-disk misalignment of $60^\circ$, shifts to after periastron at $90^\circ$, and occurs at an even later post-periastron phase at $120^\circ$. We further show that the two pulses are accompanied by a time-dependent response of the circumstellar disks (CSDs) and by different distributions of accreting material within the cavity and around the CSDs. The qualitative pre- versus post-periastron distinction is also present in individual binary orbits despite variations in pulse amplitude. Our results motivate future tests of whether pulse timing is related to binary-disk orientation.

astro-ph.EP↗

Emergence of Nanoscale Modulation in Liquid Crystals as a Result of Short-Range Order Parameter Condensation

The discovery of the twist-bend nematic phase in liquid crystals composed of bent-core and dimeric molecules has revealed an unexpected mechanism for the spontaneous formation of nanoscale periodic structures in soft condensed matter. Unlike conventional liquid-crystalline phases, the twist-bend phase exhibits a nanoscale heliconical modulation despite the absence of molecular chirality. In this note, dedicated to the late R. Meyer, we discuss a Landau phenomenological interpretation of this phenomenon. The central idea is that a short-range orientational order parameter undergoes condensation, giving rise to a heliconical structure characterized by a finite wave vector. We emphasize that the conventional nematic order is already long-ranged, whereas the additional order parameter describes local orientational correlations hidden within the nematic state. The resulting phase transition bears a close analogy to the de Gennes theory of the nematic--smectic-A transition. Fluctuation effects are expected to drive the transition weakly first order. The theory also predicts a new Goldstone mode associated with the spontaneously broken continuous symmetry of the heliconical state

cond-mat.soft↗

Representation-Aware Transport-Information Measure for Non-inclusive Discrete Supports

Information-theoretic measures for comparing probability distributions are widely used across physics and other fields. When two discrete distributions have non-inclusive supports, however, the Kullback-Leibler (KL) divergence is in general not directly applicable, and various alternative divergences and distances have been introduced. These measures compare the resulting distributions themselves, but do not generally retain information about the representation transformations by which the discrete distributions are generated from underlying continuous ones. Here we introduce a representation-aware transport-information measure for discrete distributions with non-inclusive supports, formulated based on the standard KL divergence. We consider two continuous reference distributions, each transformed into a discrete representation through its own discretization scheme. Rather than comparing only the resulting discrete distributions or their continuous references, we additionally retain local information associated with the representation-change schemes. The resulting measure can therefore distinguish discrete representations that may have identical discrete probability landscapes but originate from different continuous references or discretization schemes. The construction is based on the transport-information cost of continuous-to-discrete representation in the framework of unavoidable canonical nonlinearity (UCN). UCN provides a non-arbitrary correspondence between the transport cost of discretization as an extrinsic geometric operation and the information-theoretic indistinguishability of nearby continuous distributions, thereby allowing a discrete representation to be associated with a local family of underlying continuous distributions on the statistical manifold.

cond-mat.stat-mech↗

Deterministic Linear-Time Modular Subset Sum

We give a deterministic $O(m)$-time algorithm for exact modular subset sum over every modulus $m$ on compact input: distinct residues with multiplicities. It reports all reachable residues and answers one target query, returning a witness when the target is reachable. The algorithm uses $O(m)$ auxiliary words on an arithmetic word-RAM. This improves Potępa's deterministic $O(m\log mα(m))$-time bound under the same input convention. The running time matches the cost of explicitly reporting all $m$ reachability bits. We represent reachable residues as runs along cycles of repeated addition. Newly reached residues pay for scans of partial runs, and processing prime factors in increasing order makes cycle rebuilding linear. A theorem on subset sums of distinct units limits the number of adaptive boundary batches to $O(m^{3/4})$; radix sorting their $O(m)$ total keys also takes $O(m)$ time.

cs.DS↗

PipeDRAM: A Data-Transposition-Free Processing-Using-DRAM Architecture with Hardware/Software Pipelining

Processing-using-DRAM (PUD) architectures exploit the analog operational properties of DRAM to perform bulk bitwise Boolean and arithmetic operations inside memory arrays by organizing data in a vertical layout, where operand bits are stacked along DRAM columns. However, modern computing systems natively employ a horizontal data layout that preserves the cache line abstraction, leverages spatial locality in row buffers, and enables high memory throughput. This fundamental mismatch forces existing PUD architectures to frequently perform data layout transformations between horizontal and vertical formats, incurring significant performance, energy, and system integration overheads. Our goal is to eliminate data transposition overheads in PUD systems at low cost. To this end, we propose PipeDRAM, a PUD architecture that eliminates the need for runtime data layout transformation, enabling PUD operations directly over horizontally laid-out data. PipeDRAM's key ideas are to (i) deterministically reorganize bits inside each memory request to enable a PUD-friendly data placement within a DRAM array in a horizontal data layout, and (ii) employ a pipeline-based execution model that overlaps bit-dependent and bit-independent in-DRAM operations to exploit bit-level parallelism across the memory array. We compare PipeDRAM to different computing platforms. PipeDRAM provides (i) 11.8x, 11.8x, and 80.4x higher performance and (ii) 25.4x, 3.0x, and 38.0x lower energy consumption than three state-of-the-art PUD systems. PipeDRAM incurs low area cost on top of a DRAM chip (1.86%) and CPU die (0.05%). To enable further research on PUD systems, we open-source PipeDRAM at https://github.com/CMU-SAFARI/PipeDRAM.

cs.AR↗

DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation

Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.

cs.RO↗

TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback

Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.

cs.CV↗

An ETH-Tight, Constructive FPT Algorithm for the Cone and Polytope Intersection Problem

In a landmark paper, Goemans and Rothvoss (2020) established an XP algorithm running in time $\text{enc}(P)^{2^{O(d)}} \cdot \text{enc}(Q)^{O(1)}$ for the Cone and Polytope Intersection problem: finding a vector $y \in \operatorname{int{.}cone}(P \cap \mathbb{Z}^d) \cap Q$ together with a sparse certificate $λ\in \mathbb{Z}_{\ge 0}^{P \cap \mathbb{Z}^d}$ supported on at most $2^{2d+1}$ generators, where $P \subseteq \mathbb{R}^d$ is a bounded rational polyhedron and $Q \subseteq \mathbb{R}^d$ is an arbitrary rational polyhedron. For high-multiplicity bin packing, this gives a running time of ${|I|}^{2^{O(d)}}$, where $|I|$ denotes the encoding length of the input. Recently, Koana and Kumabe (2026) proved that the decision variant of this problem is fixed-parameter tractable (FPT) parameterized by the number of item types $d$ with running time $2^{d^{O(d)}} \cdot {|I|}^{O(1)} = 2^{2^{O(d \log d)}} \cdot {|I|}^{O(1)}$. In this work, we generalize the framework of Koana and Kumabe from standard bin packing to the full Cone and Polytope Intersection Problem of Goemans and Rothvoss, directly encompassing high-multiplicity bin packing, point-in-cone, and scheduling. Secondly, by combining Carathéodory-type integer cone bounds (Eisenbrand and Shmonin, 2006) with active support enumeration, we reduce the running time to: $$2^{2^{O(d)}} \cdot (\text{enc}(P) + \text{enc}(Q))^{O(1)}.$$ Under the Exponential Time Hypothesis (ETH), the double-exponential lower bound of Kowalik, Lassota, Majewski, Pilipczuk, and Sokołowski (2024) for point-in-cone and Jansen, Ohnesorge, and Pirotton (2026) for high-multiplicity bin packing implies that this parameter dependence is asymptotically optimal. Finally, we provide an explicit decompression algorithm that extracts a solution with sparse support $|\text{supp}(λ)| \le 2^{2d+1}$ in single-exponential FPT time.

cs.DS↗

AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy

Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle when tissue type, species, modality, or acquisition conditions change. We present AxonSynth, a domain-randomized synthetic-data framework for training 3D axon segmentation models without manually annotated real training volumes. AxonSynth generates dense synthetic axon labels with orientation priors that reflect realistic fiber configurations and renders them with randomized density, contrast, bias fields, blur, and noise. A three-class 3D U-Net is trained to predict background, axon sheath and intra-axonal space. We evaluate zero-shot transfer on 10 held-out light-sheet microscopy (LSM) patches from macaque and human brain samples labeled with one of three axonal markers, comparing against calibrated thresholding and Frangi filtering using overlap, corrected detection, false-positive, and topology metrics. On macaque samples, AxonSynth achieved the best corrected Dice and corrected precision (0.826 and 0.851), compared with 0.765 and 0.754 for thresholding and 0.685 and 0.762 for Frangi. On human samples, corrected Dice was comparable to thresholding (0.857 vs. 0.868), while component-count error decreased from 22,504 to 3,377. Across all held-out patches, AxonSynth reduced component-count error in 10/10 patches and Euler-characteristic error in 8/10. These results show that synthetic-label domain randomization can reduce dependence on manual axon annotation while supporting synthetic-to-real 3D segmentation.

cs.CV↗

Singular Backward SDEs for Optimal Control with State Constraints

We investigate a class of backward stochastic differential equations (BSDEs) with at most quadratic growth which explode at a possibly unbounded random horizon, defined through the first hitting of zero of an adapted Ito process. In contrast with the classical theory of singular BSDEs, the explosion is generated by the nonlinear dependence of the generator on the martingale integrand, rather than by a superlinear coercivity condition in the solution component. We construct a minimal singular solution and derive two-sided estimates on its explosion, together with weighted BMO estimates for the martingale integrand. We also obtain the uniqueness and exact explosion rates under additional structural assumptions. For Hamiltonian generators, we establish a verification theorem for an infinite-horizon stochastic optimal control problem with possible non-Markovian state constraints, showing that the BSDE feedback induces the unique optimal constrained law. We finally specialize the theory to exit times of uniformly elliptic Markov diffusions and recover the connection with large solutions of viscous Hamilton-Jacobi equations.

math.PR↗

Riesz kernels of hyperbolic polynomials: positivity, admissible exponents and Jordan rigidity

Scott and Sokal asked whether every homogeneous polynomial with the half-plane property has a completely monotone negative power. We answer this question affirmatively and prove the Riesz-positivity conjecture of Michałek, Sturmfels, Uhler and Zwiernik, restated by Kozhasov, Michałek and Sturmfels: in $n$ variables, every exponent $α\ge4096n^2$ is admissible, independently of the degree and coefficients. For complete hyperbolic polynomials the Riesz density is strictly log-concave, with relative Gaussian error at most $512n^2/α$ and explicit curvature bounds. We characterize admissible exponents by a common spectral Dirichlet law, prove $n\le m+αm(m-1)$ with its equality case, and obtain the sharp degree-dependent gap $0<α<1/(2(m-1))$ whenever the degree-$m$ polynomial has a nonlinear irreducible factor. A nonnegative fourth-order defect of $-\log p$ vanishes at one point precisely for products of positive integer powers of Euclidean Jordan determinants. We classify the corresponding logarithmic Monge-Amp{è}re equation and answer the question of Etingof, Kazhdan and Polishchuk about polynomial multiplicative Legendre transforms within irreducible complete hyperbolic polynomials; the general question has counterexamples, the Clifford quartics of Kogiso and Sato. The classification extends to reducible polynomials under a boundary-visibility hypothesis, and asymptotic common-power formulas for the Riesz densities force exact Jordan formulas.

math.CO↗

Gap-free Differentially Private PCA for Gaussian Data

We give a gap-free $(ε,δ)$-differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data. The algorithm is based on a private variant of the power iteration method, and it is computationally efficient.

cs.DS↗

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.

cs.LG↗