Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion

Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize principle identifies a path that existing frameworks for identification, planning, and diagnosis do not make explicit: prior formulations treat the mapping from fault models to acceptable policies as a given, whereas LCPI makes it a function of the agent's own actions. We formalize this principle through Last-Chance Policy Identification (LCPI), where correctness is evaluated at the state the agent reaches rather than at the initial state. The Last Identifiable Margin (LIM) marks the feasibility boundary between distinguishing and homogenizing. For deterministic diagnostic graphs we provide the Exact-LIM recursion; for noisy finite-horizon recovery we propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint. Across incident recovery on abstract microservice topologies and latent-damage navigation in MiniGrid, RBCP improves risk-feasible recovery while satisfying the failure budget. A sham control--cost-matched actions that preserve model incompatibility--eliminates the gain entirely, confirming that the benefit comes from changing which policies are acceptable for which models, not from extra search or additional budget.

cs.AI↗

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.

cs.AI↗

Broadband reverberation mapping of quasars at cosmic noon

Broadband reverberation mapping provides a powerful tool to investigate the structure of the accretion flow around supermassive black holes. The ultraviolet region of the quasar spectrum is of particular interest, as it enables the disentanglement of the relative contributions from the accretion disc and the broad-line region to the observed variability of quasars in the local Universe. Observations of quasars at cosmic noon ($z\sim 1-2$) offer the opportunity to probe the UV emission across a diverse population, spanning a broad range of black hole masses and luminosities. Here we re-analysed the 36 multi-year quasar lightcurves from the Dark Energy Survey (DES) between $z\sim0.7-2$ to retrieve interband delays and analyse their spectral energy distribution. Using PyROA, we were able to fit all multiple years of observations and the four optical bands simultaneously, significantly improving the lag detection as this method is able to interpolate over the seasonal gaps. We find robust detections of interband delays in 31/36 quasars, suggesting high-Eddington ratios for a large fraction of objects. Their spectral energy distribution is consistent with a geometrically thin disk $f_ν\proptoλ^{-1/3}$, with an ensemble slope of $-0.28\pm0.02$. The full sample of quasars do not show excess around the Balmer jump region, suggesting the impact of the broad-line region to the reverberation signals in high-Eddington sources could be small as documented in some AGN in the local Universe. The prospect of LSST to observe thousands of quasars in this redshift range should motivate the detailed study of the UV-rest frame of high-Eddington quasars to account for the BLR contribution and properly infer disc sizes.

astro-ph.GA↗

Aggregation of Time Preferences and Comparability of Intergenerational Utilities

Experimental evidence reveals that, in aggregating time preferences, a decision maker exhibits mild aversion to inequality across individuals' evaluations of a stream. Neither the utilitarian nor the maxmin aggregation rule can account for this finding. Building on principles concerning the comparability of intergenerational utilities and intergenerational fairness, this paper axiomatizes a family of aggregation rules that allow the decision maker to display varying degrees of inequality aversion. These rules fit well with experimental data.

econ.TH↗

Efficient Transmit Waveform Design for MIMO-OFDM DFRC Systems with 1-Bit DACS

This paper studies efficient waveform design for MIMO-OFDM dual-functional radar-communication (DFRC) systems with 1-bit digital-to-analog converters (DACs), which remains underexplored in the literature. We formulate the DFRC waveform design problem by maximizing the trace of the target angular Fisher information matrix (FIM), subject to symbol-error probability constraints realized via constructive interference (CI) and a time-domain 1-bit transmit alphabet constraint. The resulting problem is a large-scale nonlinear integer programming problem. We develop an inexact alternating direction method of multipliers (ADMM) algorithm with alternating time- and frequency-domain updates. The proposed scheme enables subcarrier-wise optimization in the frequency domain and closed-form primal and dual updates in the time domain. In particular, the CI constraints are deliberately retained in both the time- and frequency-domain subproblems to promote the feasibility of intermediate iterates and improve the convergence behavior of ADMM. Simulation results demonstrate promising DFRC performance.

eess.SP↗

EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents

Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5--41.0% fewer skills, improves skill retrieval by 36.3--38.7% and measured edit utility by 51.8--60.0%, and reduces deployment-time inference tokens by 9.1--14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.

cs.AI↗

Fabrication framework for three-dimensional colloidal particles with decoupled geometry and material composition

Shape-programmable particles offer significant opportunities for microrobotic systems at the individual level and for hierarchical materials with emergent functionalities arising from collective particle behavior. However, fabricating shape-changing stimuli-responsive particles with complex three-dimensional geometries at colloidal length scales remains a major challenge. Here, we introduce a general fabrication framework that decouples particle geometry from material composition to produce free-standing three-dimensional colloidal particles with complex architectures. Our approach combines soft lithography, swelling-assisted extraction, and sacrificial adhesive transfer to fabricate particles with high geometric fidelity. We establish a predictive framework that defines the accessible design space for increasingly complex particle geometries. We further extend our framework to fabricate high-aspect-ratio pillar arrays with intricate three-dimensional architectures at colloidal length scales. To demonstrate material versatility, we fabricate particles from both liquid crystal elastomers and hydrogels. We show reversible shape-changing behavior of liquid crystal elastomer particles with cylindrical and chiral shapes under thermal and optical stimuli. In suspension, these particles display collective optical dynamic behavior arising from coupling between changes in the programmed liquid crystal organization within the particles and stimulus-induced geometric reconfiguration of the particles. Collectively, this work establishes a versatile platform for geometry-programmable colloidal particles with emergent collective functionalities, providing a route toward materials and fluids with dynamically programmable properties.

cond-mat.mtrl-sci↗

Generalizable Lifelong Model Editing via Preference Optimization

Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model's general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.

cs.AI↗

Study of LAGP coating on polyethylene separator for polysulfide suppression in thin Li-S batteries

The use of a thin solid electrolyte layer on polyethylene (PE) separators has been explored as a potential strategy to mitigate the polysulfide shuttle effect in lithium-sulfur (Li-S) batteries, a phenomenon responsible for significantly reducing their lifespan. While some solid electrolytes, such as LAGP , have demonstrated the ability to effectively suppress the shuttle effect, their brittle ceramic nature limits their application, particularly in flexible electronics. Coating PE separators with LAGP nanofilms offers a promising solution by combining the mechanical flexibility and thinness of polymer separators with the electrochemical advantages of LAGP, due to its NASICON-like structure and high ionic conductivity. In this study, PE separators were coated with LAGP nanofilms using the ion sputtering technique and further modified with ferroic elements such as Ni and Co. The surface morphology of the coated separators was analyzed using scanning electron microscopy (SEM), and staircase voltammetry was employed to evaluate the electrochemical performance and modifications induced by the coatings. The results demonstrate that the use of PE separators with a thin LAGP layer significantly influences polysulfide diffusion and suppresses the shuttle effect, thereby reducing unwanted processes in Li-S batteries.

physics.app-ph↗

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

cs.LG↗

Degeneracy-Orthogonal Geometric Constraints for LiDAR SLAM

Autonomous robot navigation relies on simultaneous localization and mapping (SLAM) to estimate motion and maintain an accurate pose within an environment. However, in axially uniform corridors such as long tunnels and pipelines, LiDAR odometry is fundamentally limited by unconstrained drift along the feature-weak travel direction. This structural degeneracy cannot be resolved by local scan matching alone. To address this challenge, we propose the Degeneracy-orthogonal Contour Offset Descriptor (DeCOD), a structure-aligned geometric descriptor for cross-sectional landmarks. Cross-sectional boundaries, such as pipe joints and structural rings, provide metric constraints along this degenerate axis, but distinguishing individual landmarks requires capturing subtle surface variations across nearly identical profiles. The descriptor parameterizes signed normal deviation from estimated boundary contours, and matching explicitly resolves heading ambiguity and decouples first-order contour errors by distortion estimation. Matched landmarks yield geometric factors that enforce agreement in cross-section position and corridor axis alignment during pose-graph optimization, correcting longitudinal drift while leaving rotation about the common axis unconstrained. On a public benchmark and in field experiments, DeCOD achieves robust landmark retrieval over standard 3D descriptors and successfully stabilizes trajectories across different odometry frontends, reliably constraining longitudinal drift under geometric degeneracy.

cs.RO↗

Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT

Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.

eess.AS↗

Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing

Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.

cs.CV↗

NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation

One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/jaiwei804/NesTok.

cs.CV↗

FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially reduces the computational cost. During training, it further adopts velocity consistency regularization and continuous trajectory learning, which improve restoration quality. Extensive experiments show that FastVR is more efficient than the evaluated diffusion baselines while achieving state-of-the-art performance on synthetic and real-world benchmarks. We hope that this work supports further progress in the community.

cs.CV↗

Synthesis and modification of thin NaSICON solid electrolytes using ion beams

Solid electrolytes (SEs) for sodium-based superionic conductors (NaSICON) were first introduced in 1976 and quickly recognized for their excellent ionic conductivity. While considerable effort has been made to develop thin electrolytes for all-solid-state batteries (ASSBs), only a few sodium-based SEs have been successfully fabricated as thin films. These thin films are particularly desirable for their reduced electrical resistance, which typically increases with the thickness of the SE. By reducing the thickness of the SEs to the nanometer scale, their ionic conductivity can be significantly enhanced. In this study, the NASICON composite was initially prepared in the form of pellets using the mixed oxide technique with a planetary ball mill and synthesized by the solid-state method at 1300 °C. The resulting pellets were used as sputtering targets in a low-energy ion facility to prepare continuous NASICON nanofilms. To explore the effect of ion implantation on the electrical properties of NASICON, the prepared films were bombarded with Ni ions at 1.1 MeV and varying fluences, using the Tandetron accelerator at the CANAM infrastructure (NPI Řež). The electrical properties of both the synthesized and implanted films were analyzed through electrochemical impedance spectroscopy (EIS). The results, describing the impact of irradiation on NASICON's properties, are presented here.

cond-mat.mtrl-sci↗

Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.

cs.CV↗

QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.

cs.LG↗