Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,081 records · Page 60Linked to original sources

Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.

cs.AI↗

Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it

Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at https://github.com/OptimAI-Lab/HiddenStateUnlearning.

cs.LG↗

Block-Transitive $5$-$(v,k,2)$ Designs with $k$ divides $v$

The classification of block-transitive 5-designs remains an open problem. The additional parameter condition that $k$ divides $v$ is called the Camina-Gagen condition. In this paper, we investigate block-transitive simple $5$-$(v,k,2)$ designs satisfying the Camina-Gagen condition. Using the classification of finite $2$-homogeneous permutation groups, we consider the affine and almost simple cases separately. We prove that no such design admits a block-transitive automorphism group of affine type. For the almost simple case, up to isomorphism, there are exactly two possibilities: a $5$-$(12,6,2)$ design admitting PGL(2,11) as a block-transitive automorphism group and a $5$-$(24,8,2)$ design admitting PGL(2,23) as a block-transitive automorphism group.

math.CO↗

Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test

A commercial incentive need not enter the final ranking algorithm to affect a shopping assistant's recommendation: it may instead influence which preference question the assistant asks. We make this distinction experimentally observable in a deliberately small, synthetic setting. Each task has two products, three verified numerical attributes, a price limit, and a private fixed preference vector. An honest simulated user answers one pairwise question. A separate recommender receives the products and this answer but not the sponsorship assignment. We contrast a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor's advantage while omitting the rival's advantage. Across 40 held-out sponsorship-assignment cases (20 distinct catalog-preference contexts), the soft instruction changes no selections. The targeted instruction raises sponsored selection by 0.30 and reduces mean synthetic utility by 0.0547 relative to neutral questioning (95% context-bootstrap interval [-0.0828, -0.0291]) for one language-model recommender. A fixed Bayesian recommender shows a similar effect; a second model makes the same choices on all 120 frozen question-answer inputs. A terminal-answer consistency judge rates all 20 sampled targeted answers consistent, although five have synthetic regret above 0.05; a separate question-coverage dimension flags their one-sided elicitation. A robust partial-preference certificate remains valid under the stipulated synthetic utility but certifies only 16 of 40 targeted cases and is not better than asking a neutral question directly. These results establish neither typical behavior under advertising incentives nor effects on actual consumers.

cs.LG↗

CI-PINN: Causal Integral Physics-Informed Neural Network for Solving Evolution Equations

Physics-informed neural networks (PINNs) solve partial differential equations (PDEs) by incorporating governing physical laws into the training loss. For evolution equations, however, their conventional pointwise space--time representation does not explicitly encode temporal dependence, which can hinder accurate prediction. To mitigate this limitation, this work proposes a novel neural architecture termed a causal integral neural network (CinNet). The core module of CinNet is a Volterra-type causal integral term, which aggregates historical features to encode temporal dependence, thereby incorporating temporal causality at the architectural level rather than through training-level modifications as in many existing methods. Building on CinNet, we further develop a causal integral physics-informed neural network (CI-PINN) for solving evolution equations. Extensive numerical experiments on benchmark evolution equations demonstrate that the presented method outperforms various baseline PINN variants in terms of solution accuracy, with pronounced superiority under sparse-collocation scenarios. Additional empirical analyses show that CI-PINN exhibits low sensitivity to hyperparameter choices, while ablation studies confirm the effectiveness of the proposed network components.

math.NA↗

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.

cs.CV↗

Generating Edit-Inducing Questions for AI Research Manuscripts

We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.

cs.CL↗

RobotEQ 3.0: Towards Personalized Social Proactive Intelligence in Embodied Agents

Social Proactive Intelligence (SPI) is an emerging research area, aiming to shift embodied agents from reactive assistance toward proactively understanding human needs and executing socially desirable actions. Prior work has largely centered on the average user. However, human expectations are inherently diverse, and prior work overlooks individual nuances. To bridge this gap, we introduce RobotEQ 3.0, a benchmark for Personalized SPI. (Dataset) We first profile participants via a structured questionnaire covering factors that are correlated with human expectations of embodied agents, such as basic demographics and personality traits. Participants then select their preferred actions from a set of candidates. Unlike prior SPI benchmarks that focus on assessing behavioral appropriateness, our task centers on predicting the actions preferred by a specific user, thereby capturing human subjectivity. The resulting dataset establishes explicit links between individual traits and behavioral preferences. (Solution) We observe substantial inter-annotator variance, confirming that user preferences over actions are highly individualized. This motivates our exploration of Personalized SPI, in which user traits serve as additional inputs to predict individual preferences. Experimental results show that incorporating user traits can aid personalized prediction. This work aims to shift the research paradigm from developing agents suited for the average user to designing systems tailored to specific individuals.

cs.HC↗

SemPSG: A Semantic Channel-Aware Foundation Model for Polysomnography Analysis

Polysomnography (PSG) integrates multiple physiological signals to provide a comprehensive characterization of human sleep, yet its heterogeneous channel configurations across centers pose substantial challenges for transferable representation learning. Existing foundation models mainly focus on physiological modeling or temporal learning, while channel identity is often treated as a fixed structural index, overlooking the physiological semantics encoded by signal modality and reference configuration. To this end, we propose SemPSG, a Semantic channel-aware foundation model for heterogeneous PSG analysis. SemPSG explicitly represents the physiological semantics of channel identity and incorporates them into both signal representation learning and channel aggregation, enabling flexible modeling across diverse data configurations. Specifically, a semantic-conditioned time-series encoder captures signal-specific temporal dynamics and cross-signal interactions, while a multi-view image encoder extracts complementary time-frequency and morphological patterns from the same physiological recordings. We evaluate SemPSG on sleep and health-related tasks, including sleep staging, sleep-disorder breathing analysis, disease prediction, cognition and emotion recognition, and demographic estimation. Extensive experiments demonstrate consistent improvements over both general-purpose time series foundation models and PSG-specific foundation models, together with generalization across heterogeneous datasets across diverse channel configurations.

cs.LG↗

Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge

Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representations of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. Across standard knowledge-graph benchmarks, NSR achieves competitive accuracy without leading on every dataset, and has lower reported training times than several neural baselines. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.

cs.AI↗

Attosecond circular-dichroism spectroscopy of hole ring currents

Ultrafast ionization of atoms by circularly polarized few-cycle laser pulses generates a hole ring currents, offering a route for ultrafast manipulation of magnetism. Here we explore the subcycle formation of these currents whose circulation direction is determined by the driving-field helicity, whereas their magnitude is governed by the quantum coherence of the residual ion. We show that the correlated ion-photoelectron dynamics can be probed with attosecond transient-absorption circular dichroism. Using the Wigner-Eckart theorem, we derive a linear relation between state-resolved dichroic absorption and the orbital and spin angular-momentum projections of the hole, thereby extending the Thole-Carra magnetic CD sum rules from static X-ray spectroscopy to attosecond electron dynamics. The modulation depth of the dichroic signal provides a direct optical measure of the coherence injected by ionization, and its dependence on the pump-pulse duration reveals a competition between coherence build-up and ionization-induced dephasing. These findings demonstrate the potential of attosecond circular dichroism for probing ionic coherence and highlight the role of quantum coherence in ultrafast angular momentum exchange between light and matter.

physics.atom-ph↗

The classification of endpoint-transitive graphs with geodesic condition

We introduce endpoint k-path-transitive graphs, in which the automorphism group fixing two prescribed vertices pointwise acts transitively on the paths of length k joining them. We study the geodesic case, where k equals the distance between these vertices. We identify the union of these geodesics with the Hasse graph of a finite bounded graded poset and reduce it to proper blocks by complete cuts andmatching compression. Assume that the induced group K acts faithfully and 2-homogeneously on the first distance layer. We first prove that every nontrivial normal subgroup of K is transitive on every internal layer if and only if soc(K) is. Under this normal-basic condition, we classify the proper blocks in the affine case and the equal-width proper blocks with a 2-transitive internal layer in the almost-simple case. The nontrivial design interfaces are Paley or affine symplectic designs in the affine case, and projective designs, the 2-(11, 5, 2) design, the HigmanSims design, or their complements in the almost-simple case. The proper blocks have reduced rank at most four, whereas almost-simple proper blocks can have arbitrarily large reduced rank without the equal-width condition. We also give a stabilizer factorization criterion for assembling blocks while preserving endpoint-geodesic transitivity.

math.GR↗

Anomalous pressure-dependent viscosity of basaltic melts and its role in asthenosphere melt accumulation

The asthenosphere's mechanical weakness enables plate tectonics, but its origin is debated. Partial melt has been proposed to cause this softening, yet recent studies suggest that the measured viscosity minimum in basaltic melts, an essential control on melt mobility, is an experimental artifact. Using quantum mechanics-based, machine learning-accelerated molecular dynamics, we extend simulation timescales by more than a factor of 1000 and achieve percent-level precision. We show that basaltic melt exhibits a robust viscosity minimum (approximately 20% below 1-bar values) at approximately 3 GPa, driven by pressure-induced reorganization of aluminum coordination that facilitates shear relaxation while silicon-oxygen polyhedra remain structurally rigid. Our results reveal a depth-dependent rheological transition: melt mobility peaks below approximately 150 km, promoting efficient extraction, but declines sharply during ascent, causing melt to stagnate beneath the lithosphere. This mechanism provides a physical basis for the dual seismic signatures of a melt-depleted deep asthenosphere and a melt-enriched layer near the lithosphere-asthenosphere boundary.

physics.chem-ph↗

DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors

Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.

cs.AI↗

Semiclassical description of quadrupole-hexadecapole correlation in nuclear shape evolution

Background: Theoretical investigations of nuclear shapes with realistic effective interactions or mean field potentials have suggested a remarkable systematics in the hexadecapole shape evolution with varying particle number: Within each single-particle shell, from one spherical magic number to the next, a diamond type (beta_4>0) appears in the first half, and then turn into oblong type (beta_4<0) in the second half. Purpose: Nuclear deformation is essentially governed by the single-particle shell structure. In semiclassical periodic orbit theory (POT), level density is expressed as the sum over contributions from classical periodic orbits, and the origins of the gross shell structures can be understood by the contribution of one or a few shortest orbits. Using the POT, I examine how the hexadecapole degree of freedom affect the contribution of the orbits to the deformed shell structure through their bifurcations to explain the mechanism of above shape evolution. Methods: Ground state deformations are systematically investigated by the shell correction method, taking account of axially symmetric quadrupole and hexadecapole shape degrees of freedom. For simplicity, I employ two simplest mean field potentials to obtain the shell corrections: the oscillator and the cavity (infinite well) potentials. After confirming the the systematics in shape evolution under these simplest mean-field models, semiclassical analyses are made, focusing on the role of PO bifurcation. Results and conclusions: Systematics in the hexadecapole shape evolution in each single-particle shell is clearly explained by the bifurcations of two different types of periodic orbits in the cavity model, which suggest strong correlation between quadrupole and hexadecapole parameters. This also ensures the mechanism of nuclear prolate-shape predominance within this simple potential model.

nucl-th↗

Semantic Projection for Continual Self-Evolution of Language Agents

Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emph{Semantic-Scope Projected Evolution} (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.

cs.AI↗

Synchronized Quantum Devices in Lossy Channels

Due to their tight temporal coupling, entangled photon pairs produced by spontaneous parametric down-conversion offer a new route to improved network synchronization among quantum-enabled devices. In the emerging space-based quantum internet, such synchronization will be pivotal to operational success. Here, we propose and experimentally demonstrate a new polarization-assisted protocol that allows quantum-enabled devices to remain synchronized under the high loss conditions anticipated for satellite-to-ground quantum channels. Our experimental results show that polarization assistance extends reliable synchronization acquisition to higher channel losses and an increase of up to 40$\%$ in satellite-to-ground distance. Supporting Monte Carlo simulations show that this performance improvement extends over a wide range of photon-starved conditions. Collectively, these results not only provide a proof of concept for polarization-assisted network synchronization but also demonstrate its pragmatic importance in significantly extending the range under which quantum-enabled devices forming the backbone of the quantum internet can remain highly synchronized.

quant-ph↗

Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning

Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and shot contexts. Our Graded Physical Alignment score (GPA) awards credit for correct content and deducts points for incorrect or fabricated actions, making these errors explicit in the aggregate score. Building on this rubric, we propose Graded Margin Direct Preference Optimization (GM-DPO), which assigns stronger preference margins and greater training weight to more severe action errors. Across three VLM backbones, GM-DPO achieves the highest substantive-action and GPA scores among the evaluated preference objectives, improving GPA over DPO by 2.02-3.40 points. On Qwen3-8B, it reduces the weighted hallucination rate by 21.3% relative to DPO. These gains accompany sustained long-form output, improved shot structure, and competitive performance on three additional multimodal benchmarks.

cs.CV↗