Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 901 records · Page 50Linked to original sources

Hybrid Work and the Restructuring of Urban Mobility in U.S. Cities

Entering the post-pandemic era, cities navigate a new normal shaped by hybrid work and space-time flexibility, but existing evidence is spatially scattered and temporally limited. Here we develop an analytical approach that examines this reorganization through three dimensions: remote work adoption via sector composition, daily travel behavior via trip-level metrics, and functional urban form via mobility-derived measures. To capture the interplay of urban dynamics, we synthesize three complementary spatial metrics. In particular, coupling overall trip centrality with work-trip concentration differentiates whether employment and daily activity organize around the same centers. Applied to population-scale data across 15 U.S. metropolitan areas spanning 2019-2024, our approach shows that daily car travel increased despite elevated remote work, a pattern consistent across cities with distinct spatial structures. Decomposition and regression identify trip frequency as the primary contributor to VKT growth, moderated by compact urban form and concentrated employment. The approach provides a replicable framework for comparative study of post-pandemic urban mobility.

physics.soc-ph↗

Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images

Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-H&E pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from H&E images.

cs.CV↗

Positive Cubature Compression and Optimal Hyperinterpolation Stability on a Conic Surface

We study the computational cost and stability of positive cubature and hyperinterpolation on a truncated conic surface with a Jacobi-type weight. A classical quadratic disk-to-cone map is used in a basis-free form as a measure-preserving $\mathbb Z_2$ quotient. It identifies the full degree-$m$ conic trace space with the even disk polynomials of degree at most $2m$ and converts positive degree-$2n$ cone cubature into centrally symmetric positive degree-$(4n+1)$ disk cubature, and conversely. This yields an exact transfer of Möller's lower bound and of the node excess above it, so that near-minimal disk formulas produce compressed non-product cone rules. The same quotient transfers reproducing kernels and hyperinterpolation operators. Every positive degree-$2n$ cone rule gives an exact $L^2$ sampling isometry on $Π_n(V)$; in particular, the weighted sampling matrix has condition number one, independently of the number and geometry of the nodes. For the unweighted radial case $γ=0$, if $Λ_n^V$ denotes the $C(V)\to C(V)$ Lebesgue constant of degree-$n$ hyperinterpolation, then every positive degree-$2n$ cone cubature rule satisfies the rule-independent sharp law $$c n \le Λ_n^V \le C n.$$ More generally, the upper bound $Λ_{n,γ}^V \le C_γn^{γ+1}$ holds when $γ$ is a nonnegative half-integer. Thus node compression preserves exact $L^2$ conditioning while the associated $L^\infty$ stability has the optimal universal order. Low-degree near-minimal rules and higher-degree optimized disk schemes illustrate the reduction in sampling cost.

math.CA↗

Transverse momentum as the counter-diabatic generator in bent waveguide couplers

Bending the axis of a mode-evolution coupler suppresses the nonadiabatic coupling between its supermodes, and bent couplers of this kind were recently shown to realize the counter-diabatic (CD) protocol. Here we identify the operator responsible. Rigid lateral displacement of a waveguide structure is generated by the transverse momentum $\hat p_x$, whose diagonal matrix elements vanish in a real supermode basis and whose off-diagonal element is purely imaginary; in a two-mode system such an operator is proportional to $\sy$. The CD term of a two-waveguide coupler is itself proportional to $\sy$, while the detuning and the coupling, the two parameters set by the waveguide widths and spacing, lie in the $σ_z$--$σ_x$ plane and cannot produce it. The transverse momentum is therefore the CD generator of the bent coupler, by necessity rather than by design. Equating the term it supplies to the nonadiabatic coupling gives the axis slope in closed form as a ratio of two matrix elements of the unperturbed supermodes at a single cross section, which a commutator identity recasts in coupled-mode variables as $\dot{\xo}=\dotθ/(γ\dbeta^{2})$. The expression agrees with the numerical offset search to $0.3\%$ and requires no search. Beam propagation simulations confirm the design and compare it, for the first time, with the CD protocol realized by unitary transformation in a straight coupler built from the same reference structure. The two devices reach $0.994$ and $1.000$ supermode fidelity where the untreated coupler reaches $0.886$, and agree in bandwidth and fabrication tolerance, while their internal trajectories differ by exactly the frame transformation that relates them.

quant-ph↗

Temporal-Aware Fusion for Robust Outdoor LiDAR Localization

LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leaving the potential of spatio-temporal consistency across scans not fully explored. In this paper, we propose a Temporal-aware Localization framework (TempLoc) designed to enhance the robustness of outdoor localization by effectively modeling sequential consistency. Specifically, a Global Coordinate Estimation module is first introduced to predict point-wise global coordinates and associated uncertainties for each LiDAR scan. A Prior Coordinate Generation module is then presented to estimate inter-frame point correspondences by the attention mechanism. Lastly, an Uncertainty-Guided Coordinate Fusion module is deployed to integrate both predictions of point correspondence in an end-to-end fashion, yielding a more temporally consistent and accurate global 6-DoF pose. Experimental results on the NCLT and Oxford RobotCar benchmarks show that our TempLoc outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of temporal-aware correspondence modeling in LiDAR relocalization.

cs.CV↗

RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance

Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictions. However, existing branch-local reuse criteria do not explicitly account for how cache errors combine under CFG or how local perturbations affect the final output. We identify two misalignments in cache control: a branch-guided mismatch, where guided error depends on both the magnitudes and alignment of branch errors, and a local-final mismatch, where the downstream impact of a local error varies across timesteps. We propose RA-CFGCache, a Risk-Aligned Caching framework under CFG that incorporates both factors while keeping the sampling schedule and guidance rule fixed. CFG-aware Guided-Risk Composition combines existing branch-wise proxies using CFG coefficients and offline-calibrated cross-branch alignment. Propagation-Aware Rescaling further weights the resulting guided-risk estimate with a timestep-dependent propagation prior calibrated from isolated reuse perturbations. An online threshold controller then determines when to jointly refresh or reuse both branches. Experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B demonstrate improved efficiency--fidelity trade-offs over evaluated training-free caching baselines. Moreover, RA-CFGCache is compatible with diverse base proxy families, including TeaCache-, DiCache-, and MagCache-style estimators, and consistently improves fidelity at nearly unchanged latency. Code is available at https://github.com/yiming-l21/RA-CFGCache.git.

cs.CV↗

The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.

cs.AI↗

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader's input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never produced. We present MemFold, which optimizes a fixed-budget soft memory by the behavior it supports. A query-conditioned textual memory is compressed into K continuous vectors that form the reader's memory interface, and the reader is then trained on its own rollouts under two complementary signals: group-relative rewards for task outcomes, and confidence-gated on-policy distillation in which a frozen textual-memory teacher re-scores the student's sampled tokens under the textual memory. The teacher is never sampled from, so supervision stays on the student's current distribution and adds no autoregressive decoding; at inference it is removed entirely. Across three Qwen backbones, MemFold attains the highest accuracy we measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length, and transfers to PrefEval and LongMemEval without target-domain training. Ablations attribute most of the task gain to the reward term and a smaller additional gain to the teacher signal, and memory interventions show that the reader depends on the instance-specific content of its soft memory.

cs.CL↗

Merlin Plus: A Large-Scale, Multi-Cancer, Image-Mask-Report Dataset

Multi-cancer segmentation in computed tomography (CT) is fundamentally limited by the scarcity of tumor masks across different organs. We present Merlin Plus, the first large-scale CT dataset with radiologist-created tumor masks across 9 organs. Merlin Plus extends the Merlin dataset by adding 1,153 per-voxel tumor masks and longitudinal metadata. To create these tumor masks, we developed a report-based active-learning framework in which radiology reports identify tumor cases for annotation and support training of a tumor segmentation model. The model generates initial masks, which radiologists review and correct to produce the final masks, reducing annotation burden while maintaining high-quality annotations. Besides tumor masks, the longitudinal metadata in Merlin Plus enables temporal modeling of cancer progression. By directly addressing the major bottleneck of limited multi-cancer segmentation masks, Merlin Plus supports scalable multi-organ cancer detection, segmentation, and longitudinal analysis in CT. Dataset is available at: https://github.com/MrGiovanni/MerlinPlus

cs.CV↗

Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference

Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.

cs.AI↗

World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving

Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.

cs.RO↗

Acoustoelectrically enhanced acousto-optic modulation in an integrated silicon nitride and thin film lithium niobate platform

Acoustoelectric interactions in piezoelectric-semiconductor heterostructures allow the propagation characteristics of microwave frequency phonons in piezoelectric media to be controlled and radically enhanced, providing electrically controllable phonon gain, large velocity tuning, isolation, and circulation, as well as extremely large electron-mediated phononic nonlinearities. Here, for the first time, we create such a piezoelectric-semiconductor heterostructure with lithium-niobate-on-insulator and InGaAs that also supports guided optical modes through the addition of a silicon nitride waveguide and modification of the acoustic materials to provide an optical lower cladding. We use this new architecture to demonstrate acoustoelectrically enhanced acousto-optic modulation, where 1 GHz phonons are piezoelectrically generated and acoustoelectrically amplified on-chip by up to 60 dB before impinging on the optical waveguide, providing pure phase modulation with a $V_πL$ figure-of-merit of 0.077 V-cm while only consuming 3.77 mW of DC electrical power to provide the amplification. We then consider future applications enabled by these functionalities and describe a novel tunable optical delay and an optoelectronic oscillator (OEO) analog---an acoustoelectrically enhanced opto-acoustic oscillator (AE-OAO). We show that using Brillouin optomechanical transduction and acoustoelectrically lossless acoustic time delay, the AE-OAO could replace kilometers of optical fiber delay used in OEOs but on a single, centimeter-scale chip.

physics.optics↗

DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing

Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects the generation dynamics that lead to hallucinated outputs. In this paper, we analyze the relationship between TF-based editing and AR generation behavior and find that TF-based editing alone may be insufficient to capture both decoding dynamics and multimodal interactions associated with hallucinations. To address this limitation, we propose DARE (Dual-path Auto-Regressive-aware Editing), a hybrid hallucination editing framework that integrates two complementary contrast pathways: textual contrasts and image contrasts, together with autoregressive-aware representation signals. Specifically, DARE constructs hallucination editing directions from (1) TF-based textual contrasts, (2) AR-aware representation transitions during decoding, and (3) controlled visual differences between paired images. Extensive experiments on multiple LVLM hallucination benchmarks demonstrate that DARE consistently reduces object hallucinations while preserving multimodal perception capability and inference efficiency. Our implementation code is available at https://github.com/KU-VGI/DARE.

cs.CV↗

Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization

Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.

eess.AS↗

Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time

Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at https://github.com/KU-VGI/Online-VIL.

cs.CV↗

Christophersen's problem for monomial algebras

Christophersen's problem predicts that the connected component of the automorphism group of a finite-dimensional local algebra $A$ of dimension $\ell$ over an algebraically closed field of characteristic zero has dimension at least $\ell-1$, with equality if and only if $A$ is isomorphic to $\mathbf{k}[t]/(t^{\ell})$. We settle this problem in the class of monomial algebras. Using the combinatorial structure of the irredundant irreducible decomposition of a monomial ideal, we prove the predicted inequality and characterize the equality case. We further obtain stronger lower bounds depending on whether the defining monomial ideal is reducible or irreducible, and we determine all monomial algebras attaining each of these bounds.

math.AC↗

Relative Suspension and Higher Delooping for Support $τ$-Tilting Modules

We construct relative suspension for the exact category generated by a support $τ$-tilting module and characterize higher delooping by the units of the resulting adjunction, which are the generalizations of Gélinas's results in Adv. Math.(2022). We prove that comparison with ordinary higher delooping requires a shift of one in the index. More precisely, for every $k,d\ge1$, there is a module whose ordinary $k$-delooping level is zero and whose relative level is $d$. We also prove that tensor induction preserves the relative constructions and higher delooping levels on induced modules. As an application, we study cyclic Nakayama algebras with arbitrary local coefficients and automorphism twists. Their left and right little and big finitistic dimensions and ordinary and derived delooping levels all equal the finitistic dimension of the underlying Nakayama algebra.

math.RT↗

Continuously Updating GMM in Linear IV Models: A Polynomial Approach

This paper characterizes the global minimum of the continuously updating generalized method of moments (CU-GMM) objective in linear instrumental variables models. We allow optimal weighting matrices under heteroskedasticity, autocorrelation, or clustering. We show that the objective is a ratio of polynomials. For one endogenous regressor, stationary points of CU-GMM objective function are real eigenvalues of a companion matrix. Comparing their objective values with the value at infinity gives the global minimum, extending the classical eigenvector approach to limited information maximum likelihood. With multiple endogenous regressors, algebraic elimination and checks for real solutions identify the minimum among finitely many candidate objective values, including boundary values. Galois theory rules out general formulas by radicals even with one endogenous regressor and two instruments, while numerical root finding remains possible. Finding the global minimum allows us to compute overidentification and likelihood ratio tests, including the conditional likelihood ratio (CLR) test.

econ.EM↗