Searcharxiv⌕ Search

arXiv subjects

Hanlin Wu

Publications and source records attributed to Hanlin Wu.

At least 37 records · Page 2Linked to original sources

When a Man Says He Is Pregnant: Event-related Potential Evidence for a Rational Account of Speaker-contextualized Language Comprehension

Spoken language is often, if not always, understood in a context formed by the identity of the speaker. For example, we can easily make sense of an utterance such as "I'm going to have a manicure this weekend" or "The first time I got pregnant I had a hard time" when spoken by a woman, but it would be harder to understand when it is spoken by a man. Previous ERP studies have shown mixed results regarding the neurophysiological responses to such speaker-content mismatches, with some reporting an N400 effect and others a P600 effect. In an EEG experiment involving 64 participants, we used social and biological mismatches as test cases to demonstrate how these distinct ERP patterns reflect different aspects of rational inference. We showed that when the mismatch involves social stereotypes (e.g., men getting a manicure), listeners can arrive at a "literal" interpretation by integrating the content with their social knowledge, though this integration requires additional effort due to stereotype violations-resulting in an N400 effect. In contrast, when the mismatch involves biological knowledge (e.g., men getting pregnant), a "literal" interpretation becomes highly implausible or impossible, leading listeners to treat the input as potentially containing errors and engage in correction processes-resulting in a P600 effect. Supporting this rational inference framework, we found that the social N400 effect decreased as a function of the listener's personality trait of openness (as more open-minded individuals maintain more flexible social expectations), while the biological P600 effect remained robust (as biological constraints are recognized regardless of individual personalities). Our findings help to reconcile empirical inconsistencies and reveal how rational inference shapes speaker-contextualized language comprehension.

q-bio.NC↗

A Synthetic Benchmark for Collaborative 3D Semantic Occupancy Prediction in V2X-Enabled Autonomous Driving

3D semantic occupancy prediction is an emerging perception paradigm in autonomous driving, providing a voxel-level representation of both geometric details and semantic categories. However, its effectiveness is inherently constrained in single-vehicle setups by occlusions, restricted sensor range, and narrow viewpoints. To address these limitations, collaborative perception enables the exchange of complementary information, thereby enhancing the completeness and accuracy of predictions. Despite its potential, research on collaborative 3D semantic occupancy prediction is hindered by the lack of dedicated datasets. To bridge this gap, we design a high-resolution semantic voxel sensor in CARLA to produce dense and comprehensive annotations. We further develop a baseline model that performs inter-agent feature fusion via spatial alignment and attention aggregation. In addition, we establish benchmarks with varying prediction ranges designed to systematically assess the impact of spatial extent on collaborative prediction. Experimental results demonstrate the superior performance of our baseline, with increasing gains observed as range expands. Our code is available at https://github.com/tlab-wide/Co3SOP}{https://github.com/tlab-wide/Co3SOP.

cs.CV↗

Probabilistic adaptation of language comprehension for individual speakers: evidence from neural oscillations

Listeners adapt language comprehension based on their mental representations of speakers, but how these representations are updated remains unclear. We investigated whether listeners probabilistically adapt comprehension based on the frequency of speakers making stereotype-incongruent statements. In two EEG experiments, participants heard speakers make stereotype-congruent or incongruent statements, with incongruency base rate manipulated. In Experiment 1, stereotype-incongruent statements decreased high-beta (21-30 Hz) and theta (4-6 Hz) oscillatory power in the low base rate condition but increased it in the high base rate condition. The theta effect varied with listeners' openness trait: less open-minded participants tended to show theta increases to stereotype incongruencies, while more open-minded participants tended to show theta decreases. In Experiment 2, we dissociated incongruency base rate from the target speaker by manipulating it using a non-target speaker and found that only the high-beta effect persisted. Our findings reveal two potential mechanisms: a speaker-general mechanism (indicated by high-beta oscillations) that adjusts overall expectations about hearing statements that violate social stereotypes, and a speaker-specific mechanism (indicated by theta oscillations) that updates a more detailed mental model specifically about an individual speaker. These findings provide evidence for how language processing interacts with social cognition.

q-bio.NC↗

Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment

Voice-based AI development faces unique challenges in processing both linguistic and paralinguistic information. This study compares how large audio-language models (LALMs) and humans integrate speaker characteristics during speech comprehension, asking whether LALMs process speaker-contextualized language in ways that parallel human cognitive mechanisms. We compared two LALMs' (Qwen2-Audio and Ultravox 0.5) processing patterns with human EEG responses. Using surprisal and entropy metrics from the models, we analyzed their sensitivity to speaker-content incongruency across social stereotype violations (e.g., a man claiming to regularly get manicures) and biological knowledge violations (e.g., a man claiming to be pregnant). Results revealed that Qwen2-Audio exhibited increased surprisal for speaker-incongruent content and its surprisal values significantly predicted human N400 responses, while Ultravox 0.5 showed limited sensitivity to speaker characteristics. Importantly, neither model replicated the human-like processing distinction between social violations (eliciting N400 effects) and biological violations (eliciting P600 effects). These findings reveal both the potential and limitations of current LALMs in processing speaker-contextualized language, and suggest differences in social-linguistic processing mechanisms between humans and LALMs.

cs.CL↗

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

Accurate interpretation of land-cover changes in multi-temporal satellite imagery is critical for real-world scenarios. However, existing methods typically provide only one-shot change masks or static captions, limiting their ability to support interactive, query-driven analysis. In this work, we introduce remote sensing image change analysis (RSICA) as a new paradigm that combines the strengths of change detection and visual question answering to enable multi-turn, instruction-guided exploration of changes in bi-temporal remote sensing images. To support this task, we construct ChangeChat-105k, a large-scale instruction-following dataset, generated through a hybrid rule-based and GPT-assisted process, covering six interaction types: change captioning, classification, quantification, localization, open-ended question answering, and multi-turn dialogues. Building on this dataset, we propose DeltaVLM, an end-to-end architecture tailored for interactive RSICA. DeltaVLM features three innovations: (1) a fine-tuned bi-temporal vision encoder to capture temporal differences; (2) a visual difference perception module with a cross-semantic relation measuring (CSRM) mechanism to interpret changes; and (3) an instruction-guided Q-former to effectively extract query-relevant difference information from visual changes, aligning them with textual instructions. We train DeltaVLM on ChangeChat-105k using a frozen large language model, adapting only the vision and alignment modules to optimize efficiency. Extensive experiments and ablation studies demonstrate that DeltaVLM achieves state-of-the-art performance on both single-turn captioning and multi-turn interactive change analysis, outperforming existing multimodal large language models and remote sensing vision-language models. Code, dataset and pre-trained weights are available at https://github.com/hanlinwu/DeltaVLM.

cs.CV↗

Single-Step Latent Consistency Model for Remote Sensing Image Super-Resolution

Recent advancements in diffusion models (DMs) have greatly advanced remote sensing image super-resolution (RSISR). However, their iterative sampling processes often result in slow inference speeds, limiting their application in real-time tasks. To address this challenge, we propose the latent consistency model for super-resolution (LCMSR), a novel single-step diffusion approach designed to enhance both efficiency and visual quality in RSISR tasks. Our proposal is structured into two distinct stages. In the first stage, we pretrain a residual autoencoder to encode the differential information between high-resolution (HR) and low-resolution (LR) images, transitioning the diffusion process into a latent space to reduce computational costs. The second stage focuses on consistency diffusion learning, which aims to learn the distribution of residual encodings in the latent space, conditioned on LR images. The consistency constraint enforces that predictions at any two timesteps along the reverse diffusion trajectory remain consistent, enabling direct mapping from noise to data. As a result, the proposed LCMSR reduces the iterative steps of traditional diffusion models from 50-1000 or more to just a single step, significantly improving efficiency. Experimental results demonstrate that LCMSR effectively balances efficiency and performance, achieving inference times comparable to non-diffusion models while maintaining high-quality output.

eess.IV↗

A Periodic Bayesian Flow for Material Generation

Generative modeling of crystal data distribution is an important yet challenging task due to the unique periodic physical symmetry of crystals. Diffusion-based methods have shown early promise in modeling crystal distribution. More recently, Bayesian Flow Networks were introduced to aggregate noisy latent variables, resulting in a variance-reduced parameter space that has been shown to be advantageous for modeling Euclidean data distributions with structural constraints (Song et al., 2023). Inspired by this, we seek to unlock its potential for modeling variables located in non-Euclidean manifolds e.g. those within crystal structures, by overcoming challenging theoretical issues. We introduce CrysBFN, a novel crystal generation method by proposing a periodic Bayesian flow, which essentially differs from the original Gaussian-based BFN by exhibiting non-monotonic entropy dynamics. To successfully realize the concept of periodic Bayesian flow, CrysBFN integrates a new entropy conditioning mechanism and empirically demonstrates its significance compared to time-conditioning. Extensive experiments over both crystal ab initio generation and crystal structure prediction tasks demonstrate the superiority of CrysBFN, which consistently achieves new state-of-the-art on all benchmarks. Surprisingly, we found that CrysBFN enjoys a significant improvement in sampling efficiency, e.g., ~100x speedup 10 v.s. 2000 steps network forwards) compared with previous diffusion-based methods on MP-20 dataset. Code is available at https://github.com/wu-han-lin/CrysBFN.

cs.LG↗

Latent Diffusion, Implicit Amplification: Efficient Continuous-Scale Super-Resolution for Remote Sensing Images

Recent advancements in diffusion models have significantly improved performance in super-resolution (SR) tasks. However, previous research often overlooks the fundamental differences between SR and general image generation. General image generation involves creating images from scratch, while SR focuses specifically on enhancing existing low-resolution (LR) images by adding typically missing high-frequency details. This oversight not only increases the training difficulty but also limits their inference efficiency. Furthermore, previous diffusion-based SR methods are typically trained and inferred at fixed integer scale factors, lacking flexibility to meet the needs of up-sampling with non-integer scale factors. To address these issues, this paper proposes an efficient and elastic diffusion-based SR model (E$^2$DiffSR), specially designed for continuous-scale SR in remote sensing imagery. E$^2$DiffSR employs a two-stage latent diffusion paradigm. During the first stage, an autoencoder is trained to capture the differential priors between high-resolution (HR) and LR images. The encoder intentionally ignores the existing LR content to alleviate the encoding burden, while the decoder introduces an SR branch equipped with a continuous scale upsampling module to accomplish the reconstruction under the guidance of the differential prior. In the second stage, a conditional diffusion model is learned within the latent space to predict the true differential prior encoding. Experimental results demonstrate that E$^2$DiffSR achieves superior objective metrics and visual quality compared to the state-of-the-art SR methods. Additionally, it reduces the inference time of diffusion-based SR methods to a level comparable to that of non-diffusion methods.

eess.IV↗

ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning

Remote sensing (RS) change analysis is vital for monitoring Earth's dynamic processes by detecting alterations in images over time. Traditional change detection excels at identifying pixel-level changes but lacks the ability to contextualize these alterations. While recent advancements in change captioning offer natural language descriptions of changes, they do not support interactive, user-specific queries. To address these limitations, we introduce ChangeChat, the first bitemporal vision-language model (VLM) designed specifically for RS change analysis. ChangeChat utilizes multimodal instruction tuning, allowing it to handle complex queries such as change captioning, category-specific quantification, and change localization. To enhance the model's performance, we developed the ChangeChat-87k dataset, which was generated using a combination of rule-based methods and GPT-assisted techniques. Experiments show that ChangeChat offers a comprehensive, interactive solution for RS change analysis, achieving performance comparable to or even better than state-of-the-art (SOTA) methods on specific tasks, and significantly surpassing the latest general-domain model, GPT-4. Code and pre-trained weights are available at https://github.com/hanlinwu/ChangeChat.

cs.CV↗

Charge order induced Dirac pockets in the nonsymmorphic crystal TaTe$_4$

The interplay between charge order (CO) and nontrivial band topology has spurred tremendous interest in understanding topological excitations beyond the single-particle description. In a quasi-one-dimensional nonsymmorphic crystal TaTe$_4$, the (2a$\times$2b$\times$3c) charge ordered ground state drives the system into a space group where the symmetry indicator features the emergence of Dirac fermions and unconventional double Dirac fermions. Using angle-resolved photoemission spectroscopy and first-principles calculations, we provide evidence of the CO induced Dirac fermion-related bands near the Fermi level. Furthermore, the band folding at the Fermi level is compatible with the new periodicity dictated by the CO, indicating that the electrons near the Fermi level follow the crystalline symmetries needed to host double Dirac fermions in this system.

cond-mat.str-el↗

Tailoring Physical Properties of Crystals through Synthetic Temperature Control: A Case Study for new Polymorphic NbFeTe2 phases

Growth parameters play a significant role in the crystal quality and physical properties of layered materials. Here we present a case study on a van der Waals magnetic NbFeTe2 material. Two different types of polymorphic NbFeTe2 phases, synthesized at different temperatures, display significantly different behaviors in crystal symmetry, electronic structure, electrical transport, and magnetism. While the phase synthesized at low temperature showing behavior consistent with previous reports, the new phase synthesized at high temperature, has completely different physical properties, such as metallic resistivity, long-range ferromagnetic order, anomalous Hall effect, negative magnetoresistance, and distinct electronic structures. Neutron diffraction reveals out-of-plane ferromagnetism below 70K, consistent with the electrical transport and magnetic susceptibility studies. Our work suggests that simply tuning synthetic parameters in a controlled manner could be an effective route to alter the physical properties of existing materials potentially unlocking new states of matter, or even discovering new materials.

cond-mat.str-el↗

Multiple intercalation stages and universal Tc enhancement through polar organic species in electron-doped 1T-SnSe2

In this work, we report multiple intercalation stages and universal Tc enhancement of superconductivity in 1T-SnSe2 through Li and organic molecules cointercalation. We observe significantly increased lattice parameters up to 40 Å and dramatically enlarged interlayer distance up to ~11Å in Li and N,N-dimethylformamide (DMF) cointercalated SnSe2. Well-separated cointercalation stages with different stacking patterns have been discovered by carefully controlled reaction time and concentration of solutions. These cointercalation stages are superconductors showing different superconducting signals. In addition, Li and various organic species such as Acetone, Dimethyl sulfoxide (DMSO) and Tetrahydrofuran (THF) have been cointercalated into SnSe2 crystals, all of which show enhanced superconducting Tc compared to solely Li intercalated SnSe2. Our findings may provide more insight to effectively tune electronic structure of the lamellar structure through organic molecules co-regulation, and open a new strategy to engineer the physical properties of these layered materials by controlling their different intercalation stages.

cond-mat.supr-con↗

The pairing symmetry in quasi-one-dimensional superconductor Rb2Mo3As3

Quasi-one-dimensional electron systems display intrinsic instability towards long-range ordered phases at sufficiently low temperatures. The superconducting orders are of particular interest as they can possess either singlet or triplet pairing symmetry and frequently compete with magnetism. Here we report on muon spin rotation and relaxation ($\mathrmμ$SR) study of Rb$_2$Mo$_3$As$_3$ characterised by one of the highest critical temperatures $T_{\rm c}=10.4\ \mathrm{K}$ among quasi-one-dimensional superconductors. The transverse-field $\mathrmμ$SR signal shows enhanced damping below $T_{\rm c}$ due to the formation of vortex lattice. Comparison of vortex lattice broadening against single gap $s-$, $p-$ and $d-$wave models shows the best agreement for the $s-$wave scenario but with the anomalously small superconducting gap, $Δ_0$, to $T_{\rm c}$ ratio of $2Δ_0/k_{\rm B}T_{\rm c}=2.74(1)$. The alternative nodal $p-$wave or $d-$wave scenarios with marginally worse goodness of fit would yield more realistic $2Δ_0/k_{\rm B}T_{\rm c}=3.50(2)$ and $2Δ_0/k_{\rm B}T_{\rm c}=4.08(1)$, respectively, and thus they cannot be ruled out when accounting for the superconducting state in Rb$_2$Mo$_3$As$_3$.

cond-mat.supr-con↗

Ideal Weak Topological Insulator and Protected Helical Saddle Points

The paradigm of classifying three-dimensional (3D) topological insulators into strong and weak ones (STI and WTI) opens the door for the discovery of various topological phases of matter protected by different symmetries and defined in different dimensions. However, in contrast to the vast realization of STIs, very few materials have been experimentally identified as being close to WTI. Even amongst those identified, none exists with topological surface states (TSS) exposed in a global bulk band gap that is stable at all temperatures. Here we report the design and observation of an ideal WTI in a quasi-one-dimensional (quasi-1D) bismuth halide, Bi$_{4}$I$_{1.2}$Br$_{2.8}$ (BIB). Via angle-resolved photoemission spectroscopy (ARPES), we identify that BIB hosts TSS on the (100)$\prime$ side surface in the form of two anisotropic $π$-offset Dirac cones (DCs) separated in momentum while topologically dark on the (001) top surface. The ARPES data fully determine a unique side-surface Hamiltonian and thereby identify two pairs of non-degenerate helical saddle points and a series of four Lifshitz transitions. The fact that both the surface Dirac and saddle points are in the global bulk band gap of 195 meV, combined with the small Dirac velocities, nontrivial spin texture, and the near-gap chemical potential, qualifies BIB to be not only an ideal WTI but also a fertile ground for topological many-body physics.

cond-mat.mtrl-sci↗

Blind Super-Resolution for Remote Sensing Images via Conditional Stochastic Normalizing Flows

Remote sensing images (RSIs) in real scenes may be disturbed by multiple factors such as optical blur, undersampling, and additional noise, resulting in complex and diverse degradation models. At present, the mainstream SR algorithms only consider a single and fixed degradation (such as bicubic interpolation) and cannot flexibly handle complex degradations in real scenes. Therefore, designing a super-resolution (SR) model that can cope with various degradations is gradually attracting the attention of researchers. Some studies first estimate the degradation kernels and then perform degradation-adaptive SR but face the problems of estimation error amplification and insufficient high-frequency details in the results. Although blind SR algorithms based on generative adversarial networks (GAN) have greatly improved visual quality, they still suffer from pseudo-texture, mode collapse, and poor training stability. In this article, we propose a novel blind SR framework based on the stochastic normalizing flow (BlindSRSNF) to address the above problems. BlindSRSNF learns the conditional probability distribution over the high-resolution image space given a low-resolution (LR) image by explicitly optimizing the variational bound on the likelihood. BlindSRSNF is easy to train and can generate photo-realistic SR results that outperform GAN-based models. Besides, we introduce a degradation representation strategy based on contrastive learning to avoid the error amplification problem caused by the explicit degradation estimation. Comprehensive experiments show that the proposed algorithm can obtain SR results with excellent visual perception quality on both simulated LR and real-world RSIs.

eess.IV↗

Lightweight Stepless Super-Resolution of Remote Sensing Images via Saliency-Aware Dynamic Routing Strategy

Deep learning-based algorithms have greatly improved the performance of remote sensing image (RSI) super-resolution (SR). However, increasing network depth and parameters cause a huge burden of computing and storage. Directly reducing the depth or width of existing models results in a large performance drop. We observe that the SR difficulty of different regions in an RSI varies greatly, and existing methods use the same deep network to process all regions in an image, resulting in a waste of computing resources. In addition, existing SR methods generally predefine integer scale factors and cannot perform stepless SR, i.e., a single model can deal with any potential scale factor. Retraining the model on each scale factor wastes considerable computing resources and model storage space. To address the above problems, we propose a saliency-aware dynamic routing network (SalDRN) for lightweight and stepless SR of RSIs. First, we introduce visual saliency as an indicator of region-level SR difficulty and integrate a lightweight saliency detector into the SalDRN to capture pixel-level visual characteristics. Then, we devise a saliency-aware dynamic routing strategy that employs path selection switches to adaptively select feature extraction paths of appropriate depth according to the SR difficulty of sub-image patches. Finally, we propose a novel lightweight stepless upsampling module whose core is an implicit feature function for realizing mapping from low-resolution feature space to high-resolution feature space. Comprehensive experiments verify that the SalDRN can achieve a good trade-off between performance and complexity. The code is available at \url{https://github.com/hanlinwu/SalDRN}.

cs.CV↗

Scale-Aware Dynamic Network for Continuous-Scale Super-Resolution

Single-image super-resolution (SR) with fixed and discrete scale factors has achieved great progress due to the development of deep learning technology. However, the continuous-scale SR, which aims to use a single model to process arbitrary (integer or non-integer) scale factors, is still a challenging task. The existing SR models generally adopt static convolution to extract features, and thus unable to effectively perceive the change of scale factor, resulting in limited generalization performance on multi-scale SR tasks. Moreover, the existing continuous-scale upsampling modules do not make full use of multi-scale features and face problems such as checkerboard artifacts in the SR results and high computational complexity. To address the above problems, we propose a scale-aware dynamic network (SADN) for continuous-scale SR. First, we propose a scale-aware dynamic convolutional (SAD-Conv) layer for the feature learning of multiple SR tasks with various scales. The SAD-Conv layer can adaptively adjust the attention weights of multiple convolution kernels based on the scale factor, which enhances the expressive power of the model with a negligible extra computational cost. Second, we devise a continuous-scale upsampling module (CSUM) with the multi-bilinear local implicit function (MBLIF) for any-scale upsampling. The CSUM constructs multiple feature spaces with gradually increasing scales to approximate the continuous feature representation of an image, and then the MBLIF makes full use of multi-scale features to map arbitrary coordinates to RGB values in high-resolution space. We evaluate our SADN using various benchmarks. The experimental results show that the CSUM can replace the previous fixed-scale upsampling layers and obtain a continuous-scale SR network while maintaining performance. Our SADN uses much fewer parameters and outperforms the state-of-the-art SR methods.

cs.CV↗

Transport anomalies in the layered compound BaPt4Se6

We report a layered ternary selenide BaPt4Se6 featuring sesqui-selenide Pt2Se3 layers sandwiched by Ba atoms. The Pt2Se3 layers in this compound can be derived from the Dirac-semimetal PtSe2 phase with Se vacancies that form a honeycomb structure. This structure results in a Pt (VI) and Pt (II) mixed-valence compound with both PtSe6 octahedra and PtSe4 square net coordination configurations. Temperature dependent electrical transport measurements suggest two distinct anomalies: a resistivity crossover, mimic to the metal-insulator (M-I) transition at ~150K, and a resistivity plateau at temperatures below 10K. The resistivity crossover is not associated with any structural, magnetic or charge order modulated phase transitions. Magnetoresistivity, Hall and heat capacity measurements concurrently suggest an existing hidden state below 5K in this system. Angle-resolved photoemission spectroscopy measurements reveal a metallic state and no dramatic reconstruction of the electronic structure up to 200K.

cond-mat.str-el↗