SearcharxivSearch

arXiv subjects

Hui Chen

Publications and source records attributed to Hui Chen.

At least 19 recordsLinked to original sources

Self-organized positron reorienting and pinching mechanism for the experimental detection of the linear Breit-Wheeler process

The linear Breit-Wheeler (LBW) process ($\gamma+\gamma\rightarrow e^{-}+e^{+}$) is a fundamental prediction of quantum electrodynamics, but yet to be observed under laboratory conditions using real photons. In recent years, a few experimental schemes utilizing high-intense ($\sim10^{22}$W/cm$^2$) laser-plasma interactions to observe the LBW process have been proposed. However, a high level of signal-to-noise-ratio are expected in these schemes, hindering the first-ever experimental detection of the LBW process by real photons. In this paper, we present a simple experimental setup which could enhance the expected positron signals by 2-3 orders of magnitude compared to previously proposed schemes, reaching the level of $10^{6}$MeV$^{-1}$str$^{-1}$. Moreover, such high positron signal is achieve in the direction opposite to the laser propagation, where a significantly quieter background is expected compared to the previously focused direction of laser propagation. The key to achieve this result is a newly discovered self-organized positron reorienting and pinching mechanism, enabled by the in-situ strong plasma fields from the laser-plasma interaction.

physics.plasm-ph

Foundation Models for Wireless Localization: Pretraining, Adaptation, and Utilization

Accurate wireless localization is a key enabler for 6G networks, yet remains challenging under diverse and rapidly changing propagation conditions. Model-based methods degrade when multipath channels are non-resolvable and model mismatches occur, while supervised deep learning demands large labeled datasets and generalizes poorly to new deployments. Inspired by foundation models (FMs) in language and vision, this article presents a unified framework for FM-based wireless localization that learns transferable channel representations from large-scale unlabeled channel state information and adapts to new environments with minimal or even no supervision. We review the fundamentals of FMs, compare the FM paradigm with existing localization approaches, and introduce a three-stage framework spanning large-scale pretraining, localization-oriented fine-tuning, and context-augmented inference, together with the location-aware applications it enables. Ray-tracing-based case studies show improved positioning accuracy and cross-environment generalization. Finally, we present an outlook on key research directions toward AI-native networks for wireless localization.

eess.SP

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

Test-time adaptation (TTA) has been widely explored in single-label recognition, effectively mitigating distribution shifts, especially when combined with vision-language models. However, real-world images often contain multiple objects, while the more practical multi-label test-time adaptation (MLTTA) has received little attention so far. Recent cache-based TTA methods have shown promising efficiency and effectiveness, yet directly extending them to multi-label scenarios suffers from a one-to-many mapping problem: a shared global representation entangling co-occurring objects is stored as class-wise cache prototypes, inducing dominant-label bias and compromised cache calibration. While introducing region-level cues helps isolate class-specific evidence, such regional evidence can also be unreliable under distribution shifts, making its identification and utilization non-trivial. To address these issues, we introduce PuRF, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models. Specifically, PuRF first performs region purification to identify reliable regions, providing comprehensive regional cues for multi-label recognition and enabling fine-grained alignment. Based on these purified regions, PuRF conducts cache purification to enhance cache representation and adaptability, where episodic purification builds a discriminative region-based cache, and temporal refreshing further promotes long-term cache adaptability. Experiments demonstrate that PuRF consistently outperforms state-of-the-art methods, achieving a notable 4.05% mAP improvement on ViT-B/32 across five datasets.

cs.CV

Realization of Air-Stable Two-Dimensional Superconductor Nb2Pd3Te5 With Quasi-One-Dimensional Pair Density Modulation

Two-dimensional (2D) superconductors provide a fertile platform for exploring reduced-dimensional superconductivity and emergent quantum phenomena. Incorporating quasi-one-dimensional (quasi-1D) structural motifs into 2D superconductors offers a powerful route to engineer strong electronic anisotropy, enabling unconventional superconducting states and anisotropic superconducting transport functionalities. However, such systems remain rarely realized. Here we report the realization of a 2D superconductor Nb2Pd3Te5, exhibiting an intrinsic quasi-1D pair density modulation. Monolayer and bilayer Nb2Pd3Te5 is synthesized via van-der-Waals epitaxy. Using ultralow-temperature scanning tunneling microscopy/spectroscopy, we observe the quasi-1D crystal structure and superconductivity below ~0.6 K with a pronounced quasi-1D pair density modulation. Remarkably, both monolayer and bilayer Nb2Pd3Te5 show strong air stability. Our findings establish atomically 2D Nb2Pd3Te5 as a robust and promising platform for exploring novel low-dimensional quantum phenomena and anisotropy-enabled superconducting devices.

cond-mat.mtrl-sci

Sci-Surf: Navigating Scientific Literature Discovery through Human Feedback and Intelligent Summarization

The rapid growth of scientific publications makes it increasingly difficult for researchers to identify relevant new studies and effectively comprehend them. Existing academic discovery platforms typically rely on static topic subscriptions or embedding-based similarity and provide only abstracts or short summaries, offering limited support for nuanced intent modeling and in-depth paper summarization. We present Sci-Surf, an intent-centric knowledge discovery system that integrates feedback-driven personalized recommendation with multi-modal blog-style paper digestion. Our approach refines user intent representations through LLM-based user profiling, while generating structured summaries that synthesize textual and visual information from full papers. The demo presents an end-to-end academic discovery pipeline and demonstrates measurable improvements in both recommendation quality and digestion quality through real-user evaluations. Specifically, the integration of verbalized profiles led to a 10.4% average improvement in predictive alignment with real-world user preferences throughout a month-long online evaluation.

cs.IR

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-grained, clinically reliable benchmarks that reflect expert interpretation. This work introduces PanDent, a large-scale, clinically grounded OPG benchmark built upon fine-grained, expert-validated tooth-level annotations. The dataset comprises 9,524 high-quality OPGs, each associated with comprehensive structured annotations produced by experienced dentists and further validated by an oral and maxillofacial radiologist, providing clinically reliable supervision for tooth-level diagnosis and reasoning. Clinically consistent radiology reports are constructed from expert-validated findings using clinician-defined reporting logic, establishing explicit correspondence between structured clinical evidence and free-text descriptions. This design enables evaluation of whether MLLMs generate reports that are not only linguistically coherent but also clinically consistent with expert-validated tooth-level findings. Experiments are conducted on diverse MLLMs, including state-of-the-art (SOTA) proprietary models, general-domain open-source models, and medical-specific models. Results show that current MLLMs can generate fluent reports, yet fail to produce clinically consistent descriptions, exhibiting substantial errors in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent significantly improves structure-language consistency, substantially enhancing visual localization accuracy and diagnostic correctness, and bringing model outputs closer to expert dental interpretation. These results establish PanDent as a rigorous benchmark for evaluating tooth-level clinical reasoning in MLLMs and a valuable resource for clinically grounded dental AI.

cs.CV

On-Site Beam Calibration for RIS-Aided Positioning Systems

High precision positioning is a key enabler for next-generation communication applications such as smart transportation and augmented reality. Reconfigurable intelligent surface (RIS) technology can enhance positioning by providing additional angular information and improving coverage under obstructed propagation conditions. However, true RIS beams can differ significantly from the simplified or ideal beam response models commonly used in RIS-aided positioning, leading to beam model mismatch and an elevated positioning error floor. This paper proposes an on-site RIS beam calibration framework that reduces this error floor by estimating a realistic 3D RIS beam response model from on-site measurements. The proposed calibration algorithm first extracts the RIS-reflected channel response from signals received by a calibration agent sampling the angular range of interest, using delay-domain sparse recovery, and then estimates the beam model parameters with a gradient-based estimator. To validate the proposed framework, 3D beam patterns under 66 phase modulations were measured and incorporated into simulations. With an angular sampling step of 1 deg, the calibrated model achieves an average beam response similarity of 88.5% with respect to the ground truth, compared with 43.7% for the ideal model. The probability that the absolute lower bound of the positioning error is below 0.5m increases from 0.52 without calibration to 0.74 after calibration, showing that on-site RIS beam calibration effectively reduces the positioning error floor caused by true beam model mismatch.

eess.SP

The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing

Single-pixel sensing encodes a scene as a short sequence of coded measurements, and image-free methods infer the task directly from that sequence. We show that removing image reconstruction relocates the central design problem to the lift: how 1D measurements become a 2D task representation. We organize this choice as a lift spectrum from a fixed-physics inverse, through a learned static projection, to content-adaptive retrieval. These are not interchangeable forms of reconstruction: the fixed-physics route reconstructs an image consumed at inference, whereas our spatiotemporal soft-fusion (STSF) network lifts measurements directly into task features, and task-prioritized loss scheduling (TPLS) uses a separate learned reconstruction branch only as scheduled training supervision. A probe-selected recurrent encoder and a parameter-matched lift ablation identify the STSF design. In simulation, STSF+TPLS exceeds the prior image-free baseline on three datasets at 3.13% sampling (+3.2 to +9.9 pp foreground mIoU) and remains competitive down to 0.39%. The strongest clean-trained reconstruct-then-segment baseline wins without measurement noise, but measurement noise reverses the ranking: the reconstructed task input carries a 20-70x larger normalized relative perturbation than the measurements themselves. Stressed to failure, the three lift regions exhibit distinct dominant signatures--collapse, imprinting, and coarsening. STSF+TPLS transfers without fine-tuning to a real single-pixel bench, where the reversal reappears as a proof of concept; inference takes about 14 ms per mask on an RTX 4090. Within the tested fixed-acquisition regime, measurement-to-space adaptivity therefore organizes both the clean-to-noisy operating envelope and the failure a system encounters. Code and pretrained weights: https://github.com/Hanyuyuan6/STSF-TPLS.

eess.IV

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach

This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS). Since localization pilot transmissions dominate the energy budget of power-constrained devices, we introduce a low-overhead feedback link from the Base Station (BS) to the user to enable dynamic uplink power control. To navigate the discrete and decentralized nature of this active sensing problem, we propose a novel Dual-Agent (DA) deep learning framework that jointly optimizes the discrete RIS phase profiles and the UE's transmit power in real time. Specifically, our approach employs a hybrid training methodology integrating the neuroevolution paradigm with supervised learning, effectively overcoming the non-differentiability of discrete phase responses from the RIS unit elements and the strict information bottleneck of single-bit feedback messages for pilot power control. The proposed DA active sensing framework can be applied with both single- and multi-antenna BSs, the latter with only minor modifications in the structure of one NN: an additional output branch with appropriate structure is included for the latter case to select a valid digital combiner from a finite set. Extensive numerical simulations demonstrate that the proposed scheme achieves highly accurate and robust tracking across diverse target motion models, outperforming extended Kalman and particle filters, as well as, machine learning-based trackers. Furthermore, in static localization, it is shown to significantly outperform traditional fingerprinting schemes, deep reinforcement learning baselines, and standard backpropagation-based estimators.

cs.IT

Maximal regularity and caloric trace estimates in mixed Lebesgue norms for the heat equation

We study the heat equation in the half-space with nonhomogeneous Dirichlet boundary data. For the caloric extension $v$ of the boundary data $g$, we prove maximal regularity estimates in mixed Lebesgue norms $L^p_tL^q_x$ for any order derivative of $v$ in terms of mixed Besov and Lizorkin--Triebel type norms of $g$. We also establish the corresponding reverse inequalities, which are caloric trace estimates recovering the boundary regularity of $g$ from the mixed-norm regularity of $v$. As a model case, our results show that the natural \[\dotc W^{1,p}\big(\R;L^q(\R^d_+)\big)\cap L^p\big(\R;\dotc W^{2,q}(\R^d_+)\big)\] regularity norm of $v$ is controlled by the \[\dotc {F}^{1-\frac{1}{2q}}_{p,q}\big(\R;\,L^{q}(\R^{d-1})\big)\cap L^{p}\big(\R;\,\dotc{B}_{q,q}^{2-\frac{1}{q}}(\R^{d-1})\big)\] norm of $g$. The maximal regularity estimate holds for $1\leq p,q<\infty$, while the caloric trace estimate holds for $1<p<\infty$ and $1\leq q\leq\infty$. In particular, the endpoint cases $p=1$ or $q=1$ in the maximal regularity estimate are included and appear to be new. These endpoint estimates may be useful in the analysis of free-boundary Navier--Stokes problems with small initial data, whereas the caloric trace estimates may be relevant to the construction of Stokes or Navier--Stokes flows exhibiting strong boundary singularities.

math.AP

HORIZON: Recoverability-Governed Curriculum for Physical-Domain Scaling

Scaling robust robot policies requires more than broader randomization, because physical-domain experience must remain organized and learnable throughout training. We study when a policy can benefit from harder physics and identify recoverability as a central constraint in on-policy physical-domain scaling. In on-policy training, new dynamics are useful only insofar as they remain close enough to the current policy to generate corrective on-policy data, rather than collapsing rollouts into unrecoverable failures. Using quadruped locomotion as a physically demanding benchmark for embodied generalization, we introduce HORIZON, a checkpointed frontier curriculum that expands physical domains only within the current policy's recoverable boundary. HORIZON uses rollback and boundary refinement to govern each expansion step, turning fixed randomization into a continual process of physical-domain growth. Experiments reveal three regularities of physical-domain expansion. First, direct domain widening is uneven across physical axes and often unlearnable without staged ordering. Second, domain composition is non-monotonic, and adding more domains beyond a compact core can dilute recoverable joint samples and reduce overall robustness. Third, offline distillation of isolated experts cannot substitute for the joint interaction generated by on-policy curriculum. Together, these results frame physical-domain generalization as a continual growth problem for embodied control, with recoverability as the organizing principle for on-policy expansion.

cs.RO

RA-LWLM: Retrieval-Augmented In-Context Localization with Wireless Foundation Models

Wireless localization is a fundamental capability of sixth-generation (6G) networks. Conventional model-based methods require accurate modeling of the propagation environment and degrade in complex multipath and non-line-of-sight scenarios, while learning-based methods couple model parameters tightly to the training scene, requiring costly retraining whenever the base station (BS) configuration or propagation environment changes. In this paper, we propose RA-LWLM, a retrieval-augmented in-context localization framework that achieves training-free cross-scene adaptation by externalizing scene-specific information into a per-scene fingerprint database rather than encoding it in model weights. The framework consists of three components: a frozen wireless foundation model (FM) encoder that maps raw channel state information into a scene-agnostic representation; a retrieval module that selects the most informative references from the per-scene database via similarity search in the representation space; and a transformer-based in-context learning (ICL) module that fuses the query with the retrieved references to predict the user equipment (UE) position. To accommodate varying retrieval quality and propagation complexity across queries, the ICL module adopts a mixture-of-experts design in which experts specialize in different context sizes and are softly combined by a learnable selector. Extensive ray-tracing-based experiments across heterogeneous scenes with diverse BS configurations show that RA-LWLM achieves nearly identical accuracy on seen and unseen scenes without any per-scene retraining, substantially outperforming end-to-end and FM-based baselines. These results validate the proposed retrieval-augmented in-context paradigm as a scalable solution for cross-scene localization in 6G networks.

eess.SP

Multicast Capacity of XL-RIS Assisted Hybrid Near- and Far-Field mmWave Communications

Multicast transmission in millimeter-wave (mmWave) networks is fundamentally limited by the weakest user, and blockages further exacerbate this problem. Large-scale reconfigurable intelligent surfaces (XL-RIS) offer a promising solution by providing high array gain to overcome blockages. However, the large aperture of XL-RIS significantly expands the near-field region, creating a hybrid-field scenario where some users lie in the near-field while others remain in the far-field. Existing hybrid-field studies on XL-RIS have primarily focused on channel estimation and deployment optimization, leaving multicast capacity analysis unexplored. This paper investigates the fundamental capacity limits of XL-RIS-assisted multicast communications in hybrid-field scenarios. For the fundamental two-user case consisting of one near-field and one far-field user, we derive the optimal closed-form covariance matrix and optimize the RIS phase shifts via manifold optimization. We establish that the multicast capacity scales as $\Theta(\log_2(MN))$ as the number of transmit antennas M and/or RIS elements N grow large, and prove this scaling is order-tight. Numerical results validate the bounds and show the impact of M, $N$, and distance on the multicast rate.

eess.SP

Flexible Rate-Splitting for Joint Unicast and Multi-Group Multicast Transmission in RIS-Assisted mmWave Networks

Joint unicast and multi-group multicast transmission with RIS and RSMA is a promising enabler for 6G services. However, existing RSMA schemes for such scenarios split only unicast messages while leaving multicast messages intact, limiting the degree of freedom of interference management. To this end, we propose a joint rate splitting framework that splits both unicast and multicast information and two RSMA schemes. The common-common fusion (CCF-RSMA) scheme encodes the unicast common part into the global multicast common stream, while the private-common fusion (PCF-RSMA) scheme merges it with the group-specific multicast private part. For each scheme, we formulate energy efficiency (EE) maximization problems under both perfect and imperfect channel state information, and jointly optimize active beamforming, RIS phase shifts and rate allocation parameters. Simulation results demonstrate that the proposed schemes significantly outperform the comparative schemes in terms of EE, thereby proving the effectiveness of the proposed framework. Moreover, CCF-RSMA is more favorable in scenarios with larger groups and higher unicast QoS demands, whereas PCF-RSMA is better suited for scenarios with smaller groups and higher multicast QoS.

eess.SP

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve these agents, but current approaches can fail, because correct answers are often sparse and score-based selection depends on model calibration. We propose FineVerify, a fine-grained self-verification framework that decomposes each question into checkable sub-questions, verifies sampled candidates against each sub-question, and selects the candidate with the highest aggregated score. This per-check structure turns selection into simpler local judgments and produces scores under the same explicit criteria. Across four agentic search benchmarks and two models, FineVerify consistently outperforms standard scaling baselines. With only four sampled trajectories, it improves GPT-5-mini by 8.2 accuracy points and Gemini-3-flash by 5.6% on average. With 12 samples, FineVerify enables GPT-5-mini to surpass frontier GPT-5 on BrowseComp-Plus. Beyond accuracy, FineVerify produces interpretable verification traces that help audit benchmark errors, suggesting broader applications for inspecting agentic search systems. Code and data are available at https://github.com/XuZhao0/fineverify

cs.CL

Modulation of charge density waves in a twisted vortex moire superlattice

Twisted moire superlattices in van-der-Waals heterostructures provide a powerful platform for engineering correlated states through moire-band reconstruction. However, whether globally coherent electronic orders can be continuously manipulated at the nanoscale remains largely unexplored. Reconstructed moire structures in small-angle and near-commensurate regime feature continuously varying local environments, offering new opportunities for nanoscale manipulation of correlated phases. Here, we report the modulation of charge density wave (CDW) states in a twisted vortex moire superlattice formed between monolayer VTe2 and superconducting NbSe2. Scanning tunneling microscopy/spectroscopy reveals that the intrinsic long-range CDW of monolayer VTe2 is reconstructed into inequivalent local phases with distinct stability and coherence within a single moire unit cell, including suppressed CDW order and enhanced short-range CDW correlations persisting to room temperature. First-principles calculations show that the reconstructed CDW landscape originates from strong local strain variation, where compressive strain substantially stabilizes the charge order. Furthermore, the modulated CDW states exhibit competing interplay with proximity-induced superconductivity. Our results establish vortex moire superlattices as a versatile platform for nanoscale manipulation of correlated electronic orders in low-dimensional quantum materials.

cond-mat.mtrl-sci

Bifurcation of the quasi-stationary velocity of strongly discrete transition waves driven by gravity

Transition waves are common in multistable mechanical metamaterials, and the dynamics of weakly discrete transition waves under driving forces have been extensively discussed. However, as lattice effects become more pronounced, strongly discrete transition waves may exhibit dynamics that cannot be predicted by the continuum limit. Here, by tilting a bistable chain, we introduce a gravitational perturbation term into the dynamical equations, under which the transition waves are continuously accelerated. In the strongly discrete regime, we find that transition waves under gravitational driving possess quasi-stationary velocity plateaus (QSVPs), and the number of these plateaus first increases and then decreases as the tilt angle increases. We theoretically elucidate that the emergence of the velocity plateaus originates from the balance between gravitational driving and phonon radiation. In further analysis, the theoretical model reveals that the balance point undergoes a bifurcation at the radiation resonance, which leads to a change in the number of velocity plateaus. Our study extends the investigation of transition waves into the strongly discrete regime, and the emergence of multiple velocity plateaus opens up new possibilities for programmable solitary waves.

nlin.PS

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing

Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens required to encode dense documents incurs prohibitive inference cost. Existing pruning methods rely on physical eviction, e.g., permanently discarding visual tokens during the prefill stage. While effective for natural images, this strategy fundamentally breaks down on OCR, where virtually every visual token may correspond to a character or structural element, and any irreversible loss leads to catastrophic accuracy degradation. We observe that, although document images appear globally dense and seemingly unprunable, the model's attention to them is in fact temporally sparse: at each decoding step it concentrates on a small region that shifts gradually across steps, much as a human reader fixates on successive words rather than perceiving an entire page at once. Motivated by this Dynamic Visual Fixation phenomenon, we recast the intractable global pruning problem as a tractable local, dynamic one and propose FastOCR, a training-free framework with two complementary modules. Specifically, Focal-Guided Pruning identifies a small set of focal layers and selects the most task-relevant visual tokens from them at each step, while Cross-Step Fixation Reuse exploits the gradual shift of fixation to warm-start each step from the previous one. By dynamically adjusting which tokens are attended rather than evicting any from the cache, FastOCR avoids permanent information loss. Extensive experiments show that FastOCR serves as a plug-and-play acceleration module, generalizing consistently across five VLMs of varying sizes and architectures. On Qwen2.5-VL, FastOCR retains 98% of the unpruned model's accuracy while attending to only 5% of the visual tokens per decoding step, reducing attention latency by 3.0$\times$.

cs.CV