SearcharxivSearch

arXiv subjects

Yifan Shao

Publications and source records attributed to Yifan Shao.

10 recordsLinked to original sources

Edge-Dominated Twist Mechanics at van der Waals Interfaces

Despite the pivotal role of twist in modulating physical properties at van der Waals (vdW) interfaces, the mechanics governing torsional response remain poorly understood. Here, we probe twist mechanics at homo- and heterogeneous vdW interfaces, together with their sliding behaviors within a unified experimental framework. For both systems, the peak torque scales nearly linearly with contact area, in contrast to predictions from linear elastic and rigid models. Remarkably, while the sliding friction of the two interfaces diverges by over three orders of magnitude owing to different scaling laws, the corresponding torque follows the same linear scaling and differs by only about twenty-fold. Large-scale atomistic simulations reveal an edge-dominated yielding mechanism for torsional motion, wherein elastic reconstruction shifts the effective load-bearing region toward the edges, eliminating torque from the contact interior. This mechanism contrasts with the bulk-mediated stress transmission governing translational sliding, a distinction rooted in the different loading geometries inherent to the two motion modes, where torsional loading necessitates perimeter actuation, whereas sliding enables center-driven loading. This symmetry-imposed divergence demonstrates that translational and torsional properties cannot be predicted from one another at vdW interfaces, providing critical insights for the design of dynamically reconfigurable micro- and nanoelectromechanical devices.

cond-mat.mtrl-sci

Observation of robust macroscale structural superlubricity

Structural superlubricity (SSL) promises nearly frictionless and wearless sliding, but has until now been considered a special and extreme interfacial phenomenon limited to micro- and nanoscale contacts. Here, we demonstrate robust macroscale SSL within a single sub-millimeter graphite contact. Previously reported near-zero friction coefficients, where friction is nearly independent of normal load, have only been observed at microscale contacts under low loads. Our system expands both contact size and load into the macroscopic regime, exhibiting friction coefficients that fluctuate around zero and reach values as low as $10^{-6}$ across a broad load range from 1 mN to 0.5 N. Negative friction coefficients are also observed. Similar behavior is observed at graphite/MoS$_2$ interfaces, indicating that macroscale SSL is a generalizable phenomenon across flat layered materials. These findings overturn long-standing scaling limitations and establish macroscale SSL as a paradigm-shifting platform for next-generation mechanical and electromechanical systems.

cond-mat.mtrl-sci

Intent-Guided Reasoning for Sequential Recommendation

Sequential recommendation systems aim to capture users' evolving preferences from their interaction histories. Recent reasoningenhanced methods have shown promise by introducing deliberate, chain-of-thought-like processes with intermediate reasoning steps. However, these methods rely solely on the next target item as supervision, leading to two critical issues: (1) reasoning instability--the process becomes overly sensitive to recent behaviors and spurious interactions like accidental clicks, and (2) surface-level reasoning--the model memorizes item-to-item transitions rather than understanding intrinsic behavior patterns. To address these challenges, we propose IGR-SR, an Intent-Guided Reasoning framework for Sequential Recommendation that anchors the reasoning process to explicitly extracted high-level intents. Our framework comprises three key components: (1) a Latent Intent Distiller (LID) that efficiently extracts multi-faceted intents using a frozen encoder with learnable tokens, (2) an Intent-aware Deliberative Reasoner (IDR) that decouples reasoning into intent deliberation and decision-making via a dual-attention architecture, and (3) an Intent Consistency Regularization (ICR) that ensures robustness by enforcing consistent representations across different intent views. Extensive experiments on three public datasets demonstrate that IGR-SR achieves an average 7.13% improvement over state-of-the-art baselines. Critically, under 20% behavioral noise, IGR-SR degrades only 10.4% compared to 16.2% and 18.6% for competing methods, validating the effectiveness and robustness of intent-guided reasoning.

cs.IR

DTRec: Learning Dynamic Reasoning Trajectories for Sequential Recommendation

Inspired by advances in LLMs, reasoning-enhanced sequential recommendation performs multi-step deliberation before making final predictions, unlocking greater potential for capturing user preferences. However, current methods are constrained by static reasoning trajectories that are ill-suited for the diverse complexity of user behaviors. They suffer from two key limitations: (1) a static reasoning direction, which uses flat supervision signals misaligned with human-like hierarchical reasoning, and (2) a fixed reasoning depth, which inefficiently applies the same computational effort to all users, regardless of pattern complexity. These rigidity lead to suboptimal performance and significant computational waste. To overcome these challenges, we propose DTRec, a novel and effective framework that explores the Dynamic reasoning Trajectory for Sequential Recommendation along both direction and depth. To guide the direction, we develop Hierarchical Process Supervision (HPS), which provides coarse-to-fine supervisory signals to emulate the natural, progressive refinement of human cognitive processes. To optimize the depth, we introduce the Adaptive Reasoning Halting (ARH) mechanism that dynamically adjusts the number of reasoning steps by jointly monitoring three indicators. Extensive experiments on three real-world datasets demonstrate the superiority of our approach, achieving up to a 24.5% performance improvement over strong baselines while simultaneously reducing computational cost by up to 41.6%.

cs.IR

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

As large language models (LLMs) evolve into tool-using agents, the ability to browse the web in real-time has become a critical yardstick for measuring their reasoning and retrieval competence. Existing benchmarks such as BrowseComp concentrate on English and overlook the linguistic, infrastructural, and censorship-related complexities of other major information ecosystems -- most notably Chinese. To address this gap, we introduce BrowseComp-ZH, a high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web. BrowseComp-ZH consists of 289 multi-hop questions spanning 11 diverse domains. Each question is reverse-engineered from a short, objective, and easily verifiable answer (e.g., a date, number, or proper noun). A two-stage quality control protocol is applied to strive for high question difficulty and answer uniqueness. We benchmark over 20 state-of-the-art language models and agentic search systems on our proposed BrowseComp-ZH. Despite their strong conversational and retrieval capabilities, most models struggle severely: a large number achieve accuracy rates below 10%, and only a handful exceed 20%. Even the best-performing system, OpenAI's DeepResearch, reaches just 42.9%. These results demonstrate the considerable difficulty of BrowseComp-ZH, where success demands not only effective retrieval strategies, but also sophisticated reasoning and information reconciliation -- capabilities that current models still struggle to master. Our dataset, construction guidelines, and benchmark results have been publicly released at https://github.com/PALIN2018/BrowseComp-ZH.

cs.CL

Local-Global Attention: An Adaptive Mechanism for Multi-Scale Feature Integration

In recent years, attention mechanisms have significantly enhanced the performance of object detection by focusing on key feature information. However, prevalent methods still encounter difficulties in effectively balancing local and global features. This imbalance hampers their ability to capture both fine-grained details and broader contextual information-two critical elements for achieving accurate object detection.To address these challenges, we propose a novel attention mechanism, termed Local-Global Attention, which is designed to better integrate both local and global contextual features. Specifically, our approach combines multi-scale convolutions with positional encoding, enabling the model to focus on local details while concurrently considering the broader global context. Additionally, we introduce a learnable parameters, which allow the model to dynamically adjust the relative importance of local and global attention, depending on the specific requirements of the task, thereby optimizing feature representations across multiple scales.We have thoroughly evaluated the Local-Global Attention mechanism on several widely used object detection and classification datasets. Our experimental results demonstrate that this approach significantly enhances the detection of objects at various scales, with particularly strong performance on multi-class and small object detection tasks. In comparison to existing attention mechanisms, Local-Global Attention consistently outperforms them across several key metrics, all while maintaining computational efficiency.

cs.CV

Pixelated Bayer Spectral Router Based on Sparse Meta-atom Array

It has long been a challenging task to improve the light collection efficiency of conventional image sensors built with color filters that inevitably cause the energy loss of out-of-band photons. Although various schemes have been proposed to address the issue, it is still very hard to make a reasonable tradeoff between device performance and practicability. In this work, we demonstrate a pixelated spectral router based on sparse meta-atom array, which can efficiently separate the incident R (600-700 nm), G (500-600 nm), and B (400-500 nm) band light to the corresponding pixels of a Bayer image sensor, providing over 56% signal enhancement above the traditional color filter scheme. The CMOS-compatible spectral router has superior characteristics of polarization insensitivity and high incident angle tolerance (over 30°), enabled by simple compound Si3N4 nanostructures which are very suitable for massive production. Imaging experiments are conducted to verify its potential for real applications. Our pixelated spectral router scheme is also found to be robust and could be freely adapted to image sensors of various pixel sizes, having great potential in building the new generation of high-performance image sensing components.

physics.optics

Pixel-scale NIR-VIS Spectral Routers Based on 2D Mie-type Metagratings

The out-of-band energy loss caused by in-built color filters significantly degrades the signal-to-noise ratio and the dynamic range of conventional image sensors, which has restricted the attempt to develop ultrahigh-density imaging devices by merely shrinking the pixel size. This issue will be more serious for security cameras which need to collect visible (VIS) light and near-infrared (NIR) photons as well. The existing solutions mostly explore complex photonic nanostructures, which are often too complicated for production. In this work, we demonstrate a pixel-scale spectral router utilizing two-dimensional (2D) Si3N4 Mie scattering metagratings that can spatially divide NIR (850 nm) and VIS (400-700 nm) light to different pixels at high efficiencies. It has a minimum feature size larger than 360 nm, highly promising for massive production. Compared with the traditional filter design, our router can gain about 42% and 30% signal enhancement for NIR and VIS band, respectively. We show that it also has good polarization insensitivity and incident angle tolerance. The NIR-VIS simultaneous imaging is inspected without any complex reconstruction algorithm. Mode analysis indicates that the multipolar scattering of our Mie-type metagratings provides the necessary degrees of freedom to spatially optimize the routing functions for broadband photons.

physics.optics

Electromagnetic Spatiotemporal Differentiators

Spatiotemporal optical computing devices which could perform mathematical operations in both spatial and temporal domains can provide unprecedented measures to build efficient and real-time information processing systems. It is particularly important to realize the comprehensive functions in a compact design for better integration with electronic components. In this work, we experimentally demonstrated an analogue spatiotemporal differentiator in microwaves based on an asymmetrical metasurface which has a phase singularity in the spatiotemporal domain. We showed that this structure could give rise to a spatiotemporal transfer function required by an ideal first-order differentiator in both spatial and temporal domains by tailoring the unidirectional excitation of spoof surface plasmon polaritons (SSPPs). The spatial edge detection was performed utilizing a metallic slit and the temporal differentiation capability of the device was examined by Gaussian-like temporal pulses of different width. We further confirmed the differentiator demonstrated here could detect sharp changes of spatiotemporal pulses even with intricate profiles and theoretically estimated the resolution limits of the spatial and temporal edge detection. We also show that the pulse input after passing the spatiotemporal differentiator implemented here could carry a transverse orbital angular momentum (OAM) with a fractal topology charge which further increases the information quantity.

physics.app-ph

Spatial Moment Pooling Improves Neural Image Assessment

In recent years, there has been widespread attention drawn to convolutional neural network (CNN) based blind image quality assessment (IQA). A large number of works start by extracting deep features from CNN. Then, those features are processed through spatial average pooling (SAP) and fully connected layers to predict quality. Inspired by full reference IQA and texture features, in this paper, we extend SAP ($1^{st}$ moment) into spatial moment pooling (SMP) by incorporating higher order moments (such as variance, skewness). Moreover, we provide learning friendly normalization to circumvent numerical issue when computing gradients of higher moments. Experimental results suggest that simply upgrading SAP to SMP significantly enhances CNN-based blind IQA methods and achieves state of the art performance.

cs.CV