Searcharxiv⌕ Search

arXiv subjects

Sen Yang

Publications and source records attributed to Sen Yang.

At least 91 records · Page 5Linked to original sources

CMTNet: Convolutional Meets Transformer Network for Hyperspectral Images Classification

Hyperspectral remote sensing (HIS) enables the detailed capture of spectral information from the Earth's surface, facilitating precise classification and identification of surface crops due to its superior spectral diagnostic capabilities. However, current convolutional neural networks (CNNs) focus on local features in hyperspectral data, leading to suboptimal performance when classifying intricate crop types and addressing imbalanced sample distributions. In contrast, the Transformer framework excels at extracting global features from hyperspectral imagery. To leverage the strengths of both approaches, this research introduces the Convolutional Meet Transformer Network (CMTNet). This innovative model includes a spectral-spatial feature extraction module for shallow feature capture, a dual-branch structure combining CNN and Transformer branches for local and global feature extraction, and a multi-output constraint module that enhances classification accuracy through multi-output loss calculations and cross constraints across local, international, and joint features. Extensive experiments conducted on three datasets (WHU-Hi-LongKou, WHU-Hi-HanChuan, and WHU-Hi-HongHu) demonstrate that CTDBNet significantly outperforms other state-of-the-art networks in classification performance, validating its effectiveness in hyperspectral crop classification.

cs.CV↗

Parameterized quasinormal frequencies and Hawking radiation for axial gravitational perturbations of a holonomy-corrected black hole

As the fingerprints of black holes, quasinormal modes are closely associated with many properties of black holes. Especially, the ringdown phase of gravitational waveforms from the merger of compact binary components can be described by quasinormal modes. Serving as a model-independent approach, the framework of parameterized quasinormal frequencies offers a universal method for investigating quasinormal modes of diverse black holes. In this work, we first obtain the Schrödinger-like master equation of the axial gravitational perturbation of a holonomy-corrected black hole. We calculate the corresponding quasinormal frequencies using the Wentzel-Kramers-Brillouin approximation and asymptotic iteration methods. We investigate the numerical evolution of an initial wave packet on the background spacetime. Then, we deduce the parameterized expression of the quasinormal frequencies and find that $r_0 \leq 10^{-2}$ is a necessary condition for the parameterized approximation to be valid. We also study the impact of the quantum parameter $r_0$ on the greybody factor and Hawking radiation. With more ringdown signals of gravitational waves detected in the future, our research will contribute to the study of the quantum properties of black holes.

gr-qc↗

Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities

With the rapid advancement of Multimodal Large Language Models (MLLMs), a variety of benchmarks have been introduced to evaluate their capabilities. While most evaluations have focused on complex tasks such as scientific comprehension and visual reasoning, little attention has been given to assessing their fundamental image classification abilities. In this paper, we address this gap by thoroughly revisiting the MLLMs with an in-depth analysis of image classification. Specifically, building on established datasets, we examine a broad spectrum of scenarios, from general classification tasks (e.g., ImageNet, ObjectNet) to more fine-grained categories such as bird and food classification. Our findings reveal that the most recent MLLMs can match or even outperform CLIP-style vision-language models on several datasets, challenging the previous assumption that MLLMs are bad at image classification \cite{VLMClassifier}. To understand the factors driving this improvement, we conduct an in-depth analysis of the network architecture, data selection, and training recipe used in public MLLMs. Our results attribute this success to advancements in language models and the diversity of training data sources. Based on these observations, we further analyze and attribute the potential reasons to conceptual knowledge transfer and enhanced exposure of target concepts, respectively. We hope our findings will offer valuable insights for future research on MLLMs and their evaluation in image classification tasks.

cs.CV↗

Energy Consumption of GEO-to-ground Beaconless Link Acquisition Against Random Vibration with Coherent Detection

The GEO satellite maintains good synchronization with the ground, reducing the priority of acquisition time in the establishment of the optical link. Whereas energy is an important resource for the satellite to execute space missions, the consumption during the acquisition process rises to the primary optimization objective. However, no previous studies have addressed this issue. Motivated by this gap, this paper first model the relationship between the transmitted power and the received SNR in the coherent detection system, with the corresponding single-field acquisition probability, the acquisition time is then calculated, and the closed-form expression of the multi-field acquisition energy consumption is further derived in scan-stare mode. Then for dual-scan technique, through the induction of the probability density function of acquisition energy, it is transformed into the equivalent form of scan-stare, thereby acquiring acquisition energy. Subsequently, optimizations are performed on these two modes. The above theoretical derivations are verified through Monte Carlo simulations. Consequently, the acquisition energy of dual-scan is lower than that of scan-stare, with the acquisition time being about half of the latter, making it a more efficient technique. Notably, the optimum beam divergence angle is the minimum that the laser can modulate, and the beaconless acquisition energy is only 6\% of that with the beacon, indicating that the beaconless is a better strategy for optical link acquisition with the goal of energy consumption optimization.

physics.optics↗

High-Accuracy Model Predictive Control with Inverse Hysteresis for High-Speed Trajectory Tracking of Piezoelectric Fast Steering Mirror

Piezoelectric fast steering mirrors (PFSM) are widely utilized in beam precision-pointing systems but encounter considerable challenges in achieving high-precision tracking of fast trajectories due to nonlinear hysteresis and mechanical dual-axis cross-coupling. This paper proposes a model predictive control (MPC) approach integrated with a hysteresis inverse based on the Hammerstein modeling structure of the PFSM. The MPC is designed to decouple the rate-dependent dual-axis linear components, with an augmented error integral variable introduced in the state space to eliminate steady-state errors. Moreover, proofs of zero steady-state error and disturbance rejection are provided. The hysteresis inverse model is then cascaded to compensate for the rate-independent nonlinear components. Finally, PFSM tracking experiments are conducted on step, sinusoidal, triangular, and composite trajectories. Compared to traditional model-free and existing model-based controllers, the proposed method significantly enhances tracking accuracy, demonstrating superior tracking performance and robustness to frequency variations. These results offer valuable insights for engineering applications.

eess.SY↗

First Principles based High-precision Modelling and Identification of Piezoelectric Fast Steering Mirror

We establish a high-precision composite model for a piezoelectric fast steering mirror (PFSM) using a Hammerstein structure. A novel asymmetric Bouc-Wen model is proposed to describe the nonlinear rate-independent hysteresis, while a dynamic model is derived to represent the linear rate-dependent component. By analyzing the physical process from the displacement of the piezoelectric actuator to the angle of the PFSM, cross-axis coupling is modeled based on first principles. Given the dynamic isolation of each module on different frequency scales, a step-by-step method for model parameter identification is carried out. Finally, experimental results demonstrate that the identified parameters can accurately represent the hysteresis, creep, and mechanical dynamic characteristics of the PFSM. Furthermore, by comparing the outputs of the identified model with the real PFSM under different excitation signals, the effectiveness of the proposed dual-input dual-output composite model is validated.

eess.SY↗

TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior

Recent advances in autonomous driving systems have shifted towards reducing reliance on high-definition maps (HDMaps) due to the huge costs of annotation and maintenance. Instead, researchers are focusing on online vectorized HDMap construction using on-board sensors. However, sensor-only approaches still face challenges in long-range perception due to the restricted views imposed by the mounting angles of onboard cameras, just as human drivers also rely on bird's-eye-view navigation maps for a comprehensive understanding of road structures. To address these issues, we propose to train the perception model to "see" standard definition maps (SDMaps). We encode SDMap elements into neural spatial map representations and instance tokens, and then incorporate such complementary features as prior information to improve the bird's eye view (BEV) feature for lane geometry and topology decoding. Based on the lane segment representation framework, the model simultaneously predicts lanes, centrelines and their topology. To further enhance the ability of geometry prediction and topology reasoning, we also use a topology-guided decoder to refine the predictions by exploiting the mutual relationships between topological and geometric features. We perform extensive experiments on OpenLane-V2 datasets to validate the proposed method. The results show that our model outperforms state-of-the-art methods by a large margin, with gains of +6.7 and +9.1 on the mAP and topology metrics. Our analysis also reveals that models trained with SDMap noise augmentation exhibit enhanced robustness.

cs.CV↗

Magnetoresistance oscillations in vertical junctions of 2D antiferromagnetic semiconductor CrPS$_4$

Magnetoresistance (MR) oscillations serve as a hallmark of intrinsic quantum behavior, traditionally observed only in conducting systems. Here we report the discovery of MR oscillations in an insulating system, the vertical junctions of CrPS$_4$ which is a two dimensional (2D) A-type antiferromagnetic semiconductor. Systematic investigations of MR peaks under varying conditions, including electrode materials, magnetic field direction, temperature, voltage bias and layer number, elucidate a correlation between MR oscillations and spin-canted states in CrPS$_4$. Experimental data and analysis point out the important role of the in-gap electronic states in generating MR oscillations, and we proposed that spin selected interlayer hopping of localized defect states may be responsible for it. Our findings not only illuminate the unusual electronic transport in CrPS$_4$ but also underscore the potential of van der Waals magnets for exploring interesting phenomena.

cond-mat.mes-hall↗

Artificial Intelligence-Enhanced Couinaud Segmentation for Precision Liver Cancer Therapy

Precision therapy for liver cancer necessitates accurately delineating liver sub-regions to protect healthy tissue while targeting tumors, which is essential for reducing recurrence and improving survival rates. However, the segmentation of hepatic segments, known as Couinaud segmentation, is challenging due to indistinct sub-region boundaries and the need for extensive annotated datasets. This study introduces LiverFormer, a novel Couinaud segmentation model that effectively integrates global context with low-level local features based on a 3D hybrid CNN-Transformer architecture. Additionally, a registration-based data augmentation strategy is equipped to enhance the segmentation performance with limited labeled data. Evaluated on CT images from 123 patients, LiverFormer demonstrated high accuracy and strong concordance with expert annotations across various metrics, allowing for enhanced treatment planning for surgery and radiation therapy. It has great potential to reduces complications and minimizes potential damages to surrounding tissue, leading to improved outcomes for patients undergoing complex liver cancer treatments.

eess.IV↗

RediSwap: MEV Redistribution Mechanism for CFMMs

Automated Market Makers (AMMs) are essential to decentralized finance, offering continuous liquidity and enabling intermediary-free trading on blockchains. However, participants in AMMs are vulnerable to Maximal Extractable Value (MEV) exploitation. Users face threats such as front-running, back-running, and sandwich attacks, while liquidity providers (LPs) incur the loss-versus-rebalancing (LVR). In this paper, we introduce RediSwap, a novel AMM designed to capture MEV at the application level and refund it fairly among users and liquidity providers. At its core, RediSwap features an MEV-redistribution mechanism that manages arbitrage opportunities within the AMM pool. We formalize the mechanism design problem and the desired game-theoretical properties. A central insight underpinning our mechanism is the interpretation of the maximal MEV value as the sum of LVR and individual user losses. We prove that our mechanism is incentive-compatible and Sybil-proof, and demonstrate that it is easy for arbitrageurs to participate. We empirically compared RediSwap with existing solutions by replaying historical AMM trades. Our results suggest that RediSwap can achieve better execution than UniswapX in 89% of trades and reduce LPs' loss to under 0.5% of the original LVR in most cases.

cs.GT↗

ChatHouseDiffusion: Prompt-Guided Generation and Editing of Floor Plans

The generation and editing of floor plans are critical in architectural planning, requiring a high degree of flexibility and efficiency. Existing methods demand extensive input information and lack the capability for interactive adaptation to user modifications. This paper introduces ChatHouseDiffusion, which leverages large language models (LLMs) to interpret natural language input, employs graphormer to encode topological relationships, and uses diffusion models to flexibly generate and edit floor plans. This approach allows iterative design adjustments based on user ideas, significantly enhancing design efficiency. Compared to existing models, ChatHouseDiffusion achieves higher Intersection over Union (IoU) scores, permitting precise, localized adjustments without the need for complete redesigns, thus offering greater practicality. Experiments demonstrate that our model not only strictly adheres to user specifications but also facilitates a more intuitive design process through its interactive capabilities.

cs.HC↗

Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning

Iterative preference learning, though yielding superior performances, requires online annotated preference labels. In this work, we study strategies to select worth-annotating response pairs for cost-efficient annotation while achieving competitive or even better performances compared with the random selection baseline for iterative preference learning. Built on assumptions regarding uncertainty and distribution shifts, we propose a comparative view to rank the implicit reward margins as predicted by DPO to select the response pairs that yield more benefits. Through extensive experiments, we show that annotating those response pairs with small margins is generally better than large or random, under both single- and multi-iteration scenarios. Besides, our empirical results suggest allocating more annotation budgets in the earlier iterations rather than later across multiple iterations.

cs.CL↗

MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction

The construction of Vectorized High-Definition (HD) map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationships between points and instances. In this work, we propose a simple yet efficient framework named MGMapNet (Multi-Granularity Map Network) to model map element with a multi-granularity representation, integrating both coarse-grained instance-level and fine-grained point-level queries. Specifically, these two granularities of queries are generated from the multi-scale bird's eye view (BEV) features using a proposed Multi-Granularity Aggregator. In this module, instance-level query aggregates features over the entire scope covered by an instance, and the point-level query aggregates features locally. Furthermore, a Point Instance Interaction module is designed to encourage information exchange between instance-level and point-level queries. Experimental results demonstrate that the proposed MGMapNet achieves state-of-the-art performance, surpassing MapTRv2 by 5.3 mAP on nuScenes and 4.4 mAP on Argoverse2 respectively.

cs.CV↗

Microwave interference from a spin ensemble and its mirror image in waveguide magnonics

We investigate microwave interference from a spin ensemble and its mirror image in a one-dimensional waveguide. Away from the mirror, the resonance frequencies of the Kittel mode (KM) inside a ferrimagnetic spin ensemble have sinusoidal shifts as the normalized distance between the spin ensemble and the mirror increases compared to the setup without the mirror. These shifts are a consequence of the KM's interaction with its own image. Furthermore, the variation of the magnon radiative decay into the waveguide shows a cosine squared oscillation and is enhanced twofold when the KM sits at the magnetic antinode of the corresponding eigenmode. We can finely tune the KM to achieve the maximum adsorption of the input photons at the critical coupling point. Moreover, by placing the KM in proximity to the node of the resonance field, its lifetime is extended to more than eight times compared to its positioning near the antinode.

physics.app-ph↗

Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving

The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregressive world model by addressing these challenges through the formulation of multiple probabilistic hypotheses. We propose LatentDriver, a framework models the environment's next states and the ego vehicle's possible actions as a mixture distribution, from which a deterministic control signal is then derived. By incorporating mixture modeling, the stochastic nature of decisionmaking is captured. Additionally, the self-delusion problem is mitigated by providing intermediate actions sampled from a distribution to the world model. Experimental results on the recently released close-loop benchmark Waymax demonstrate that LatentDriver surpasses state-of-the-art reinforcement learning and imitation learning methods, achieving expert-level performance. The code and models will be made available at https://github.com/Sephirex-X/LatentDriver.

cs.RO↗

Broad-line Region of the Quasar PG 2130+099. II. Doubling the Size Over Four Years?

Over the past three decades, multiple reverberation mapping (RM) campaigns conducted for the quasar PG 2130+099 have exhibited inconsistent findings with time delays ranging from $\sim$10 to $\sim$200 days. To achieve a comprehensive understanding of the geometry and dynamics of the broad-line region (BLR) in PG 2130+099, we continued an ongoing high-cadence RM monitoring campaign using the Calar Alto Observatory 2.2m optical telescope for an extra four years from 2019 to 2022. We measured the time lags of several broad emission lines (including He II, He I, H$β$, and Fe II) with respect to the 5100 Å continuum, and their time lags continuously vary through the years. Especially, the H$β$ time lags exhibited approximately a factor of two increase in the last two years. Additionally, the velocity-resolved time delays of the broad H$β$ emission line reveal a back-and-forth change between signs of virial motion and inflow in the BLR. The combination of negligible ($\sim$10%) continuum change and substantial time-lag variation (over two times) results in significant scatter in the intrinsic $R_{\rm Hβ}-L_{\rm 5100}$ relationship for PG 2130+099. Taking into account the consistent changes in the continuum variability time scale and the size of the BLR, we tentatively propose that the changes in the measurement of the BLR size may be affected by 'geometric dilution'.

astro-ph.GA↗

Studying Critical Parameters of Superconductor via Diamond Quantum Sensors

Critical parameters are the key to superconductivity research, and reliable instrumentations can facilitate the study. Traditionally, one has to use several different measurement techniques to measure critical parameters separately. In this work, we develop the use of a single species of quantum sensor to determine and estimate several critical parameters with the help of independent simulation data. We utilize the nitrogen-vacancy (NV) center in the diamond, which recently emerged as a promising candidate for probing exotic features in condensed matter physics. The non-invasive and highly stable nature provides extraordinary opportunities to solve scientific problems in various systems. Using a high-quality single-crystalline YBa$_{2}$Cu$_{4}$O$_{8}$ (YBCO) as a platform, we demonstrate the use of diamond particles and a bulk diamond to probe the Meissner effect. The evolution of the vector magnetic field, the $H-T$ phase diagram, and the map of fluorescence contour are studied via NV sensing. Our results reveal different critical parameters, including lower critical field $H_{c1}$, upper critical field $H_{c2}$, and critical current density $j_{c}$, as well as verifying the unconventional nature of this high-temperature superconductor YBCO. Therefore, NV-based quantum sensing techniques have huge potential in condensed matter research.

cond-mat.mes-hall↗

Demonstration of a variational quantum eigensolver with a solid-state spin system under ambient conditions

Quantum simulators offer the potential to utilize the quantum nature of a physical system to study another physical system. In contrast to conventional simulation, which experiences an exponential increase in computational complexity, quantum simulation cost increases only linearly with increasing size of the problem, rendering it a promising tool for applications in quantum chemistry. The variational-quantum-eigensolver algorithm is a particularly promising application for investigating molecular electronic structures. For its experimental implementation, spin-based solid-state qubits have the advantage of long decoherence time and high-fidelity quantum gates, which can lead to high accuracy in the ground-state finding. This study uses the nitrogen-vacancy-center system in diamond to implement the variational-quantum-eigensolver algorithm and successfully finds the eigenvalue of a specific Hamiltonian without the need for error-mitigation techniques. With a fidelity of 98.9% between the converged state and the ideal eigenstate, the demonstration provides an important step toward realizing a scalable quantum simulator in solid-state spin systems.

quant-ph↗