SearcharxivSearch

arXiv subjects

Hang Ruan

Publications and source records attributed to Hang Ruan.

12 recordsLinked to original sources

IPEC: Test-Time Incremental Prototype Enhancement Classifier for Few-Shot Learning

Metric-based few-shot approaches have gained significant popularity due to their relatively straightforward implementation, high interpret ability, and computational efficiency. However, stemming from the batch-independence assumption during testing, which prevents the model from leveraging valuable knowledge accumulated from previous batches. To address these challenges, we propose a novel test-time method called Incremental Prototype Enhancement Classifier (IPEC), a test-time method that optimizes prototype estimation by leveraging information from previous query samples. IPEC maintains a dynamic auxiliary set by selectively incorporating query samples that are classified with high confidence. To ensure sample quality, we design a robust dual-filtering mechanism that assesses each query sample based on both global prediction confidence and local discriminative ability. By aggregating this auxiliary set with the support set in subsequent tasks, IPEC builds progressively more stable and representative prototypes, effectively reducing its reliance on the initial support set. We ground this approach in a Bayesian interpretation, conceptualizing the support set as a prior and the auxiliary set as a data-driven posterior, which in turn motivates the design of a practical "warm-up and test" two-stage inference protocol. Extensive empirical results validate the superior performance of our proposed method across multiple few-shot classification tasks.

cs.LG

InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery

Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for real-world, open-world applications. However, existing methods often rely on pseudo-labeling, or two-stage clustering, which lack a principled mechanism to explicitly disentangle essential, category-defining signals from instance-specific noise. In this paper, we address this fundamental limitation by re-framing GCD from an information-theoretic perspective, grounded in the Information Bottleneck (IB) principle. We introduce InfoSculpt, a novel framework that systematically sculpts the representation space by minimizing a dual Conditional Mutual Information (CMI) objective. InfoSculpt uniquely combines a Category-Level CMI on labeled data to learn compact and discriminative representations for known classes, and a complementary Instance-Level CMI on all data to distill invariant features by compressing augmentation-induced noise. These two objectives work synergistically at different scales to produce a disentangled and robust latent space where categorical information is preserved while noisy, instance-specific details are discarded. Extensive experiments on 8 benchmarks demonstrate that InfoSculpt validating the effectiveness of our information-theoretic approach.

cs.CV

EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers

Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful representational capacity. However, fine-tuning such large models demands extensive GPU memory and prolonged training time, making them impractical for many real-world low-resource scenarios. To bridge this gap, we propose EfficientFSL, a query-only fine-tuning framework tailored specifically for few-shot classification with ViT, which achieves competitive performance while significantly reducing computational overhead. EfficientFSL fully leverages the knowledge embedded in the pre-trained model and its strong comprehension ability, achieving high classification accuracy with an extremely small number of tunable parameters. Specifically, we introduce a lightweight trainable Forward Block to synthesize task-specific queries that extract informative features from the intermediate representations of the pre-trained model in a query-only manner. We further propose a Combine Block to fuse multi-layer outputs, enhancing the depth and robustness of feature representations. Finally, a Support-Query Attention Block mitigates distribution shift by adjusting prototypes to align with the query set distribution. With minimal trainable parameters, EfficientFSL achieves state-of-the-art performance on four in-domain few-shot datasets and six cross-domain datasets, demonstrating its effectiveness in real-world applications.

cs.CV

ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining

Remote sensing image change detection is one of the fundamental tasks in remote sensing intelligent interpretation. Its core objective is to identify changes within change regions of interest (CRoI). Current multimodal large models encode rich human semantic knowledge, which is utilized for guidance in tasks such as remote sensing change detection. However, existing methods that use semantic guidance for detecting users' CRoI overly rely on explicit textual descriptions of CRoI, leading to the problem of near-complete performance failure when presented with implicit CRoI textual descriptions. This paper proposes a multimodal reasoning change detection model named ReasonCD, capable of mining users' implicit task intent. The model leverages the powerful reasoning capabilities of pre-trained large language models to mine users' implicit task intents and subsequently obtains different change detection results based on these intents. Experiments on public datasets demonstrate that the model achieves excellent change detection performance, with an F1 score of 92.1\% on the BCDD dataset. Furthermore, to validate its superior reasoning functionality, this paper annotates a subset of reasoning data based on the SECOND dataset. Experimental results show that the model not only excels at basic reasoning-based change detection tasks but can also explain the reasoning process to aid human decision-making.

cs.CV

RIS-Assisted Near-Field ISAC for Multi-Target Indication in NLoS Scenarios

Enabling multi-target sensing in near-field integrated sensing and communication (ISAC) systems is a key challenge, particularly when line-of-sight paths are blocked. This paper proposes a beamforming framework that leverages a reconfigurable intelligent surface (RIS) to achieve multi-target indication. Our contribution is the extension of classic beampattern gain and inter-target cross-correlation metrics to the near-field, leveraging both angle and distance information to discriminate between multiple users and targets. We formulate a problem to maximize the worst-case sensing performance by jointly designing the beamforming at the base station and the phase shifts at the RIS, while guaranteeing communication rates. The non-convex problem is solved via an efficient alternating optimization (AO) algorithm that utilizes semidefinite relaxation (SDR). Simulations demonstrate that our RIS-assisted framework enables high-resolution sensing of co-angle targets in blocked scenarios.

eess.SP

Near-Field Integrated Sensing and Communication for Multi-Target Indication

Integrated sensing and communication (ISAC) in the near-field regime offers the potential to jointly support high-rate downlink transmission and high-resolution multi-target detection by exploiting the spherical-wave nature of electromagnetic propagation. In this paper, we propose a unified beamforming framework for a multi-user multi-target near-field ISAC system. In this system, a multi-antenna base station simultaneously serves multiple single-antenna users and senses multiple point-targets without prior knowledge of their radar cross sections. By optimizing the transmit covariance matrix, our design maximizes the minimum weighted transmit beampattern gain across all targets to ensure accurate sensing while strictly limiting inter-target cross-correlations and guaranteeing per-user communication rate and total power constraints. We extend classical far-field beampattern and cross-correlation measures to the near-field by incorporating both angle and range dependencies, enabling discrimination of targets along the same direction but at different distances. The resulting non-convex program is efficiently relaxed to a semidefinite program via rank-one lifting. We then develop a closed-form reconstruction to recover optimal rank-one beamformers. Numerical simulations demonstrate that our near-field ISAC design can simultaneously resolve and serve users/targets along the same direction but at different distances, achieving significant gains over far-field and single-target benchmarks.

eess.SP

Task-Based Quantizer Design for Sensing With Random Signals

In integrated sensing and communication (ISAC) systems, random signaling is used to convey useful information as well as sense the environment. Such randomness poses challenges in various components in sensing signal processing. In this paper, we investigate quantizer design for sensing in ISAC systems. Unlike quantizers for channel estimation in massive multiple-input-multiple-out (MIMO) communication systems, sensing in ISAC systems needs to deal with random nonorthogonal transmitted signals rather than a fixed orthogonal pilot. Considering sensing performance and hardware implementation, we focus on task-based hardware-limited quantization with spatial analog combining. We propose two strategies of quantizer optimization, i.e., data-dependent (DD) and data-independent (DI). The former achieves optimized sensing performance with high implementation overhead. To reduce hardware complexity, the latter optimizes the quantizer with respect to the random signal from a stochastic perspective. We derive the optimal quantizers for both strategies and formulate an algorithm based on sample average approximation (SAA) to solve the optimization in the DI strategy. Numerical results show that the optimized quantizers outperform digital-only quantizers in terms of sensing performance. Additionally, the DI strategy, despite its lower computational complexity compared to the DD strategy, achieves near-optimal sensing performance.

cs.IT

Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime

Predictive combinatorial optimization, where the parameters of combinatorial optimization (CO) are unknown at the decision-making time, is the precise modeling of many real-world applications, including energy cost-aware scheduling and budget allocation on advertising. Tackling such a problem usually involves a prediction model and a CO solver. These two modules are integrated into the predictive CO pipeline following two design principles: "Predict-then-Optimize (PtO)", which learns predictions by supervised training and subsequently solves CO using predicted coefficients, while the other, named "Predict-and-Optimize (PnO)", directly optimizes towards the ultimate decision quality and claims to yield better decisions than traditional PtO approaches. However, there lacks a systematic benchmark of both approaches, including the specific design choices at the module level, as well as an evaluation dataset that covers representative real-world scenarios. To this end, we develop a modular framework to benchmark 11 existing PtO/PnO methods on 8 problems, including a new industrial dataset for combinatorial advertising that will be released. Our study shows that PnO approaches are better than PtO on 7 out of 8 benchmarks, but there is no silver bullet found for the specific design choices of PnO. A comprehensive categorization of current approaches and integration of typical scenarios are provided under a unified benchmark. Therefore, this paper could serve as a comprehensive benchmark for future PnO approach development and also offer fast prototyping for application-focused development. The code is available at https://github.com/Thinklab-SJTU/PredictiveCO-Benchmark.

cs.LG

Designing the Waveform Bandwidth and Time Duration of Automotive Radars for Better Collision Warning Performance

Automotive radar is a key component in an ADAS. The increasing number of radars implemented in vehicles makes interference between them a noteworthy issue. One method of interference mitigation is to limit the TBP of radar waveforms. However, the problems of how much TBP is necessary and how to optimally utilize the limited TBP have not been addressed. We take CWS as an example and propose a method of designing the radar waveform parameters oriented by the performance of CWS We propose a metric to quantify the CWS performance and study how the radar waveform parameters (bandwidth and duration) influence this metric. Then, the waveform parameters are designed with a limit on the TBP to optimize the system performance. Numerical results show that the proposed design outperforms the state-of-the-art parameter settings in terms of system performance and resource or energy efficiency.

eess.SP

Study of Joint MSINR and Relay Selection Algorithms for Distributed Beamforming

This paper presents joint maximum signal-to-interference-plus-noise ratio (MSINR) and relay selection algorithms for distributed beamforming. We propose a joint MSINR and restricted greedy search relay selection (RGSRS) algorithm with a total relay transmit power constraint that iteratively optimizes both the beamforming weights at the relays nodes, maximizing the SINR at the destination. Specifically, we devise a relay selection scheme that based on greedy search and compare it to other schemes like restricted random relay selection (RRRS) and restricted exhaustive search relay selection (RESRS). A complexity analysis is provided and simulation results show that the proposed joint MSINR and RGSRS algorithm achieves excellent bit error rate (BER) and SINR performances.

cs.IT

The Random Frequency Diverse Array: A New Antenna Structure for Uncoupled Direction-Range Indication in Active Sensing

In this paper, we propose a new type of array antenna, termed the Random Frequency Diverse Array (RFDA), for an uncoupled indication of target direction and range with low system complexity. In RFDA, each array element has a narrow bandwidth and a randomly assigned carrier frequency. The beampattern of the array is shown to be stochastic but thumbtack-like, and its stochastic characteristics, such as the mean, variance, and asymptotic distribution are derived analytically. Based on these two features, we propose two kinds of algorithms for signal processing. One is matched filtering, due to the beampattern's good characteristics. The other is compressive sensing, because the new approach can be regarded as a sparse and random sampling of target information in the spatial-frequency domain. Fundamental limits, such as the Cramér-Rao bound and the observing matrix's mutual coherence, are provided as performance guarantees of the new array structure. The features and performances of RFDA are verified with numerical results.

cs.IT

Robust Adaptive Beamforming Based on Low-Complexity Shrinkage-Based Mismatch Estimation

In this work, we propose a low-complexity robust adaptive beamforming (RAB) technique which estimates the steering vector using a Low-Complexity Shrinkage-Based Mismatch Estimation (LOCSME) algorithm. The proposed LOCSME algorithm estimates the covariance matrix of the input data and the interference-plus-noise covariance (INC) matrix by using the Oracle Approximating Shrinkage (OAS) method. LOCSME only requires prior knowledge of the angular sector in which the actual steering vector is located and the antenna array geometry. LOCSME does not require a costly optimization algorithm and does not need to know extra information from the interferers, which avoids direction finding for all interferers. Simulations show that LOCSME outperforms previously reported RAB algorithms and has a performance very close to the optimum.

cs.IT