SearcharxivSearch

arXiv subjects

Min Gao

Publications and source records attributed to Min Gao.

At least 19 recordsLinked to original sources

Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning

Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models toward higher-confidence reasoning states without external verifiers or reward models. However, higher confidence does not necessarily imply correctness, as LLMs may remain highly confident along incorrect reasoning trajectories. We observe that high-confidence reasoning is more likely to be correct when confidence remains stable under local perturbations. Based on this observation, we propose Test-Time Adaptation via Stability-Aware Confidence Optimization (TASCO), a framework that incorporates local stability into confidence-based test-time adaptation while keeping the LLM frozen. TASCO operationalizes local stability by optimizing a lightweight task-level prefix under two alternative perturbation strategies: Random Perturbation promotes distributional stability across trajectories induced by nearby perturbed prefixes, whereas Sharpness-Aware Perturbation targets worst-case local sensitivity. Experiments demonstrate that TASCO improves reasoning accuracy and token efficiency across diverse LLMs and reasoning benchmarks, while behavioral analyses show that it maintains stable confidence under local perturbations without prematurely concentrating the model's predictive distribution.

cs.AI

Lightweight Soft X-ray Imager (LSXI) with glass-based coded mask

The coded mask technique has been widely used in X-ray and Gamma-ray imagers, especially in space astronomy. However, the traditional design of coded mask imagers usually has problems with large size and heavy weight. Here, we propose a novel design for a coded mask imager made of glass based on microchannel plate (MCP) technology, making it lightweight (within 1 kg), compact, and self-supporting, which is very suitable for space exploration satellites. The design is initially demonstrated by reconstructing encoded patterns from measurements at the X-ray beamline. Monte Carlo simulations of the module integrated into a 6U CubeSat are conducted to assess its in-orbit detector performance, including sensitivity and source localization.

astro-ph.IM

Personalized Communication Skills for Agentic Recommender Systems

Agentic recommender systems increasingly employ large language model-based UserAgents to evaluate candidate items through simulated feedback before recommendations are delivered. However, existing UserAgents typically reason in isolation based on limited personal histories, which may lead to perspective narrowing: the agent evaluates candidates from a local and incomplete view, overlooks relevant preference facets, and consequently produces inaccurate judgments. A natural way to alleviate this problem is to introduce other users as advisor agents, whose diverse histories provide complementary evidence that helps the target user reconsider overlooked preference signals. Nevertheless, a generic user-advisor communication process is insufficient, as different user decision states require different forms of external advice. Based on this insight, we propose AgentCom, a personalized communication skill framework for agentic recommender systems. AgentCom organizes reusable communication skills into a shared why--what--how--who skill bank: why identifies the decision deficiency that necessitates communication, what specifies the information task, how determines the advisor interaction protocol, and who retrieves advisors capable of executing that protocol. To make the shared skill bank personalized at use time and adaptive over time, AgentCom introduces two complementary mechanisms: personalized skill routing and failure-driven skill evolution. Personalized skill routing constructs a communication path by sequentially selecting suitable skills for each user and recommendation context. Failure-driven skill evolution learns from unsuccessful communication cases and enriches the shared bank with reusable skills that address previously uncovered communication needs. Experiments show that AgentCom consistently improves recommendation performance across traditional, social, and agentic recommenders.

cs.IR

CCAT: Design and Characterization of the 350 GHz Instrument Module

The CCAT Collaboration's Prime-Cam instrument will soon be deployed to the Fred Young Submillimeter Telescope (FYST) in Chile's Atacama Desert. Featuring prominently in Prime-Cam's calibration and early science observations will be the 350 GHz instrument module, a broadband camera that will field more than 10,000 microwave kinetic inductance detectors (KIDs) across three detector arrays. Forecasts show this module will be capable of making the most sensitive to-date measurements of polarized dust emission over a large fraction of the sky at this frequency, enabling new galactic polarization science and improved understanding of cosmological foregrounds. In this work we discuss the design of the 350 GHz instrument module, covering aspects of the optics, readout, and detector arrays. We then report on the results of in-lab testing of the fully-integrated module, achieving stable cryogenic performance with a 100 mK focal plane, high detector yield, and a passband comparable to designed specifications. Upon completion of these tests, this module was shipped to the telescope site in Chile for integration in Prime-Cam.

astro-ph.IM

Where Reasoning Matters: Rethinking Latent Reasoning in Semantic ID-based Generative Recommendation

Semantic ID-based generative recommendation predicts an item by generating a short sequence of semantic ID tokens, where each token is produced autoregressively. Latent reasoning has recently been introduced to improve this process through additional hidden-state computation before each token decision. This raises a practical question: when one item is represented by a sequence of semantic ID tokens, should each token receive the same fixed number of latent refinement steps, or should these steps be allocated more effectively across positions? We study this question through position-wise information-gain (IG), which measures how much each semantic ID position reduces the uncertainty of the target item. We observe that earlier semantic ID positions usually provide higher information-gain, while later positions contribute less additional information. We further analyze that applying more refinement to high-IG positions tends to bring larger expected benefits. Based on this observation, we propose IBA, an Information-Gain Budget Allocation framework for semantic ID-based generative recommendation. IBA treats latent refinement steps as a limited computational resource and learns how to allocate them across semantic ID positions, assigning more refinement to informative positions and less to positions with smaller contribution. Experiments on multiple public datasets show that IBA consistently improves strong generative recommendation baselines and achieves a better accuracy--computation trade-off than fixed or poorly matched step allocations.

cs.IR

Three-Dimensional Retinal Microvasculature Restoration in OCT Angiography

Optical coherence tomographic angiography (OCTA) is a powerful technique for imaging retinal microvasculature. However, acquiring reliable quantification of retinal blood flow and areas of retinal nonperfusion is challenging because of imaging artifacts. Existing methods primarily focus on noise suppression, projection artifact removal, or signal enhancement to improve the image quality of OCTA in cross-sectional or two-dimensional (2D) en face projections, while neglecting the intrinsic three-dimensional vascular architecture. In this study, we propose a deep learning-based algorithm for restoring capillary anatomical vasculature from a single OCTA volume. The network consists of an EfficientNet-B5 encoder and a decoder incorporating concurrent spatial and channel squeeze-and-excitation modules, connected via skip connections to preserve spatial resolution. Three adjacent B-frames are used as input to predict the restored middle B-frame. We evaluated the performance of the model using the peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) against ground truth generated from averaging multiple scans. The results show that the proposed model significantly (both p < 0.001) improved image quality compared with the original single OCTA volume, with a PSNR of 26.16 +/- 1.26 vs. 22.23 +/- 0.78 and an SSIM of 0.91 +/- 0.02 vs. 0.72 +/- 0.03. The proposed model also significantly (p < 0.001) improved microvascular fidelity, measured by the Dice coefficient overlap between the model output and ground truth, in both 2D and 3D by at least 3.8% and 51.2%, respectively, across several different vascular slabs.

cs.CV

Deep Learning-assisted AMD Staging based on OCT and OCT Angiography

To develop and evaluate deep learning models for automated grading of age-related macular degeneration (AMD) severity using optical coherence tomography (OCT) and OCT angiography (OCTA) data. Two hundred seventy-one participants aged >= 50 years with varying AMD severities. Central macular 6 x 6 mm OCT/OCTA volumes were acquired using a swept-source OCTA system (SOLIX; Visionix/Optovue Inc., CA). AMD severity was graded into four stages (No AMD, Early AMD, Intermediate AMD, and Advanced AMD) according to the AREDS simplified severity scale. Three deep learning models were developed using different input modalities: (1) biomarker maps derived from segmented pathological features, including retinal fluid, drusen, geographic atrophy (GA), and macular neovascularization (MNV); (2) two-dimensional (2D) en face OCT and OCTA projections; and (3) three-dimensional (3D) OCT/OCTA volumes. EfficientNet-based architectures were trained using normalized inputs, data augmentation, and five-fold cross-validation. A total of 2,030 OCT/OCTA volumes from 351 eyes of 271 participants were analyzed. All models demonstrated strong AMD staging performance with substantial agreement with the reference standard (QWK >= 0.83). The biomarker-based model achieved the highest overall performance (QWK = 0.85 +/- 0.03, mean +/- standard deviation) and the best detection of early AMD (F1-score = 0.59 +/- 0.14). The 3D model achieved performance comparable to the 2D OCT/OCTA model (QWK = 0.83 +/- 0.04 vs. 0.83 +/- 0.09), while the 2D OCT/OCTA model showed the highest precision (0.79 +/- 0.06) and most accurately identified eyes without AMD. Deep learning models using OCT/OCTA data can accurately and automatically grade AMD severity. Among the evaluated approaches, the biomarker-based model provided the most balanced performance and showed particular value for early AMD detection.

cs.CV

Robust and Reliable AI for Predictive Quality in Semiconductor Materials Manufacturing with MLOps and Uncertainty Quantification

Semiconductor materials manufacturing presents unique challenges for machine learning deployment due to evolving process conditions, equipment degradation, and raw material variability that can cause model performance deterioration over time. This study benchmarks machine learning operations (MLOps) retraining strategies using five years of real manufacturing data to identify optimal retraining approaches for quality prediction. We evaluate various retraining frequencies and hyperparameter optimization strategies using control limit normalized residuals as key performance metric. Results demonstrate that a fixed retraining cadence every five production batches without hyperparameter retuning achieves superior performance across all drift conditions while significantly reducing computational overhead compared to strategies incorporating hyperparameter optimization. This approach effectively maintains model accuracy during both abrupt process changes and gradual equipment degradation patterns. To address the critical need for uncertainty quantification in manufacturing decision-making, we implement conformal prediction to generate prediction confidence intervals with strong statistical guarantees. This enables proactive quality control by identifying when prediction intervals fall within acceptable control limits, transforming traditional reactive quality management into a predictive framework. The findings provide practical guidelines for implementing robust MLOps strategies in manufacturing environments where computational efficiency and reliable uncertainty quantification are paramount for operational success.

cs.LG

Disagreement as Signals: Dual-view Calibration for Sequential Recommendation Denoising

Sequential recommendation seeks to model the evolution of user interests by capturing temporal user intent and item-level transition patterns. Transformer-based recommenders demonstrate a strong capacity for learning long-range and interpretable dependencies, yet remain vulnerable to behavioral noise that is misaligned with users' true preferences. Recent large language model (LLM)-based approaches attempt to denoise interaction histories through static semantic editing. Such methods neglect the learning dynamics of recommendation models and fail to account for the evolving nature of user interests. To address this limitation, we propose a Dual-view Calibration framework for Sequential Recommendation denoising (DC4SR). Specifically, we introduce a semantic prior, derived from an LLM fine-tuned via labeled historical interactions, to estimate the noise distribution from a semantic perspective. From the learning perspective, we further employ a model-side posterior that infers the noise distribution based on the model's learning dynamics. The disagreement between the two distributions is then leveraged to jointly refine semantic understanding and learning-aware model-side representations. Through iterative updates, dynamic dual-view calibration is achieved for both the global semantic prior and the model-side posterior, enabling consistent alignment with evolving user interests. Extensive experiments demonstrate that DC4SR consistently outperforms strong Transformer-based recommenders and LLM-based denoising methods, exhibiting enhanced robustness across training stages and noise conditions.

cs.IR

GRM Scientific Pipeline

The Gamma-Ray Monitor (GRM) is a key payload of the Space-based multiband astronomical Variable Objects Monitor (SVOM) mission, which is designed to detect gamma ray bursts (GRBs) within the energy range of 15 keV to 5 MeV. The GRM Instrument Center (GRM\_IC) features real-time data processing through the X-band, enabling rapid response of high-energy GRB events. The system employs an event-driven architecture and distributed design, achieving efficient processing and real-time monitoring of massive observational data. Through comprehensive data production processes and scientific data product management, the system achieves efficient production of scientific data products of the L1B / C level through the submission of jobs to the task scheduling system. Through modular architecture design and automated processing workflow, the GRM data processing system realizes precise conversion and scientific analysis of GRB detection data, providing robust technical support for future system upgrades and cross-platform collaboration.

astro-ph.IM

Green-Red Watermarking for Recommender Systems

The widespread open-sourcing of advanced recommendation algorithms and the rising threat of model extraction attacks have made safeguarding the intellectual property of recommender systems an imperative task. While watermarking serves as a potent defense, existing methods primarily rely on forcing models to memorize pre-defined interaction patterns. Such memorization-based approaches often require excessive synthetic data injection and are vulnerable to removal attacks due to their detectable statistical deviations from natural user behavior. To address these limitations, we propose GREW, a novel Green-REd Watermarking framework for recommender systems. GREW leverages a secret key to partition the item space into "green" items for soft promotion and "red" items as anchors, thereby shifting the paradigm from fragile memorization to a stealthy, key-controlled output bias. By integrating watermark signals directly into the intrinsic ranking process, GREW employs three recommendation-tailored modules: (1) Semantic-Consistent Hashing, which utilizes the secret key to cluster green items for performance-aware stealthiness; (2) Decision-Aligned Masking, which confines signal injection to the competitive item subset to preserve ranking logic; and (3) Confidence-Aware Scaling, which dynamically modulates injection intensity based on model uncertainty. Ownership verification is performed via statistical hypothesis testing on aggregated black-box outputs, enabled by the keyed re-partitioning of the item space. Experiments on multiple base models demonstrate that GREW achieves strong ownership verification and robustness against extraction attacks compared to existing baselines while requiring no data injection. Our code is available at https://github.com/Loche2/GREW.

cs.IR

The Gamma-Ray Monitor onboard the SVOM satellite

The Gamma-Ray Monitor (GRM) is a key scientific payload onboard the Space-based Multi-band Variable Object Monitor (SVOM) satellite, designed specifically for the detection and study of gamma-ray bursts (GRBs). Launched into a 625 km low-Earth orbit on 22 June 2024, GRM serves as a large-area, wide-field-of-view instrument capable of observing the hard X-ray and soft gamma-ray emissions in the energy range of 15 keV to 5 MeV. Its primary scientific objectives include: promptly triggering and localizing GRBs (with particular sensitivity to short-hard GRBs), measuring spectral and temporal properties of bursts, monitoring charged particle fluxes in orbit. GRM successfully detected its first GRB (GRB 240627B) on 27 June 2024, and has since maintained a detection rate of more than 100 GRBs per year. Cross-instrument comparisons with detectors such as GECAM and Fermi/GBM have validated the performance and data quality of GRM. This paper provides a comprehensive overview of GRM instrument design, reliability verification through ground testing, in-orbit triggering and localization algorithms, performance calibration, and preliminary in-orbit results, demonstrating its capability as a versatile gamma-ray all-sky monitor.

astro-ph.IM

Design and preliminary performance study of the broad-band spectrometer detector for POLAR-2

POLAR-2, the successor of the POLAR experiment aboard China's Tiangong-2 space lab, is set to be deployed on the China Space Station. The POLAR-2 mission aims to conducting high-precision polarization measurements of high-energy transients with a primary focus on Gamma-Ray Bursts (GRBs), following POLAR's pioneering accurate polarization measurements of GRB prompt emission. One of the key advancements in POLAR-2 is the inclusion of a dedicated Broad-band Spectrometer Detector (BSD) instrument, designed to provide precise measurements of GRB location and spectral parameters, which are critical inputs for accurate polarization analysis of POLAR-2's dedicated High-energy Polarimetry Detector (HPD), which is made of plastic scintillator bars array. BSD employs a coded-aperture mask imaging technique and pixelated GAGG scintillation crystals, offering a wide half-coded field of view of ~132{\deg} x 125{\deg} and an operational energy range of 10-1000 keV. Simulation results indicate that the instrument can achieve a localization accuracy of approximately 1.5{\deg} for faint GRBs similar to GRB 170817A, satisfying the core requirements of GRB polarimetry with HPD. BSD also has moderate capability for GRB polarimetry, particularly at several hundred keV energy. This paper outlines the preliminary design of BSD and presents an overall evaluation of its expected scientific performance, based on extensive Monte Carlo simulations and preliminary ground-based calibration tests.

astro-ph.IM

Study on the detector energy response of SVOM/GRM

The SVOM mission is specifically designed to for the detection and localization of Gamma-Ray Bursts (GRBs) and subsequent follow-up observations. Among the four telescopes installed on the SVOM satellite, the Gamma-Ray Monitor (GRM) plays a crucial role in capturing the prompt emission of GRBs due to its wide field of view (FOV) and broad energy range. Accurate determination of the detector's energy response is vital for analyzing GRM data, particularly considering the significant impact of the atmospheric albedo effect on this response. This research focuses on deriving the detector's energy response and establishing a calibration database for the GRM, with particular emphasis on investigating the atmospheric albedo effect. The study shows that the contribution of albedo photons to the detector's effective area depends strongly on the orientation of the GRD line of sight (LoS) relative to Earth and on the incident direction of the GRB. When the GRD LoS is anti-Earth oriented, the albedo effect is minimal, with the highest proportion of albedo effective area accounting for approximately 10% of the total effective area. This occurs when the incident angle of the GRB is nearly perpendicular to the LoS. Conversely, if the GRD LoS is not pointing away from Earth and the GRB arrives from angles greater than about 90$^{\circ}$, the albedo component can become predominant, contributing up to around 100% of the total effective area. This is especially pronounced in the 8-20 keV range, where the direct effective area drops to zero due to the large GRB injection angle. Our results show that, it is necessary for GRM to consider the atmospheric albedo effects in detector response, otherwise the spectral and localization analyses will result in biased measurements.

astro-ph.HE

Multi-LLM Token Filtering and Routing for Sequential Recommendation

Large language models (LLMs) have recently shown promise in recommendation by providing rich semantic knowledge. While most existing approaches rely on external textual corpora to align LLMs with recommender systems, we revisit a more fundamental yet underexplored question: Can recommendation benefit from LLM token embeddings alone without textual input? Through a systematic empirical study, we show that directly injecting token embeddings from a single LLM into sequential recommenders leads to unstable or limited gains, due to semantic misalignment, insufficient task adaptation, and the restricted coverage of individual LLMs. To address these challenges, we propose MLTFR, a Multi-LLM Token Filtering and Routing framework for corpus-free sequential recommendation. MLTFR follows an interaction-guided LLM knowledge integration paradigm, where task-relevant token embeddings are selected via user-guided token filtering to suppress noisy and irrelevant vocabulary signals. To overcome the limitations of single-LLM representations, MLTFR integrates multiple LLM token spaces through a Mixture-of-Experts architecture, with a Fisher-weighted semantic consensus expert to balance heterogeneous experts and prevent domination during training. By jointly filtering informative tokens and aggregating complementary semantic knowledge across multiple LLMs, MLTFR enables stable and effective utilization of LLM token embeddings without textual inputs or backbone modification. Extensive experiments demonstrate that MLTFR consistently outperforms state-of-the-art sequential recommendation baselines and existing alignment methods. Our code is available at: https://github.com/ccwwhhh/MLTFR.

cs.IR

The trigger and localization system of SVOM-GRM

The Space multi-band Variable Object Monitor (SVOM) is an astronomical satellite jointly developed by China and France, primarily focused on the detection of gamma-ray bursts (GRBs) and transient sources. The SVOM satellite was launched on 22nd June, 2024 with four payloads installed onboard. As one of payload, GRM comprises 3 gamma-ray detectors (each detector has an effective area of approximately 200~cm$^{2}$) with distinct pointing directions, enabling the temporal and spectral measurements as well as localization of GRBs in the energy range of 15-5000 keV. This article firstly introduces the on-board localization algorithm design for GRM and presents preliminary test results. Then, leveraging abundant ground-based computational resources, a joint fitting method for spectral and localization analysis using Monte Carlo Markov Chain (MCMC) is implemented. In contrast to the on-board localization algorithm, the on-ground MCMC method comprehensively considers the influence of spectral characteristics, thereby mitigating systematic biases. Finally, a systematic analysis based on this method is provided, highlighting the localization and spectral measurement capabilities of GRM. The preliminary localization analysis result for the on-board detected GRB 240629A by both GRM and Fermi/GBM shows that the localization result (error$\sim$4.14$^{\circ}$) of GRM is consistent with the Fermi/GBM result.

astro-ph.IM

Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems

Large language model-empowered agentic recommender systems (ARS) reformulate recommendation as a multi-turn interaction between a recommender agent and a user agent, enabling iterative preference elicitation and refinement beyond conventional one-shot prediction. However, existing ARS are mainly optimized in a Reflexion-style paradigm, where past interaction trajectories are stored as textual memory and retrieved as prompt context for later reasoning. Although this design allows agents to recall prior feedback and observations, the accumulated experience remains external to model parameters, leaving agents reliant on generic reasoning rather than progressively acquiring recommendation-specific decision-making ability through learning. Reinforcement learning (RL) therefore provides a natural way to internalize such interaction experience into parameters. Yet existing RL methods for ARS still suffer from two key limitations. First, they fail to capture the interactive nature of ARS, in which the recommender agent and the user agent continuously influence each other and can naturally generate endogenous supervision through interaction feedback. Second, they reduce a rich multi-turn interaction process to final outcomes, overlooking the dense supervision embedded throughout the trajectory. To this end, we propose CoARS, a self-distilled reinforcement learning framework for co-evolving agentic recommender systems. CoARS introduces two complementary learning schemes: interaction reward, which derives coupled task-level supervision for the recommender agent and the user agent from the same interaction trajectory, and self-distilled credit assignment, which converts historical trajectories into token-level credit signals under teacher-student conditioning. Experiments on multiple datasets show that CoARS outperforms representative ARS baselines in recommendation performance and user alignment.

cs.IR

Comprehensive Measurement of Spectral Evolution in a GRB Flare: High Time-Resolution Insights into the "Double-Tracking" Phenomenon

The spectral evolution characteristics of the prompt emission in gamma-ray bursts (GRBs) have been extensively studied, but detailed investigations of spectral evolution in a GRB flare remain lacking. In this work, we present the first analysis of spectral parameter evolution in a GRB flare through high time-resolved spectral fitting of the Brightest Flare in GRB 221009A. We find that the $\alpha$-Flux, $E_p$-Flux, and $E_p$-$\alpha$ relationships during both the overall phase and the rise phase of flare can be well described by simple power-law model, showing positive correlations. Therefore, we conclude that Brightest Flare exhibits "Double-tracking" behavior. Since values of $\alpha$ do not exceed the synchrotron "death line" (-2/3), we explain this phenomenon using a magnetic dissipation synchrotron radiation model. In the decay phase of flare, the $E_p$-Flux and $E_p$-$\alpha$ correlations become notably flatter, with their power-law indices decreasing significantly compared to those in the rise phase. This may be due to the fact that the next flare begins to erupt before the Brightest Flare has completely ended, resulting in the combined effects of both two flares. Our study of spectral parameter relations of the Brightest Flare provides new insights into the radiation mechanisms of both GRB prompt emission and flares.

astro-ph.HE