SearcharxivSearch

arXiv subjects

Yangyang Li

Publications and source records attributed to Yangyang Li.

At least 19 recordsLinked to original sources

Synergising Local Geo-Environmental Characteristics with Spatial Context for Enhancing Landslide Susceptibility Mapping

Data-driven methods are widely used in landslide susceptibility mapping (LSM) because they can effectively model the complex relationships between landslides and geo-environmental conditions. Existing data-driven approaches generally follow two types of data representations. Pixel-based models focus solely on the geo-environmental characteristics of a specific landslide but neglect the influence of its surrounding environment. Patch-based models incorporate surrounding spatial context but may include pixels with weak or no spatial relevance to the target landslide location. To address this limitation, this study proposes a Local-Geo and Spatial Context Fusion (LGSCF) strategy, which synergises the geo-environmental characteristics of landslide points with their corresponding spatial context through a feature-wise modulation mechanism. We tested the LGSCF strategy by integrating it into several representative convolutional neural network (CNN) architectures, creating nine different LGSCF-based models. The primary study area covers approximately 2644 km2 across Jenai and Sinyi Townships in Nantou County, Taiwan, and the dataset comprises 5332 landslide samples and an equal number of non-landslide samples. The results show that LGSCF-based models consistently outperform their corresponding baselines, achieving F1-scores up to 87.09% and AUC values up to 0.9472. Furthermore, the susceptibility maps produced by LGSCF-based models show that known landslides are more accurately concentrated in "very high" susceptibility zones with fewer misclassifications. These findings demonstrate that our fusion strategy can significantly improve the accuracy of landslide susceptibility mapping.

cs.CV

RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation

Language-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as streams and profiles must be incrementally updated under limited capacity. This requires maintaining a consistent bounded profile state and constructing profile-targeted semantic feedback from industrial implicit behavior logs. We propose RECAP, an offline closed-loop framework for optimizing streaming structured semantic profiles with historical recommendation feedback. RECAP maintains each profile as a bounded structured memory by combining LLM-based semantic updates with deterministic lifecycle and capacity control. RECAP constructs profile-targeted semantic feedback by filtering label-consistent behavior pairs with an LLM judge and training a dual-tower evaluator whose matching score serves as a GRPO reward. Experiments on Kuaishou short-video data show that RECAP improves uAUC by 0.0084 and Recall@2000 by about 4.9% over the base generator. Further analyses confirm the benefits of feedback construction and policy optimization, and show more grounded refinement and user-level abstraction in profile updates. A seven-day online A/B test further shows a statistically significant 0.139% improvement in average application usage time per user.

cs.IR

Enhancing Photometric Redshift Estimation for LSST with a Hybrid LSTM-Mixture Density Network

Accurate photometric redshift (photo-$z$) estimation and robust uncertainty quantification are essential for the LSST to achieve its precision cosmology goals. Traditional machine learning algorithms are largely restricted to point estimates, struggling to characterize the multimodal nature of redshift PDFs and the degeneracies within the color-redshift space. To address this, we present and validate the LSTM-MDNz architecture, which integrates sequential feature extraction with flexible probability density modeling to enhance both prediction accuracy and uncertainty calibration across a broad redshift range, thereby meeting the stringent data quality requirements necessitated by next-generation cosmological analysis. The LSTM-MDNz framework treats multi-band photometry as wavelength-ordered sequences, utilizing LSTM networks to capture non-linear evolutionary correlations across the SED. A Mixture Density Network (MDN) is then employed to explicitly model posterior PDFs via Gaussian mixture models (GMMs). Performance is evaluated on the HSC GalaxiesML dataset (which serves as a small-scale proxy for next-generation surveys like LSST) and benchmarked against the BNN architecture established by Jones et al. (2024). The proposed model consistently outperforms the BNN baseline, achieving a $\sim 10\%$ improvement in point-estimation accuracy (specifically across RMSE, MAE, scatter, and $\sigma_{\text{NMAD}}$) and a $\sim 20\%$ reduction in the rates of both general and catastrophic outliers. A uniform probability integral transform (PIT) distribution confirms well-calibrated probabilistic outputs. Furthermore, the PDF-based confidence metric $z_{\text{conf}}$ enables high-purity catalog construction: excluding just approximately $4\%$ of extremely low-confidence ($z_{\text{conf}} < 0.05$) samples reduces the overall outlier rate by $\sim 48\%$.

astro-ph.GA

From Empathy to Personalized Empathy: Adapting Empathetic Strategies to Individual Users

As Large Language Models (LLMs) are increasingly deployed in long-term interactions with users, empathy has become an increasingly important capability. However, existing research overlooks the influence of users' personality traits on empathetic strategies during long-term interactions. To address this gap, we introduce the task of personalized empathy, which focuses on adapting empathetic strategies according to users' personalized characteristics derived from history. To study and enhance this capability, we construct PersonaEmp, a personalized empathy dataset built from long-term user-AI interactions, featuring rich user histories, persona information, and empathy-seeking queries. We further propose PereGRM, a reward modeling framework that combines the empathy evaluation structure with dynamic evaluation criteria generation for fine-grained reward modeling. Experimental results across different settings and multiple judge models show that PereGRM consistently achieves the strongest performance improvements, indicating its effectiveness for enhancing personalized empathetic capabilities.

cs.CL

Multimodal Emotion Recognition with Large Language Models

Multimodal Emotion Recognition (MER) focuses on identifying and interpreting emotions from modality-compound inputs. Closely mirroring human cognitive processes in real-world environments, MER has drawn substantial attention from both academia and industry. Recently, a paradigm shift has been unveiled in MER, from leveraging small-scale, task-specific models to Large Language Models (LLMs). We refer to the latter as the MER-with-LLMs paradigm, which offers unprecedented generality, spurring numerous empirical attempts, even alongside speculation about LLMs' potential to achieve general emotional intelligence. However, with these new opportunities come new challenges, including the scarcity of emotionally annotated data, the affective gap both within and across modalities, and the opacity of affective interpretation. To systematically review existing research and guide future exploration, this paper categorizes prior works according to their focus on addressing these challenges into three directions: Affective Data Augmentation, Multimodal Affective Representation, and Multimodal Affective Reasoning. By thoroughly tracing the development, emerging trends, and remaining issues within each direction, this paper aims to provide a clear academic map of the MER-with-LLMs paradigm and foster its structured advancement.

cs.MM

Rigidity and flexibility under spectral Ricci lower bounds and mean-convex boundary

We study Riemannian manifolds $(M^n,g)$ with mean-convex boundary whose Ricci curvature is nonnegative in a spectral sense. Our first main result is a sharp spectral extension of a rigidity theorem by Kasue: we prove that under the conditions \[ \lambda_1(-\gamma\Delta+\mathrm{Ric})\geq 0,\qquad H_{\partial M}\geq 0, \] and in the sharp range $0\leq \gamma<4$ if $n=2$, and $0\leq\gamma<\frac{n-1}{n-2}$ if $n\geq3$, a (possibly noncompact) complete manifold with disconnected boundary, with at least one compact boundary component, must split isometrically as a product $[0,L]\times \Sigma$. Our second main contribution is a topological rigidity result for the relative fundamental group $\pi_1(M,\partial M)$, combined with a deep theorem of Lawson--Michelsohn. We prove that, in dimensions $n\neq4$, any compact manifold with boundary satisfying the two inequalities above, with at least one of them strict, admits a metric with positive sectional curvature and strictly mean-convex boundary, provided $\gamma\geq0$ if $n=2$, and $0\leq\gamma\leq\frac{n-1}{n-2}$ if $n\geq3$. This range of $\gamma$ is sharp for the latter result to hold.

math.DG

Breaking User-Centric Agency: A Tri-Party Framework for Agent-Based Recommendation

Large language models (LLMs) have spurred interest in agent-based recommender systems, yet most agentic approaches remain user-centric: items stay passive entities whose exposure is a by-product of relevance ranking, which exacerbates exposure concentration and long-tail under-representation. We break this user-centric allocation of agency with a Tri-party LLM-agent Recommendation framework (TriRec). Responsibility is split deliberately: Stage 1 has each item generate self-promotion conditioned on the target user, which lowers cold-start barriers, while the exposure budget stays with the platform, whose Stage 2 sequential re-ranker balances relevance, item utility, and exposure fairness. On four public datasets TriRec improves accuracy, fairness, and item-level utility, with the accuracy gain significant on three of the four. A three-arm ablation at 50 candidates separates two levels of the mechanism on items that received no exposure during training: self-promotion drives the accuracy gain, and conditioning it on the target user adds further exposure, together raising these items' share of top-ranked exposure by 43.6% relative. Restricting promotions to catalogue-verifiable attributes cuts strong exaggeration to 0.5%/2.0% on two datasets while retaining 88.6%/92.0% of the accuracy gain. Our code is available at https://github.com/Marfekey/TriRec.

cs.IR

Beyond Colors: Probing Redshifts from Galaxy Morphology in Single-band Images with ViT-MDNz

To address the challenge of estimating redshifts when only single-band images are available, this study introduces a deep learning model named ViT-MDNz. Leveraging robust statistical priors learned from large-scale data concerning the correlation between redshift and morphology, the model can directly estimate redshifts and their associated uncertainties from single-band galaxy images. It integrates a Vision Transformer (ViT) to extract deep morphological features and a Mixture Density Network (MDN) to predict the full redshift probability density function. Trained and evaluated on approximately 300,000 single-band images from the DESI Legacy Imaging Surveys (DESI-LS), the model achieves a normalized median absolute deviation $\sigma_{\rm NMAD} = 0.034$ and an outlier fraction $f_{\rm out} = 2.6\%$ in the $r$-band for redshifts up to $z \lesssim 1$. Evaluations using probability integral transform (PIT) and continuous ranked probability score (CRPS) confirm that the predicted probability density functions are well calibrated and closely match the true distribution. These results demonstrate that competitive redshift estimates can be obtained using morphological features alone, and that incorporating color information further enhances the accuracy and robustness of the estimation. Therefore, ViT-MDNz provides a practical approach for redshift estimation of galaxy samples with limited photometric band coverage, contributing to improved completeness and usability of redshift catalogs for future large-scale surveys such as DESI and LSST.

astro-ph.GA

Dynamic Graph Structure Learning via Resistance Curvature Flow

Geometric Representation Learning (GRL) aims to approximate the non-Euclidean topology of high-dimensional data through discrete graph structures, grounded in the manifold hypothesis. However, traditional static graph construction methods based on Euclidean distance often fail to capture the intrinsic curvature characteristics of the data manifold. Although Ollivier-Ricci Curvature Flow (OCF) has proven to be a powerful tool for dynamic topological optimization, its core reliance on Optimal Transport (Wasserstein distance) leads to prohibitive computational complexity, severely limiting its application in large-scale datasets and deep learning frameworks. To break this bottleneck, this paper proposes a novel geometric evolution framework: Resistance Curvature Flow (RCF). Leveraging the concept of effective resistance from circuit physics, RCF transforms expensive curvature optimization into efficient matrix operations. This approach achieves over 100x computational acceleration while maintaining geometric optimization capabilities comparable to OCF. We provide an in-depth exploration of the theoretical foundations and dynamical principles of RCF, elucidating how it guides the redistribution of edge weights via curvature gradients to eliminate topological noise and strengthen local cluster structures. Furthermore, we provide a mechanistic explanation of RCF's role in manifold enhancement and noise suppression, as well as its compatibility with deep learning models. We design a graph optimization algorithm, DGSL-RCF, based on this framework. Experimental results across deep metric learning, manifold learning, and graph structure learning demonstrate that DGSL-RCF significantly improves representation quality and downstream task performance.

cs.LG

An enumerative min-max theorem for minimal surfaces

We prove an enumerative min-max theorem that relates the number of genus g minimal surfaces in 3-manifolds of positive Ricci curvature to topological properties of the set of embedded surfaces of genus $\leq g$, possibly with finitely many singularities. This completes a central component of our program of using topological methods to enumerating minimal surfaces with prescribed genus. As an application, we show that every 3-sphere of positive Ricci curvature contains at least 4 embedded minimal surfaces of genus 2.

math.DG

BALNet: Deep Learning-Based Detection and Measurement of Broad Absorption Lines in Quasar Spectra

Broad absorption line (BAL) quasars serve as critical probes for understanding active galactic nucleus (AGN) outflows, black hole accretion, and cosmic evolution. To address the limitations of manual classification in large-scale spectroscopic surveys - where the number of quasar spectra is growing exponentially - we propose BALNet, a deep learning approach consisting of a one-dimensional convolutional neural network (1D-CNN) and bidirectional long short-term memory (Bi-LSTM) networks to automatically detect BAL troughs in quasar spectra. BALNet enables both the identification of BAL quasars and the measurement of their BAL troughs. We construct a simulated dataset for training and testing by combining non-BAL quasar spectra and BAL troughs, both derived from SDSS DR16 observations. Experimental results in the testing set show that: (1) BAL trough detection achieves 83.0% completeness, 90.7% purity, and an F1-score of 86.7%; (2) BAL quasar classification achieves 90.8% completeness and 94.4% purity; (3) the predicted BAL velocities agree closely with simulated ground truth labels, confirming BALNet's robustness and accuracy. When applied to the SDSS DR16 data within the redshift range 1.5<z<5.7, at least one BAL trough is detected in 20.4% of spectra. Notably, more than a quarter of these are newly identified sources with significant absorption, 8.8% correspond to redshifted systems, and some narrow/weak absorption features were missed. BALNet greatly improves the efficiency of large-scale BAL trough detection and enables more effective scientific analysis of quasar spectra.

astro-ph.GA

Efficient Curvature-aware Graph Network

Graph curvature provides geometric priors for Graph Neural Networks (GNNs), enhancing their ability to model complex graph structures, particularly in terms of structural awareness, robustness, and theoretical interpretability. Among existing methods, Ollivier-Ricci curvature has been extensively studied due to its strong geometric interpretability, effectively characterizing the local geometric distribution between nodes. However, its prohibitively high computational complexity limits its applicability to large-scale graph datasets. To address this challenge, we propose a novel graph curvature measure--Effective Resistance Curvature--which quantifies the ease of message passing along graph edges using the effective resistance between node pairs, instead of the optimal transport distance. This method significantly outperforms Ollivier-Ricci curvature in computational efficiency while preserving comparable geometric expressiveness. Theoretically, we prove the low computational complexity of effective resistance curvature and establish its substitutability for Ollivier-Ricci curvature. Furthermore, extensive experiments on diverse GNN tasks demonstrate that our method achieves competitive performance with Ollivier-Ricci curvature while drastically reducing computational overhead.

cs.LG

When Semantics Connect the Swarm: LLM-Driven Fuzzy Control for Cooperative Multi-Robot Underwater Coverage

Underwater multi-robot cooperative coverage remains challenging due to partial observability, limited communication, environmental uncertainty, and the lack of access to global localization. To address these issues, this paper presents a semantics-guided fuzzy control framework that couples Large Language Models (LLMs) with interpretable control and lightweight coordination. Raw multimodal observations are compressed by the LLM into compact, human-interpretable semantic tokens that summarize obstacles, unexplored regions, and Objects Of Interest (OOIs) under uncertain perception. A fuzzy inference system with pre-defined membership functions then maps these tokens into smooth and stable steering and gait commands, enabling reliable navigation without relying on global positioning. Then, we further coordinate multiple robots by introducing semantic communication that shares intent and local context in linguistic form, enabling agreement on who explores where while avoiding redundant revisits. Extensive simulations in unknown reef-like environments show that, under limited sensing and communication, the proposed framework achieves robust OOI-oriented navigation and cooperative coverage with improved efficiency and adaptability, narrowing the gap between semantic cognition and distributed underwater control in GPS-denied, map-free conditions.

cs.RO

Never Too Rigid to Reach: Adaptive Virtual Model Control with LLM- and Lyapunov-Based Reinforcement Learning

Robotic arms are increasingly deployed in uncertain environments, yet conventional control pipelines often become rigid and brittle when exposed to perturbations or incomplete information. Virtual Model Control (VMC) enables compliant behaviors by embedding virtual forces and mapping them into joint torques, but its reliance on fixed parameters and limited coordination among virtual components constrains adaptability and may undermine stability as task objectives evolve. To address these limitations, we propose Adaptive VMC with Large Language Model (LLM)- and Lyapunov-Based Reinforcement Learning (RL), which preserves the physical interpretability of VMC while supporting stability-guaranteed online adaptation. The LLM provides structured priors and high-level reasoning that enhance coordination among virtual components, improve sample efficiency, and facilitate flexible adjustment to varying task requirements. Complementarily, Lyapunov-based RL enforces theoretical stability constraints, ensuring safe and reliable adaptation under uncertainty. Extensive simulations on a 7-DoF Panda arm demonstrate that our approach effectively balances competing objectives in dynamic tasks, achieving superior performance while highlighting the synergistic benefits of LLM guidance and Lyapunov-constrained adaptation.

cs.RO

DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models

Large language models are highly sensitive to prompt wording. However, popular automatic prompt search methods, including InstructZero, often degrade under distribution shift and adversarial evaluation because they optimize expected performance under a single evaluation distribution. Consequently, prompts that work in one setting frequently fail to transfer. To address this, DRO-InstructZero formulates zero-shot prompt optimization as robust Bayesian optimization. Specifically, an f-divergence ball defines an ambiguity set around the evaluation distribution, and a robust acquisition rule maximizes worst-case expected utility while retaining the query efficiency of Bayesian search. Therefore, the search explicitly targets reliability under distribution shift rather than average behavior alone. Experiments follow the instruction-induction protocol with matched query budgets across formality rewriting, code debugging, and translation. For example, on BIG-Bench informative-to-formal rewriting, accuracy improves from 61.3 +/- 0.7% to approximately 85-90%, yielding an absolute gain of about 25-30 points. Moreover, auto-debugging shows about +25-point gains under domain shift. Meanwhile, stable tasks such as cause-and-effect remain above 96%, indicating no loss on in-distribution cases. Furthermore, improvements are consistent across divergence choices and decoding temperatures. Overall, DRO-InstructZero connects distributionally robust optimization with prompt learning, offering a plug-and-play and general approach for reliable, transferable prompt alignment under real-world uncertainty.

cs.LG

A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease

Early detection of Alzheimer's Disease (AD) is greatly beneficial to AD patients, leading to early treatments that lessen symptoms and alleviating financial burden of health care. As one of the leading signs of AD, language capability changes can be used for early diagnosis of AD. In this paper, I develop a robust classification method using hybrid word embedding and fine-tuned hyperparameters to achieve state-of-the-art accuracy in the early detection of AD. Specifically, we create a hybrid word embedding based on word vectors from Doc2Vec and ELMo to obtain perplexity scores of the sentences. The scores identify whether a sentence is fluent or not and capture semantic context of the sentences. I enrich the word embedding by adding linguistic features to analyze syntax and semantics. Further, we input an embedded feature vector into logistic regression and fine tune hyperparameters throughout the pipeline. By tuning hyperparameters of the machine learning pipeline (e.g., model regularization parameter, learning rate and vector size of Doc2Vec, and vector size of ELMo), I achieve 91% classification accuracy and an Area Under the Curve (AUC) of 97% in distinguishing early AD from healthy subjects. Based on my knowledge, my model with 91% accuracy and 97% AUC outperforms the best existing NLP model for AD diagnosis with an accuracy of 88% [32]. I study the model stability through repeated experiments and find that the model is stable even though the training data is split randomly (standard deviation of accuracy = 0.0403; standard deviation of AUC = 0.0174). This affirms our proposed method is accurate and stable. This model can be used as a large-scale screening method for AD, as well as a complementary examination for doctors to detect AD.

cs.CL

The Impact of Medicaid Coverage on Mental Health, Why Insurance Makes People Happier in OHIE: by Spending Less or by Spending More?

The Oregon Health Insurance Experiment (OHIE) offers a unique opportunity to examine the causal relationship between Medicaid coverage and happiness among low-income adults, using an experimental design. This study leverages data from comprehensive surveys conducted at 0 and 12 months post-treatment. Previous studies based on OHIE have shown that individuals receiving Medicaid exhibited a significant improvement in mental health compared to those who did not receive coverage. The primary objective is to explore how Medicaid coverage impacts happiness, specifically analyzing in which direction variations in healthcare spending significantly improve mental health: higher spending or lower spending after Medicaid. Utilizing instrumental variable (IV) regression, I conducted six separate regressions across subgroups categorized by expenditure levels and happiness ratings, and the results reveal distinct patterns. Enrolling in OHP has significantly decreased the probability of experiencing unhappiness, regardless of whether individuals had high or low medical spending. Additionally, it decreased the probability of being pretty happy and having high medical expenses, while increasing the probability among those with lower expenses. Concerning the probability of being very happy, the OHP only had a positive effect on being very happy and spending less, and its effect on those with high expenses was insignificant. These findings align with the benefit of Medicaid: alleviating financial burden, contributing to the well-being of distinct subgroups.

econ.GN

Source-Free Object Detection with Detection Transformer

Source-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existing SFOD approaches are either confined to conventional object detection (OD) models like Faster R-CNN or designed as general solutions without tailored adaptations for novel OD architectures, especially Detection Transformer (DETR). In this paper, we introduce Feature Reweighting ANd Contrastive Learning NetworK (FRANCK), a novel SFOD framework specifically designed to perform query-centric feature enhancement for DETRs. FRANCK comprises four key components: (1) an Objectness Score-based Sample Reweighting (OSSR) module that computes attention-based objectness scores on multi-scale encoder feature maps, reweighting the detection loss to emphasize less-recognized regions; (2) a Contrastive Learning with Matching-based Memory Bank (CMMB) module that integrates multi-level features into memory banks, enhancing class-wise contrastive learning; (3) an Uncertainty-weighted Query-fused Feature Distillation (UQFD) module that improves feature distillation through prediction quality reweighting and query feature fusion; and (4) an improved self-training pipeline with a Dynamic Teacher Updating Interval (DTUI) that optimizes pseudo-label quality. By leveraging these components, FRANCK effectively adapts a source-pre-trained DETR model to a target domain with enhanced robustness and generalization. Extensive experiments on several widely used benchmarks demonstrate that our method achieves state-of-the-art performance, highlighting its effectiveness and compatibility with DETR-based SFOD models.

cs.CV