SearcharxivSearch

arXiv subjects

Min Huang

Publications and source records attributed to Min Huang.

At least 19 recordsLinked to original sources

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.

cs.CV

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.

cs.CV

Quantum cluster algebra realization for stated ${\rm SL}_n$-skein algebras and rotation-invariant bases for polygons

We construct a quantum cluster structure on the skew-field of fractions ${\rm Frac}({\mathscr S}_\omega(\mathfrak{S}))$ of the stated ${\rm SL}_n$-skein algebra ${\mathscr S}_\omega(\mathfrak{S})$, where $\mathfrak{S}$ is a triangulable pb surface without interior punctures. This work complements the construction for the projected stated skein algebra $\widetilde{\mathscr S}_\omega(\mathfrak{S})$ given by the last two authors. Let ${\mathscr S}_\omega^{\rm fr}(\mathfrak{S})$ denote the localization of ${\mathscr S}_\omega(\mathfrak{S})$ at the multiplicative set generated by all frozen variables. Let ${\mathscr A}_\omega^{\rm fr}(\mathfrak{S})$ and ${\mathscr U}_\omega^{\rm fr}(\mathfrak{S})$ (respectively $\overline{\mathscr A}_\omega(\mathfrak{S})$ and $\overline{\mathscr U}_\omega(\mathfrak{S})$) denote the quantum cluster algebra and quantum upper cluster algebra associated to ${\rm Frac}({\mathscr S}_\omega(\mathfrak{S}))$ (respectively ${\rm Frac}(\widetilde{\mathscr S}_\omega(\mathfrak{S}))$). We prove that \[ \widetilde{\mathscr S}_\omega(\mathfrak{S}) = \overline{\mathscr A}_\omega(\mathfrak{S}) = \overline{\mathscr U}_\omega(\mathfrak{S}) \quad \text{and} \quad {\mathscr S}_\omega^{\rm fr}(\mathfrak{S}) = {\mathscr A}_\omega^{\rm fr}(\mathfrak{S}) = {\mathscr U}_\omega^{\rm fr}(\mathfrak{S}) \] whenever $\mathfrak{S}$ is a polygon. As a consequence, when $\mathfrak{S}$ is a polygon, we show that the theta basis of $\overline{\mathscr U}_\omega(\mathfrak{S})$ (respectively ${\mathscr U}_\omega^{\rm fr}(\mathfrak{S})$) yields a rotation-invariant basis of $\overline{\mathscr S}_\omega(\mathfrak{S})$ (respectively ${\mathscr S}_\omega^{\rm fr}(\mathfrak{S})$) with several desirable properties, including positivity and a natural parametrization.

math.QA

ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving

Vision-Language Models (VLMs) have become central to autonomous driving systems, yet their deployment is severely bottlenecked by the massive computational overhead of multi-view camera and multi-frame video input. Existing token pruning methods, primarily designed for single-image inputs, treat each frame or view in isolation and thus fail to exploit the inherent spatio-temporal redundancies in driving scenarios. To bridge this gap, we propose ST-Prune, a training-free, plug-and-play framework comprising two complementary modules: Motion-aware Temporal Pruning (MTP) and Ring-view Spatial Pruning (RSP). MTP addresses temporal redundancy by encoding motion volatility and temporal recency as soft constraints within the diversity selection objective, prioritizing dynamic trajectories and current-frame content over static historical background. RSP further resolves spatial redundancy by exploiting the ring-view camera geometry to penalize bilateral cross-view similarity, eliminating duplicate projections and residual background that temporal pruning alone cannot suppress. These two modules together constitute a complete spatio-temporal pruning process, preserving key scene information under strict compression. Validated across four benchmarks spanning perception, prediction, and planning, ST-Prune establishes new state-of-the-art for training-free token pruning. Notably, even at 90\% token reduction, ST-Prune achieves near-lossless performance with certain metrics surpassing the full-model baseline, while maintaining inference speeds comparable to existing pruning approaches.

cs.CV

Orderings of Generalized k-Markov Numbers

A $k$-Markov number is a positive integer that appears in a positive integral solution to the Diophantine equation $x^2 + y^2 + z^2 + k(xy + xz + yz) = (3+3k)xyz$. This equation was introduced by Gyoda and Matsushita. When $k =0$, this definition recovers that of ordinary Markov numbers. The set of $k$-Markov numbers can be indexed by pairs of coprime positive integers. There is a consistent way to label non-coprime pairs with positive integers as well, yielding a larger set of ``generalized $k$-Markov numbers.'' In this paper, we classify lines along which the generalized $k$-Markov numbers grow monotonically, extending work in the ordinary case by Lee-Li-Rabideau-Schiffler and by the second author. We find that, as $k$ grows, the $k$-Markov numbers are more likely to be monotonic along a random line. This gives evidence that a $k$-version of Frobenius' uniqueness conjecture, which has been proposed by Gyoda and Maruyama, could be true.

math.NT

CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis

Thematic analysis is difficult to scale: manual workflows are labor-intensive, while fully automated pipelines often lack controllability and transparent evaluation. We present \textbf{CentaurTA Studio}, a web-based system for self-improving human--agent collaboration in open coding and theme construction. The system integrates (1) a two-stage human feedback pipeline separating simulator drafting and expert validation, (2) persistent prompt optimization that distills validated feedback into reusable alignment principles, and (3) rubric-based evaluation with early stopping for process control. Across three domains, CentaurTA achieves the strongest performance in both Open Coding and Theme Construction, reaching up to 92.12\% accuracy and consistently outperforming baseline systems. Agreement between the rubric-based LLM judge and human annotators reaches substantial reliability (average $\kappa = 0.68$). Ablation studies show that removing the feedback loop reduces performance from 90\% to 81\%, while eliminating the Critic or early stopping degrades accuracy or increases interaction cost. The full system reaches peak performance within 10 iterative rounds (about 25 minutes), demonstrating improved efficiency over expert-only refinement.

cs.HC

DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis

Aspect-Based Sentiment Intensity Analysis (ABSIA) has garnered increasing attention, though research largely focuses on domain-specific, sentence-level settings. In contrast, document-level ABSIA--particularly in addressing complex tasks like extracting Aspect-Category-Opinion-Sentiment-Intensity (ACOSI) tuples--remains underexplored. In this work, we introduce DanceHA, a multi-agent framework designed for open-ended, document-level ABSIA with informal writing styles. DanceHA has two main components: Dance, which employs a divide-and-conquer strategy to decompose the long-context ABSIA task into smaller, manageable sub-tasks for collaboration among specialized agents; and HA, Human-AI collaboration for annotation. We release Inf-ABSIA, a multi-domain document-level ABSIA dataset featuring fine-grained and high-accuracy labels from DanceHA. Extensive experiments demonstrate the effectiveness of our agentic framework and show that the multi-agent knowledge in DanceHA can be effectively transferred into student models. Our results highlight the importance of the overlooked informal styles in ABSIA, as they often intensify opinions tied to specific aspects.

cs.CL

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness

We present ACE-Bench (Azure SDK Coding Evaluation Benchmark), an execution-free benchmark that provides fast, reproducible pass or fail signals for whether large language model (LLM)-based coding agents use Azure SDKs correctly-without provisioning cloud resources or maintaining fragile end-to-end test environments. ACE-Bench turns official Azure SDK documentation examples into self-contained coding tasks and validates solutions with task-specific atomic criteria: deterministic regex checks that enforce required API usage patterns and reference-based LLM-judge checks that capture semantic workflow constraints. This design makes SDK-centric evaluation practical in day-to-day development and CI: it reduces evaluation cost, improves repeatability, and scales to new SDKs and languages as documentation evolves. Using a lightweight coding agent, we benchmark multiple state-of-the-art LLMs and quantify the benefit of retrieval in an MCP-enabled augmented setting, showing consistent gains from documentation access while highlighting substantial cross-model differences.

cs.DC

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy to deploy but often fragile under distribution shift and strict parsers, while continual parameter-efficient fine-tuning improves reliability at the cost of training, maintenance, and potential forgetting. We identify a critical Lazy Agent failure mode where tool necessity is nearly perfectly decodable from mid-layer activations, yet the model remains conservative in entering tool mode, revealing a representation-behavior gap. We propose Activation Steering Adapter (ASA), a training-free, inference-time controller that performs a single-shot mid-layer intervention and targets tool domains via a router-conditioned mixture of steering vectors with a probe-guided signed gate to amplify true intent while suppressing spurious triggers. On MTU-Bench with Qwen2.5-1.5B, ASA improves strict tool-use F1 from 0.18 to 0.50 while reducing the false positive rate from 0.15 to 0.05, using only about 20KB of portable assets and no weight updates.

cs.SE

A mutation invariant for skew-symmetrizable matrices

Matrix mutation of skew-symmetrizable matrices is foundational in cluster algebra theory. Effective mutation invariants are essential for determining whether two matrices lie in the same mutation class. Casals~\cite{Casals} introduced a binary mutation invariant for skew-symmetric matrices. In this paper, we extend Casals' construction to the skew-symmetrizable setting. When the skew-symmetrizer $d_1,\dots, d_n$ is pairwise coprime, we obtain two distinct extensions of this invariant.

math.CO

Regulatory Hub Discovery in MDD Methylome: Hypotheses for Molecular Subtypes via Computational Analysis

Major Depressive Disorder (MDD) is a clinically heterogeneous syndrome with diverse etiological pathways. Traditional Epigenome-Wide Association Studies (EWAS) have successfully identified risk loci based on differential methylation magnitude. As a complementary perspective, effect-size-based ranking alone may not fully capture regulatory nodes that exhibit modest methylation changes but occupy critical upstream positions in biological networks. Here, we report findings and hypotheses from a two-tier computational analysis of DNA methylation data (GSE198904; \(n=206\) ), combining conventional statistical approaches with machine learning-assisted regulatory inference.

cs.CE

SpikGPT: A High-Accuracy and Interpretable Spiking Attention Framework for Single-Cell Annotation

Accurate and scalable cell type annotation remains a challenge in single-cell transcriptomics, especially when datasets exhibit strong batch effects or contain previously unseen cell populations. Here we introduce SpikGPT, a hybrid deep learning framework that integrates scGPT-derived cell embeddings with a spiking Transformer architecture to achieve efficient and robust annotation. scGPT provides biologically informed dense representations of each cell, which are further processed by a multi-head Spiking Self-Attention mechanism for energy-efficient feature extraction. Across multiple benchmark datasets, SpikGPT consistently matches or exceeds the performance of leading annotation tools. Notably, SpikGPT uniquely identifies unseen cell types by assigning low-confidence predictions to an "Unknown" category, allowing accurate rejection of cell states absent from the training reference. Together, these results demonstrate that SpikGPT is a versatile and reliable annotation tool capable of generalizing across datasets, resolving complex cellular heterogeneity, and facilitating discovery of novel or disease-associated cell populations.

q-bio.QM

The First Scientific Flight and Observations of the 50-mm Balloon-Borne White-Light Coronagraph

A 50-mm balloon-borne white-light coronagraph (BBWLC) to observe whitelight solar corona over the altitude range from 1.08 to 1.50 solar radii has recently been indigenously developed by Yunnan Observatories in collaboration with Shangdong University (in Weihai) and Changchun Institute of Optics, Fine Mechanics and Physics, which will significantly improve the ability of China to detect and measure inner corona. On 2022 October 4, its first scientific flight took place at the Dachaidan area in Qinghai province of China. We describe briefly the BBWLC mission including its optical design, mechanical structure, pointing system, the first flight and results associated with the data processing approach. Preliminary analysis of the data shows that BBWLC imaged the Kcorona with three streamer structures on the west limb of the Sun. To further confirm the coronal signals obtained by BBWLC, comparisonswere made with observations of the Kcoronagraph of the High Altitude Observatory and the Atmospheric ImagingAssembly on board the Solar Dynamics Observatory. We conclude that BBWLC eventually observed the white-light corona in its first scientific flight.

astro-ph.SR

Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics

Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech production, which also encompasses speaker traits, including emotional states. To bridge this gap, we conducted a series of experiments to investigate the potential of the phonation excitation information and articulatory kinematics for SER. Due to the scarcity of training data for this purpose, we introduce a portrayed emotional dataset, STEM-E2VA, which includes audio and physiological data such as electroglottography (EGG) and electromagnetic articulography (EMA). EGG and EMA provide information of phonation excitation and articulatory kinematics, respectively. Additionally, we performed emotion recognition using estimated physiological data derived through inversion methods from speech, instead of collected EGG and EMA, to explore the feasibility of applying such physiological information in real-world SER. Experimental results confirm the effectiveness of incorporating physiological information about speech production for SER and demonstrate its potential for practical use in real-world scenarios.

cs.SD

Towards Quantum Simulations of Sphaleron Dynamics at Colliders

Sphaleron dynamics in the Standard Model at high-energy particle collisions remains experimentally unobserved, with theoretical predictions hindered by its nonperturbative real-time nature. In this work, we investigate a quantum simulation approach to this challenge. Taking the $1+1$D $O(3)$ model as a protocol towards studying dynamics of sphaleron in the electroweak theory, we identify the sphaleron configuration and establish lattice parameters that reproduce continuum sphaleron energies with controlled precision. We then develop quantum algorithms to simulate sphaleron evolutions where quantum effects can be included. This work lays the ground to establish quantum simulations for studying the interaction between classical topological objects and particles in the quantum field theory that are usually inaccessible to classical methods and computations.

hep-ph

Quantum cluster realization for projected stated ${\rm SL}_n$-skein algebras

We introduce a quantum cluster algebra structure $\mathscr A_\omega(\mathfrak{S})$ inside the skew-field fractions ${\rm Frac}\bigl(\widetilde{\mathscr{S}}_\omega(\mathfrak{S})\bigr)$ of the projected stated ${\rm SL}_n$-skein algebra $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})$ (the quotient of the reduced stated ${\rm SL}_n$-skein algebra by the kernel of the quantum trace map) for any triangulable pb surface $\mathfrak{S}$ without interior punctures. To study the relationships among the projected ${\rm SL}_n$-skein algebra $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})$, the quantum cluster algebra $\mathscr A_\omega(\mathfrak{S})$, and its quantum upper cluster algebra $\mathscr U_\omega(\mathfrak{S})$, we construct a splitting homomorphism for $\mathscr U_\omega(\mathfrak{S})$ and show that it is compatible with the splitting homomorphism for $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})$. When every connected component of $\mathfrak{S}$ contains at least two punctures, this compatibility allows us to prove that $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})$ embeds into $\mathscr A_\omega(\mathfrak{S})$ by showing that the stated arcs joining two distinct boundary components of $\mathfrak{S}$ (which generate $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})$) are, up to multiplication by a Laurent monomial in the frozen variables, exchangeable cluster variables. We further conjecture that these exchangeable cluster variables generate the quantum upper cluster algebra $\mathscr U_\omega(\mathfrak{S})$, which, if true, would imply the equality $\widetilde{\mathscr{S}}_\omega(\mathfrak{S})=\mathscr A_\omega(\mathfrak{S})=\mathscr U_\omega(\mathfrak{S})$.

math.QA

Noncommutative marked surfaces II: tagged triangulations, clusters, and their symmetries

The aim of the paper is to define noncommutative cluster structure on several algebras ${\mathcal A}$ related to marked surfaces possibly with orbifold points of various orders, which includes noncommutative clusters, i.e., embeddings of a given group $G$ into the multiplicative monoid ${\mathcal A}^\times$ and an action of a certain braid-like group $Br_{\mathcal A}$ by automorphisms of each cluster group in a compatible way. For punctured surfaces we construct new symmetries, noncommutative tagged clusters and establish a noncommutative Laurent Phenomenon.

math.RT