SearcharxivSearch

arXiv subjects

Qian Huang

Publications and source records attributed to Qian Huang.

At least 19 recordsLinked to original sources

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tokens. By applying classical Doeblin--Dobrushin contraction theory, we show that a rollout operator with a small Dobrushin coefficient is quantitatively close to a rank-one stochastic matrix whose common row is determined by its normalized column sums. This result gives a structural interpretation to the corresponding rollout propagation profile. In Transformers trained for metabolomic age prediction, the measured rollout contraction strengthens with depth. Trained and randomly initialized models also exhibit different propagation profiles, although the present experiments do not establish the predictive relevance of individual rollout-ranked variables. Exploratory comparisons with PCA and GradientExplainer approximations to SHAP reveal localized agreement among highly ranked variables but weak agreement across complete rankings. Attention rollout is therefore used here as a diagnostic of attention-mediated propagation, not as a causal explanation or faithful attribution of the complete Transformer.

cs.LG

Direct Evidence for Outflow Driven by Wolf-Rayet Stars in the Nearby Galaxy PGC44685

Wolf--Rayet (WR) stars are evolved massive stars which can drive strong stellar winds, injecting energy and momentum into the interstellar medium (ISM). However, the geometry and kinematics of WR-dominated outflows, specially in low-metallicity environments, is still poorly constrained by observations. We present a spatially resolved spectroscopic study of a WR region in a nearby dwarf galaxy, PGC\,44685, using high-resolution MEGARA IFU data from the Gran Telescopio Canarias (GTC). After decomposing the [\textsc{O iii}]~$\lambda5007$ emission line with narrow and broad components, we verify a WR-driven outflow with a velocity reaching up to $20\,\mathrm{km\,s^{-1}}$ relative to the systemic velocity. By use of the velocity and flux of the [\textsc{O iii}] broad component, we estimate an outflow mass of $(8.25 \pm 3.03)\times10^{3}\,M_\odot$ and a mass-loss rate of $(9.47 \pm 3.48)\times10^{-4}\,M_\odot\,\mathrm{yr}^{-1}$. The corresponding kinetic power and momentum injection rate are $(4.77 \pm 1.77)\times10^{41}\,\mathrm{erg\,s^{-1}}$ and $(8.20 \pm 3.02)\times10^{28}\,\mathrm{g\,cm\,s^{-2}}$, respectively. The inferred low energy-loading efficiency ($\sim0.35\%$), together with the low metallicity of the WR region ($\sim0.1\,Z_\odot$), suggests that the system is observed in an early feedback phase in which stellar winds have not yet efficiently coupled their energy into the ISM. These results support the ability of WR feedback to shape the ISM on sub-kiloparsec scales, while these winds fail to launch galactic-scale outflows.

astro-ph.GA

FusionRelight: Relighting Portraits in Real Time via Hybrid Domain Knowledge Fusion

Portrait relighting is a low-level vision problem in which physically plausible illumination transfer, identity preservation, and compact real-time inference must be considered together. Iterative diffusion-style methods can synthesize fine detail, but stochastic inference and cost complicate deterministic live video creation; physically grounded relighting preserves identity, but controlled synthetic or light-stage supervision transfers poorly to unconstrained cameras. We present Hybrid Domain Knowledge Fusion (HDKF), a relighting-specific training framework that learns complementary physics, reflectance, and realism priors from synthetic, One-Light-at-A-Time (OLAT), and in-the-wild data, then distills their source-routed supervision into a compact student with clean teacher labels and degraded student inputs. The framework is trained with pixel-aligned RGB, albedo, and normal supervision, providing a simulation substrate for physically grounded low-level relighting. On a held-out OLAT benchmark, HDKF obtains the best MSE, PSNR, and SSIM among evaluated methods while remaining competitive in LPIPS. The distilled model runs in real time at 512x512, reaching 11.89 ms on an RTX 2060 and 1.82 ms on an RTX 4090.

cs.CV

A data-driven approach for 2D vorticity PDF equations by a new conditional average estimation

We consider the statistics for the vorticity field in two-dimensional homogeneous isotropic turbulence (HIT). First, we exploit the invariance properties to derive dimensionally reduced governing equations for the one-point and two-point probability density functions (PDFs). These take the form of linear kinetic transport equations, but with an unclosed operator in terms of a conditional average. To solve the PDF equation numerically we suggest a hybrid data-driven method that relies on carefully selected samples of DNS data and a sampling estimator for the conditional average. The method is applied to DNS data for both decaying and forced HIT, demonstrating good agreement with the direct evaluation of the PDFs using the DNS data.

physics.flu-dyn

SuperSkillsStack: Agency, Domain Knowledge, Imagination, and Taste in Human-AI Design Education

This study examines how students integrate generative artificial intelligence (AI) into design projects through the lens of the SuperSkillsStack framework, which identifies four key human competencies for effective human-AI collaboration: Agency, Domain Knowledge, Imagination, and Taste. As generative AI increasingly transforms creative practice, design education must consider how human capabilities are cultivated alongside technological tools. Using qualitative thematic analysis, this study analyzes reflective writings from 80 student design teams participating in a human-centered design course. The findings show that students primarily used AI during the early stages of the design process, including brainstorming, information synthesis, and problem framing. However, students consistently relied on human judgment to interpret contextual information, validate AI-generated outputs, and refine design solutions. Domain knowledge derived from field observations enabled students to detect inaccuracies in AI suggestions, while taste played a critical role in evaluating and selecting meaningful ideas. The results suggest that generative AI functions primarily as a cognitive accelerator rather than a replacement for human creativity. The study highlights the importance of cultivating higher-order human capabilities to support effective human-AI collaboration in design education.

cs.CY

The Trilingual Triad Framework: Integrating Design, AI, and Domain Knowledge in No-code AI Smart City Course

This paper introduces the "Trilingual Triad" framework, a model that explains how students learn to design with generative artificial intelligence (AI) through the integration of Design, AI, and Domain Knowledge. As generative AI rapidly enters higher education, students often engage with these systems as passive users of generated outputs rather than active creators of AI-enabled knowledge tools. This study investigates how students can transition from using AI as a tool to designing AI as a collaborative teammate. The research examines a graduate course, Creating the Frontier of No-code Smart Cities at the Singapore University of Technology and Design (SUTD), in which students developed domain-specific custom GPT systems without coding. Using a qualitative multi-case study approach, three projects - the Interview Companion GPT, the Urban Observer GPT, and Buddy Buddy - were analyzed across three dimensions: design, AI architecture, and domain expertise. The findings show that effective human-AI collaboration emerges when these three "languages" are orchestrated together: domain knowledge structures the AI's logic, design mediates human-AI interaction, and AI extends learners' cognitive capacity. The Trilingual Triad framework highlights how building AI systems can serve as a constructionist learning process that strengthens AI literacy, metacognition, and learner agency.

cs.AI

CH3CCH as a thermometer in warm molecular gas

Kinetic temperature is a fundamental parameter in molecular clouds. Symmetric top molecules, such as NH$_3$ and CH$_3$CCH, are often used as thermometers. However, at high temperatures, NH$_3$(2,2) can be collisionally excited to NH$_3$(2,1) and rapidly decay to NH$_3$(1,1), which can lead to an underestimation of the kinetic temperature when using rotation temperatures derived from NH$_3$(1,1) and NH$_3$(2,2). In contrast, CH$_3$CCH is a symmetric top molecule with lower critical densities of its rotational levels than those of NH$_3$, which can be thermalized close to the kinetic temperature at relatively low densities of about 10$^{4}$ cm$^{-3}$. To compare the rotation temperatures derived from NH$_3$(1,1)$\&$(2,2) and CH$_3$CCH rotational levels in warm molecular gas, we used observational data toward 55 massive star-forming regions obtained with Yebes 40m and TMRT 65m. Our results show that rotation temperatures derived from NH$_3$(1,1)$\&$(2,2) are systematically lower than those from CH$_3$CCH 5-4. This suggests that CH$_3$CCH rotational lines with the same $J$+1$\rightarrow$$J$ quantum number may be a more reliable thermometer than NH$_3$(1,1)$\&$(2,2) in warm molecular gas located in the surroundings of massive young stellar objects or, more generally, in massive star-forming regions.

astro-ph.GA

Convergence of a two-parameter hyperbolic relaxation system toward the incompressible Navier-Stokes equations

We investigate a two-parameter hyperbolic relaxation approximation to the incompressible Navier-Stokes equations, incorporating a first-order relaxation and the artificial compressibility method. With vanishingly small perturbations of initial velocity, we rigorously prove the simultaneous convergence of fluid velocity and pressure toward the Navier-Stokes limit in the three-dimensional case by constructing an intermediate affine system to obtain the necessary error estimates for the pressure. Furthermore, we extend the velocity convergence analysis to the case of $\mathcal O(1)$ initial velocity perturbations, and establish the global-in-time recovery of the velocity field using a modulated energy structure and delicate bootstrap arguments in both two- and three-dimensional settings.

math.AP

Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning

Medical vision-language models (VLMs) achieve strong performance in diagnostic reporting and image-text alignment, yet their underlying reasoning mechanisms remain fundamentally correlational, exhibiting reliance on superficial statistical associations that fail to capture the causal pathophysiological mechanisms central to clinical decision-making. This limitation makes them fragile, prone to hallucinations, and sensitive to dataset biases. Retrieval-augmented generation (RAG) offers a partial remedy by grounding predictions in external knowledge. However, conventional RAG depends on semantic similarity, introducing new spurious correlations. We propose Multimodal Causal Retrieval-Augmented Generation, a framework that integrates causal inference principles with multimodal retrieval. It retrieves clinically relevant exemplars and causal graphs from external sources, conditioning model reasoning on counterfactual and interventional evidence rather than correlations alone. Applied to radiology report generation, diagnosis prediction, and visual question answering, it improves factual accuracy, robustness to distribution shifts, and interpretability. Our results highlight causal retrieval as a scalable path toward medical VLMs that think beyond pattern matching, enabling trustworthy multimodal reasoning in high-stakes clinical settings.

cs.LG

FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.

cs.CV

A moment model of shallow granular flows with variable friction laws

In this work, we develop a modelling framework for granular flows based on the shallow water moment equations on inclined planes. Under the assumption of a polynomial expansion of the velocity field, the model extends the classical shallow water equations to vertically variable velocity profiles. The friction effects, which are captured through the strain-rate tensor, are incorporated into the model in two terms, the bulk and bottom friction. We propose a modelling procedure to incorporate general friction laws into our framework and exemplify this combining the Manning, Coulomb, Savage-Hutter, and $μ(I)$-rheology friction models in our modeling framework. Moreover, we develop a path-conservative finite volume numerical scheme based on the polynomial viscosity matrix method to properly handle the stiffness of the source terms. Numerical simulations are presented for different models of friction, including the case of wet-dry fronts.

math.NA

To Use or to Refuse? Re-Centering Student Agency with Generative AI in Engineering Design Education

This pilot study traces students' reflections on the use of AI in a 13-week foundational design course enrolling over 500 first-year engineering and architecture students at the Singapore University of Technology and Design. The course was an AI-enhanced design course, with several interventions to equip students with AI based design skills. Students were required to reflect on whether the technology was used as a tool (instrumental assistant), a teammate (collaborative partner), or neither (deliberate non-use). By foregrounding this three-way lens, students learned to use AI for innovation rather than just automation and to reflect on agency, ethics, and context rather than on prompt crafting alone. Evidence stems from coursework artefacts: thirteen structured reflection spreadsheets and eight illustrated briefs submitted, combined with notes of teachers and researchers. Qualitative coding of these materials reveals shared practices brought about through the inclusion of Gen-AI, including accelerated prototyping, rapid skill acquisition, iterative prompt refinement, purposeful "switch-offs" during user research, and emergent routines for recognizing hallucinations. Unexpectedly, students not only harnessed Gen-AI for speed but (enabled by the tool-teammate-neither triage) also learned to reject its outputs, invent their own hallucination fire-drills, and divert the reclaimed hours into deeper user research, thereby transforming efficiency into innovation. The implications of the approach we explore shows that: we can transform AI uptake into an assessable design habit; that rewarding selective non-use cultivates hallucination-aware workflows; and, practically, that a coordinated bundle of tool access, reflection, role tagging, and public recognition through competition awards allows AI based innovation in education to scale without compromising accountability.

cs.CY

Designing Knowledge Tools: How Students Transition from Using to Creating Generative AI in STEAM classroom

This study explores how graduate students in an urban planning program transitioned from passive users of generative AI to active creators of custom GPT-based knowledge tools. Drawing on Self-Determination Theory (SDT), which emphasizes the psychological needs of autonomy, competence, and relatedness as foundations for intrinsic motivation, the research investigates how the act of designing AI tools influences students' learning experiences, identity formation, and engagement with knowledge. The study is situated within a two-term curriculum, where students first used instructor-created GPTs to support qualitative research tasks and later redesigned these tools to create their own custom applications, including the Interview Companion GPT. Using qualitative thematic analysis of student slide presentations and focus group interviews, the findings highlight a marked transformation in students' roles and mindsets. Students reported feeling more autonomous as they chose the functionality, design, and purpose of their tools, more competent through the acquisition of AI-related skills such as prompt engineering and iterative testing, and more connected to peers through team collaboration and a shared sense of purpose. The study contributes to a growing body of evidence that student agency can be powerfully activated when learners are invited to co-design the very technologies they use. The shift from AI tool users to AI tool designers reconfigures students' relationships with technology and knowledge, transforming them from consumers into co-creators in an evolving educational landscape.

cs.CY

Provably realizability-preserving finite volume method for quadrature-based moment models of kinetic equations

Quadrature-based moment methods (QBMM) provide tractable closures for multiscale kinetic equations, with diverse applications across aerosols, sprays, and particulate flows, etc. However, for the derived hyperbolic moment-closure systems, seeking numerical schemes preserving moment realizability is essential yet challenging due to strong nonlinear coupling and the lack of explicit conservative-to-flux maps. This paper proposes and analyzes a provably realizability-preserving finite-volume method for five-moment systems closed by the two-node Gaussian-EQMOM and three-point HyQMOM. Rather than relying on kinetic fluxes, we recast the realizability condition into a nonnegative quadratic form in the moment vector, reducing the original nonlinear constraints to bilinear inequalities amenable to analysis. On this basis, we construct a tailored Harten--Lax--van Leer (HLL) flux with rigorously derived wave speeds and intermediate states that embed realizability directly into the flux evaluation. We prove sufficient realizability-preserving conditions under explicit Courant--Friedrichs--Lewy (CFL) constraints in the collisionless case, and for BGK relaxation, we obtain coupled time-step conditions involving a realizability radius; a semi-implicit BGK variant inherits the collisionless CFL. From a multiscale perspective, the analysis yields stability conditions uniform in the relaxation time and supports stiff-to-kinetic transitions. A practical limiter enforces strict realizability of reconstructed interface states without degrading accuracy. Numerical experiments demonstrate the accuracy, robustness in low-density regions, and realizability for both closures. This framework unifies realizability preservation for solving hyperbolic moment systems with complex closures and extends naturally to higher-order space--time discretizations.

math.NA

Numerical approximations to statistical conservation laws for scalar hyperbolic equations

Motivated by the statistical description of turbulence, we study statistical conservation laws in the form of kinetic-type PDEs for joint probability density functions (PDFs) and cumulative distribution functions (CDFs) associated with solutions of scalar balance laws. Starting from viscous balance laws, the resulting PDF/CDF equations involve unclosed conditional averages arising in the viscous terms. We show that these terms exhibit a dissipative anomaly: they remain non-negligible in the vanishing viscosity limit and are essential to preserve the nonnegativity of evolving PDFs. To approximate these PDF/CDF equations in a unified framework, we propose a novel sampling-based estimator for the unclosed terms, constructed from numerical or exact realizations of the underlying balance-law solutions. In certain cases, a priori error bounds can be derived, demonstrating that the deviation between the true and approximate CDFs is controlled by the estimation error of the unclosed terms. Numerical experiments with analytically solvable test problems confirm that the sampling-based approximation converges satisfactorily with the number of samples.

math.NA

Language-Enhanced Mobile Manipulation for Efficient Object Search in Indoor Environments

Enabling robots to efficiently search for and identify objects in complex, unstructured environments is critical for diverse applications ranging from household assistance to industrial automation. However, traditional scene representations typically capture only static semantics and lack interpretable contextual reasoning, limiting their ability to guide object search in completely unfamiliar settings. To address this challenge, we propose a language-enhanced hierarchical navigation framework that tightly integrates semantic perception and spatial reasoning. Our method, Goal-Oriented Dynamically Heuristic-Guided Hierarchical Search (GODHS), leverages large language models (LLMs) to infer scene semantics and guide the search process through a multi-level decision hierarchy. Reliability in reasoning is achieved through the use of structured prompts and logical constraints applied at each stage of the hierarchy. For the specific challenges of mobile manipulation, we introduce a heuristic-based motion planner that combines polar angle sorting with distance prioritization to efficiently generate exploration paths. Comprehensive evaluations in Isaac Sim demonstrate the feasibility of our framework, showing that GODHS can locate target objects with higher search efficiency compared to conventional, non-semantic search strategies. Website and Video are available at: https://drapandiger.github.io/GODHS

cs.RO

Statistical conservation laws for scalar model problems: Hierarchical evolution equations

The probability density functions (PDFs) for the solution of the incompressible Navier-Stokes equation can be represented by a hierarchy of linear equations. This article develops new hierarchical evolution equations for PDFs of a scalar conservation law with random initial data as a model problem. Two frameworks are developed, including multi-point PDFs and single-point higher-order derivative PDFs. These hierarchies capture statistical correlations and guide closure strategies.

math.AP

Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems

Advances in artificial intelligence (AI) are fueling a new paradigm of discoveries in natural sciences. Today, AI has started to advance natural sciences by improving, accelerating, and enabling our understanding of natural phenomena at a wide range of spatial and temporal scales, giving rise to a new area of research known as AI for science (AI4Science). Being an emerging research paradigm, AI4Science is unique in that it is an enormous and highly interdisciplinary area. Thus, a unified and technical treatment of this field is needed yet challenging. This work aims to provide a technically thorough account of a subarea of AI4Science; namely, AI for quantum, atomistic, and continuum systems. These areas aim at understanding the physical world from the subatomic (wavefunctions and electron density), atomic (molecules, proteins, materials, and interactions), to macro (fluids, climate, and subsurface) scales and form an important subarea of AI4Science. A unique advantage of focusing on these areas is that they largely share a common set of challenges, thereby allowing a unified and foundational treatment. A key common challenge is how to capture physics first principles, especially symmetries, in natural systems by deep learning methods. We provide an in-depth yet intuitive account of techniques to achieve equivariance to symmetry transformations. We also discuss other common technical challenges, including explainability, out-of-distribution generalization, knowledge transfer with foundation and large language models, and uncertainty quantification. To facilitate learning and education, we provide categorized lists of resources that we found to be useful. We strive to be thorough and unified and hope this initial effort may trigger more community interests and efforts to further advance AI4Science.

cs.LG