Searcharxiv⌕ Search

arXiv subjects

Yating Wang

Publications and source records attributed to Yating Wang.

At least 37 records · Page 2Linked to original sources

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs

Open-ended question answering (QA) is a key task for evaluating the capabilities of large language models (LLMs). Compared to closed-ended QA, it demands longer answer statements, more nuanced reasoning processes, and diverse expressions, making refined and interpretable automatic evaluation both crucial and challenging. Traditional metrics like ROUGE and BERTScore struggle to capture semantic similarities due to different patterns between model responses and reference answers. Current LLM-based evaluation approaches, such as pairwise or listwise comparisons of candidate answers, lack intuitive interpretability. While pointwise scoring of each response provides some descriptions, it fails to adapt across different question contents. Most notably, existing methods overlook the distinction between factoid and non-factoid questions. To address these challenges, we propose \textbf{MinosEval}, a novel evaluation method that first distinguishes open-ended questions and then ranks candidate answers using different evaluation strategies. For factoid questions, it applies an adaptive key-point scoring strategy, while for non-factoid questions, it uses an instance-aware listwise ranking strategy. Experiments on multiple open-ended QA datasets, including self-built ones with more candidate responses to complement community resources, show that MinosEval better aligns with human annotations and offers more interpretable results.

cs.CL↗

3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations

Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose a novel method that addresses all the aforementioned demands. In specific, we introduce an expressive and compact representation that encodes texture-related attributes of the 3D Gaussians in the tensorial format. We store appearance of neutral expression in static tri-planes, and represents dynamic texture details for different expressions using lightweight 1D feature lines, which are then decoded into opacity offset relative to the neutral face. We further propose adaptive truncated opacity penalty and class-balanced sampling to improve generalization across different expressions. Experiments show this design enables accurate face dynamic details capturing while maintains real-time rendering and significantly reduces storage costs, thus broadening the applicability to more scenarios.

cs.CV↗

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning

Learning from multiple domains is a primary factor that influences the generalization of a single unified robot system. In this paper, we aim to learn the trajectory prediction model by using broad out-of-domain data to improve its performance and generalization ability. Trajectory model is designed to predict any-point trajectories in the current frame given an instruction and can provide detailed control guidance for robotic policy learning. To handle the diverse out-of-domain data distribution, we propose a sparsely-gated MoE (\textbf{Top-1} gating strategy) architecture for trajectory model, coined as \textbf{Tra-MoE}. The sparse activation design enables good balance between parameter cooperation and specialization, effectively benefiting from large-scale out-of-domain data while maintaining constant FLOPs per token. In addition, we further introduce an adaptive policy conditioning technique by learning 2D mask representations for predicted trajectories, which is explicitly aligned with image observations to guide action prediction more flexibly. We perform extensive experiments on both simulation and real-world scenarios to verify the effectiveness of Tra-MoE and adaptive policy conditioning technique. We also conduct a comprehensive empirical study to train Tra-MoE, demonstrating that our Tra-MoE consistently exhibits superior performance compared to the dense baseline model, even when the latter is scaled to match Tra-MoE's parameter count.

cs.RO↗

SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with intrinsic spatial understanding. We present the most comprehensive evaluation of embodied representation learning to date, covering 268 tasks across 8 simulators with diverse policies in both single-task and language-conditioned multi-task scenarios. The results are compelling: SPA consistently outperforms more than 10 state-of-the-art representation methods, including those specifically designed for embodied AI, vision-centric tasks, and multi-modal applications, while using less training data. Furthermore, we conduct a series of real-world experiments to confirm its effectiveness in practical scenarios. These results highlight the critical role of 3D spatial awareness for embodied representation learning. Our strongest model takes more than 6000 GPU hours to train and we are committed to open-sourcing all code and model weights to foster future research in embodied representation learning. Project Page: https://haoyizhu.github.io/spa/.

cs.CV↗

ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait Image

While recent works have achieved great success on image-to-3D object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by text descriptions and image-based methods struggled to produce high-quality head geometry. To handle this challenging problem, we propose a novel framework, ID-Sculpt, to generate high-quality 3D heads while preserving their identities. Our work incorporates the identity information of the portrait image into three parts: 1) geometry initialization, 2) geometry sculpting, and 3) texture generation stages. Given a reference portrait image, we first align the identity features with text features to realize ID-aware guidance enhancement, which contains the control signals representing the face information. We then use the canny map, ID features of the portrait image, and a pre-trained text-to-normal/depth diffusion model to generate ID-aware geometry supervision, and 3D-GAN inversion is employed to generate ID-aware geometry initialization. Furthermore, with the ability to inject identity information into 3D head generation, we use ID-aware guidance to calculate ID-aware Score Distillation (ISD) for geometry sculpting. For texture generation, we adopt the ID Consistent Texture Inpainting and Refinement which progressively expands the view for texture inpainting to obtain an initialization UV texture map. We then use the ID-aware guidance to provide image-level supervision for noisy multi-view images to obtain a refined texture map. Extensive experiments demonstrate that we can generate high-quality 3D heads with accurate geometry and texture from a single in-the-wild portrait image.

cs.CV↗

Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation

Text-to-motion generation is a crucial task in computer vision, which generates the target 3D motion by the given text. The existing annotated datasets are limited in scale, resulting in most existing methods overfitting to the small datasets and unable to generalize to the motions of the open domain. Some methods attempt to solve the open-vocabulary motion generation problem by aligning to the CLIP space or using the Pretrain-then-Finetuning paradigm. However, the current annotated dataset's limited scale only allows them to achieve mapping from sub-text-space to sub-motion-space, instead of mapping between full-text-space and full-motion-space (full mapping), which is the key to attaining open-vocabulary motion generation. To this end, this paper proposes to leverage the atomic motion (simple body part motions over a short time period) as an intermediate representation, and leverage two orderly coupled steps, i.e., Textual Decomposition and Sub-motion-space Scattering, to address the full mapping problem. For Textual Decomposition, we design a fine-grained description conversion algorithm, and combine it with the generalization ability of a large language model to convert any given motion text into atomic texts. Sub-motion-space Scattering learns the compositional process from atomic motions to the target motions, to make the learned sub-motion-space scattered to form the full-motion-space. For a given motion of the open domain, it transforms the extrapolation into interpolation and thereby significantly improves generalization. Our network, $DSO$-Net, combines textual $d$ecomposition and sub-motion-space $s$cattering to solve the $o$pen-vocabulary motion generation. Extensive experiments demonstrate that our DSO-Net achieves significant improvements over the state-of-the-art methods on open-vocabulary motion generation. Code is available at https://vankouf.github.io/DSONet/.

cs.CV↗

Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning

In robot learning, the observation space is crucial due to the distinct characteristics of different modalities, which can potentially become a bottleneck alongside policy design. In this study, we explore the influence of various observation spaces on robot learning, focusing on three predominant modalities: RGB, RGB-D, and point cloud. We introduce OBSBench, a benchmark comprising two simulators and 125 tasks, along with standardized pipelines for various encoders and policy baselines. Extensive experiments on diverse contact-rich manipulation tasks reveal a notable trend: point cloud-based methods, even those with the simplest designs, frequently outperform their RGB and RGB-D counterparts. This trend persists in both scenarios: training from scratch and utilizing pre-training. Furthermore, our findings demonstrate that point cloud observations often yield better policy performance and significantly stronger generalization capabilities across various geometric and visual conditions. These outcomes suggest that the 3D point cloud is a valuable observation modality for intricate robotic tasks. We also suggest that incorporating both appearance and coordinate information can enhance the performance of point cloud methods. We hope our work provides valuable insights and guidance for designing more generalizable and robust robotic models. Codes are available at https://github.com/HaoyiZhu/PointCloudMatters.

cs.RO↗

Thermodynamic Geometric Control of Active Matter

Active matter represents a class of non-equilibrium systems that constantly dissipate energy to produce directed motion. The thermodynamic control of active matter holds great potential for advancements in synthetic molecular motors, targeted drug delivery, and adaptive smart materials. However, the inherently non-equilibrium nature of active matter poses a significant challenge in achieving optimal control with minimal energy cost. In this work, we extend the concept of thermodynamic geometry, traditionally applied to passive systems, to active matter, proposing a systematic geometric framework for minimizing energy cost in non-equilibrium driving processes. We derive a cost metric that defines a Riemannian manifold for control parameters, enabling the use of powerful geometric tools to determine optimal control protocols. The geometric perspective reveals that, unlike in passive systems, minimizing energy cost in active systems involves a trade-off between intrinsic and external dissipation, leading to an optimal transportation speed that coincides with the self-propulsion speed of active matter. This insight enriches the broader concept of thermodynamic geometry. We demonstrate the application of this approach by optimizing the performance of an active monothermal engine within this geometric framework.

cond-mat.stat-mech↗

Physical Similarity of Fluid Flow in Bimodal Porous Media: Part 1 -- Basic Model and Solution Characteristics

Fluid flow through bimodal porous media, characterized by a distinct separation in pore size distribution, is critical in various scientific and engineering applications, including groundwater management, oil and gas production, and carbon sequestration. This note delves into the physical similarity of fluid flow within such media, bridging the gap between microscale phenomena and macroscale observations. We present a representative mathematical model that conceptualizes bimodal porous media as a double-continuum system, distinguishing between macroporous and microporous regions. The model captures the complex interactions between these regions, particularly focusing on the challenges of modeling fluid flow when there is significant disparity in pore sizes. By employing a heuristic approach grounded in pore-scale tomography, we derive governing equations that describe fluid flow and analyze the solution characteristics. The results reveal unique features of the fluid flow in bimodal systems, such as the occurrence of boundary discontinuities and the delayed transient response, which are not observed in conventional porous media. This work provides ground for further studies in bimodal porous media, offering insights that could enhance predictive modeling and optimization in various applications concerning porous media with similar bimodal pore size distributions.

physics.flu-dyn↗

Asymptotic-preserving neural networks for the semiconductor Boltzmann equation and its application on inverse problems

In this paper, we develop the Asymptotic-Preserving Neural Networks (APNNs) approach to study the forward and inverse problem for the semiconductor Boltzmann equation. The goal of the neural network is to resolve the computational challenges of conventional numerical methods and multiple scales of the model. To guarantee the network can operate uniformly in different regimes, it is desirable to carry the Asymptotic-Preservation (AP) property in the learning process. In a micro-macro decomposition framework, we design such an AP formulation of loss function. The convergence analysis of both the loss function and its neural network is shown, based on the Universal Approximation Theorem and hypocoercivity theory of the model equation. We show a series of numerical tests for forward and inverse problems of both the semiconductor Boltzmann and the Boltzmann-Poisson system to validate the effectiveness of our proposed method, which addresses the significance of the AP property when dealing with inverse problems of multiscale Boltzmann equations especially when only sparse or partially observed data are available.

math-ph↗

Learning-based Multi-continuum Model for Multiscale Flow Problems

Multiscale problems can usually be approximated through numerical homogenization by an equation with some effective parameters that can capture the macroscopic behavior of the original system on the coarse grid to speed up the simulation. However, this approach usually assumes scale separation and that the heterogeneity of the solution can be approximated by the solution average in each coarse block. For complex multiscale problems, the computed single effective properties/continuum might be inadequate. In this paper, we propose a novel learning-based multi-continuum model to enrich the homogenized equation and improve the accuracy of the single continuum model for multiscale problems with some given data. Without loss of generalization, we consider a two-continuum case. The first flow equation keeps the information of the original homogenized equation with an additional interaction term. The second continuum is newly introduced, and the effective permeability in the second flow equation is determined by a neural network. The interaction term between the two continua aligns with that used in the Dual-porosity model but with a learnable coefficient determined by another neural network. The new model with neural network terms is then optimized using trusted data. We discuss both direct back-propagation and the adjoint method for the PDE-constraint optimization problem. Our proposed learning-based multi-continuum model can resolve multiple interacted media within each coarse grid block and describe the mass transfer among them, and it has been demonstrated to significantly improve the simulation results through numerical experiments involving both linear and nonlinear flow equations.

math.NA↗

DGRC: An Effective Fine-tuning Framework for Distractor Generation in Chinese Multi-choice Reading Comprehension

When evaluating a learner's knowledge proficiency, the multiple-choice question is an efficient and widely used format in standardized tests. Nevertheless, generating these questions, particularly plausible distractors (incorrect options), poses a considerable challenge. Generally, the distractor generation can be classified into cloze-style distractor generation (CDG) and natural questions distractor generation (NQDG). In contrast to the CDG, utilizing pre-trained language models (PLMs) for NQDG presents three primary challenges: (1) PLMs are typically trained to generate ``correct'' content, like answers, while rarely trained to generate ``plausible" content, like distractors; (2) PLMs often struggle to produce content that aligns well with specific knowledge and the style of exams; (3) NQDG necessitates the model to produce longer, context-sensitive, and question-relevant distractors. In this study, we introduce a fine-tuning framework named DGRC for NQDG in Chinese multi-choice reading comprehension from authentic examinations. DGRC comprises three major components: hard chain-of-thought, multi-task learning, and generation mask patterns. The experiment results demonstrate that DGRC significantly enhances generation performance, achieving a more than 2.5-fold improvement in BLEU scores.

cs.CL↗

MotionMaster: Training-free Camera Motion Transfer For Video Generation

The emergence of diffusion models has greatly propelled the progress in image and video generation. Recently, some efforts have been made in controllable video generation, including text-to-video generation and video motion control, among which camera motion control is an important topic. However, existing camera motion control methods rely on training a temporal camera module, and necessitate substantial computation resources due to the large amount of parameters in video generation models. Moreover, existing methods pre-define camera motion types during training, which limits their flexibility in camera control. Therefore, to reduce training costs and achieve flexible camera control, we propose COMD, a novel training-free video motion transfer model, which disentangles camera motions and object motions in source videos and transfers the extracted camera motions to new videos. We first propose a one-shot camera motion disentanglement method to extract camera motion from a single source video, which separates the moving objects from the background and estimates the camera motion in the moving objects region based on the motion in the background by solving a Poisson equation. Furthermore, we propose a few-shot camera motion disentanglement method to extract the common camera motion from multiple videos with similar camera motions, which employs a window-based clustering technique to extract the common features in temporal attention maps of multiple videos. Finally, we propose a motion combination method to combine different types of camera motions together, enabling our model a more controllable and flexible camera control. Extensive experiments demonstrate that our training-free approach can effectively decouple camera-object motion and apply the decoupled camera motion to a wide range of controllable video generation tasks, achieving flexible and diverse camera motion control.

cs.CV↗

Partially explicit splitting scheme with explicit-implicit-null method for nonlinear multiscale flow problems

In this work, we present an efficient approach to solve nonlinear high-contrast multiscale diffusion problems. We incorporate the explicit-implicit-null (EIN) method to separate the nonlinear term into a linear term and a damping term, and then utilise the implicit and explicit time marching scheme for the two parts respectively. Due to the multiscale property of the linear part, we further introduce a temporal partially explicit splitting scheme and construct suitable multiscale subspaces to speed up the computation. The approximated solution is splitted into these subspaces associated with different physics. The temporal splitting scheme employs implicit discretization in the subspace with small dimension that representing the high-contrast property and uses explicit discretization for the other subspace. We exploit the stability of the proposed scheme and give the condition for the choice of the linear diffusion coefficient. The convergence of the proposed method is provided. Several numerical tests are performed to show the efficiency and accuracy of the proposed approach.

math.NA↗

On a neural network approach for solving potential control problem of the semiclassical Schrödinger equation

Robust control design for quantum systems is a challenging and key task for practical technology. In this work, we apply neural networks to learn the control problem for the semiclassical Schrödinger equation, where the control variable is the potential given by an external field that may contain uncertainties. Inspired by a relevant work [29], we incorporate the sampling-based learning process into the training of networks, while combining with the fast time-splitting spectral method for the Schrödinger equation in the semiclassical regime. The numerical results have shown the efficiency and accuracy of our proposed deep learning approach.

math.NA↗

Interactions enhance dispersion in fluctuating channels via emergent flows

Understanding particle motion in narrow channels is essential to guide progress in numerous applications, from filtration to vascular transport. Thermal or active fluctuations of channel walls for fluid-filled channels can slow down or increase the dispersion of tracer particles. Entropic trapping in the wall bulges slows dispersion, and hydrodynamic flows induced by wall fluctuations enhance dispersion. Previous studies primarily concentrated on the case of a single Brownian tracer either embedded in an incompressible fluid or in the ideal case where the presence of fluid is ignored. Here we address the question of what happens when there is a large ensemble of interacting Brownian tracers -- a common situation in applications. Introducing repulsive interactions between the tracer particles, while ignoring the presence of a background fluid, leads to an effective flow field. This flow field enhances tracer dispersion, a phenomenon strongly reminiscent of that seen in incompressible background fluid. We characterise the dispersion by the long-time diffusion coefficient of a single tracer, numerically and analytically with a mean-field density functional analysis. We find a surprising effect where an increased particle density enhances the diffusion coefficient, challenging the notion that crowding effects tend to reduce diffusion. Here, inter-particle interactions push particles closer to the fluctuating channel walls. Then, interactions between the fluctuating wall and the now-nearby particles drive particle mixing. Our mechanism is sufficiently general that we expect it to apply to various systems. In addition, the perturbation theory we derive quantifies dispersion in generic advection-diffusion systems with arbitrary spatiotemporal drift.

cond-mat.soft↗

Escape rate of an active Brownian particle in a rough potential

We discuss escape problem with the consideration of both the activity of particles and the roughness of potentials. we derive analytic expressions for the escape rate of a Brownian particle (ABP) in two types of rough potentials by employing the effective equilibrium approach and the Zwanzig method. We find that activity enhances the escape rate, but both the oscillating perturbation and the random amplitude hinder escaping.

cond-mat.stat-mech↗

Entropy rate of random walks on complex networks under stochastic resetting

Stochastic processes under resetting at random times have attracted a lot of attention in recent years and served as illustrations of nontrivial and interesting static and dynamic features of stochastic dynamics. In this paper, we aim to address how the entropy rate is affected by stochastic resetting in discrete-time Markovian processes, and explore nontrivial effects of the resetting in the mixing properties of a stochastic process. In particular, we consider resetting random walks on complex networks and compute the entropy rate as a function of the resetting probability. Interestingly, we find that the entropy rate can show a nonmonotonic dependence on the resetting probability. There exists an optimal resetting probability for which the entropy rate reaches a maximum. We also show that the maximum entropy rate can be larger than that of the maximal-entropy random walks on the same topology. Our study provides a new nontrivial effect of stochastic resetting on nonequilibrium statistical physics.

cond-mat.stat-mech↗