SearcharxivSearch

arXiv subjects

Dawei Su

Publications and source records attributed to Dawei Su.

8 recordsLinked to original sources

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

cs.CV

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation

Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities. Typically, multiple 2D views are sampled from a 3D point cloud and fed into pre-trained VLMs to answer a given question. This paradigm highlights the critical role of input context quality and raises the challenge of retaining as many task-relevant 3D details as possible under a limited input budget. We propose \texttt{KeyVT}, a hierarchical approach for input context collection at both the view and token levels. Specifically, we combine pixel features with camera parameters and assess view importance based on both semantic content and geometric position, resulting in spatially consistent and task-relevant views. Furthermore, we address redundancy among patches across selected views by identifying representative tokens under the optimal transport (OT) framework, where view tokens and key tokens are formulated as two discrete distributions in the embedding space. These key tokens are expected to cover all view features by minimizing the OT distance. We evaluate our framework on three widely used benchmarks, demonstrating significant improvements over existing tuning-free methods and performance comparable to training-based approaches.

cs.CV

Improving Sparse Autoencoder with Dynamic Attention

Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts. However, identifying the optimal level of sparsity for each neuron remains challenging in practice: excessive sparsity can lead to poor reconstruction, whereas insufficient sparsity may harm interpretability. While existing activation functions such as ReLU and TopK provide certain sparsity guarantees, they typically require additional sparsity regularization or cherry-picked hyperparameters. We show in this paper that dynamically sparse attention mechanisms using sparsemax can bridge this trade-off, due to their ability to determine the activation numbers in a data-dependent manner. Specifically, we first explore a new class of SAEs based on the cross-attention architecture with the latent features as queries and the learnable dictionary as the key and value matrices. To encourage sparse pattern learning, we employ a sparsemax-based attention strategy that automatically infers a sparse set of elements according to the complexity of each neuron, resulting in a more flexible and general activation function. Through comprehensive evaluation and visualization, we show that our approach successfully achieves lower reconstruction loss while producing high-quality concepts, particularly in top-n classification tasks.

cs.LG

RETLLM: Training and Data-Free MLLMs for Multimodal Information Retrieval

Multimodal information retrieval (MMIR) has gained attention for its flexibility in handling text, images, or mixed queries and candidates. Recent breakthroughs in multimodal large language models (MLLMs) boost MMIR performance by incorporating MLLM knowledge under the contrastive finetuning framework. However, they suffer from pre-training inconsistency and require large datasets. In this work, we introduce a novel framework, RetLLM, designed to query MLLMs for MMIR in a training- and data-free manner. Specifically, we formulate MMIR as a similarity score generation task and prompt MLLMs to directly predict retrieval scores in a coarse-then-fine pipeline. At the coarse stage, a top-k filtering strategy builds a small yet high-quality candidate pool for each query, enabling MLLMs to focus on semantically relevant candidates. Subsequently, the retrieval score is predicted by feeding both the query and candidate into MLLMs at the fine stage. Importantly, we propose a visual enhancement module during reasoning to help MLLMs re-pick forgotten visuals, improving retrieval. Extensive experiments on MMIR benchmarks show that RetLLM outperforms fine-tuned models. Ablation studies further verify each component. Our work demonstrates that MLLMs can achieve strong MMIR performance without any training, highlighting their inherent multimodal reasoning ability in a simple, scalable framework. We release our code at: https://github.com/alivecat05/RETLLM

cs.IR

Pressure-mediated crystalline g-C$_3$N$_4$ with enhanced spatial charge transport for solar H$_2$ evolution and photocathodic protection of 304 stainless steels

Conjugated polymeric g-C$_3$N$_4$ has emerged as a leading semiconductor for solar-to-chemical energy conversion due to its unique electronic band structure, robust physicochemical stability, and environmental benignity. However, defect engineering-while effective at enhancing visible-light absorption and charge separation-often introduces excessive dangling bonds and lattice disorder, which exacerbate carrier recombination and impair light harvesting. High crystallinity offers a complementary route to improve spatial charge transport, yet strategies that concurrently optimize crystallinity and surface defects remain underexplored. Here we report a pressure-mediated ion thermal synthesis of high-crystalline g-C$_3$N$_4$ (CCN-P) using a NaCl/KCl eutectic salt under elevated pressure. The molten salt facilitates in-plane and cross-plane crystal growth, while applied pressure reduces interlayer spacing and shortens photocarrier pathways. This dual modulation yields CCN-P with balanced surface defects (-CN and -NHx), an electron-trapping resistance (Rtrap) of 11.36 k$\Omega$ cm$^2$ and a photocarrier decay rate constant of 0.013 s$^{-1}$. CCN-P achieves a hydrogen evolution rate of 2168.8 $\mu$mol g$^{-1}$ h$^{-1}$ and delivers 78.5% dark photocathodic protection of 304 stainless steel over 7500 s, outperforming bulk and conventionally crystalline g-C$_3$N$_4$. This straightforward pressure-ion thermal approach provides a versatile platform for tailoring crystalline frameworks and defect distributions in polymeric semiconductors for efficient solar energy conversion.

cond-mat.mtrl-sci

Adaptive Robust Energy Management Strategy for Campus-Based Commercial Buildings Considering Comprehensive Comfort Levels

Neglecting consumers' comfort always leads to failure or slow-response to demand response request. In this paper, we propose several comprehensive comfort level models for various appliances in campus-based commercial buildings (CBs). The objective of the proposed system is to minimize O\&M costs of campus-based CBs and maximize various comfort levels simultaneously under the worst-case scenarios. Adaptive robust optimization (ARO) is leveraged to handle various uncertainties within the proposed system: (i) demand response signals sending from the distribution system operator (DSO); (ii) arrival state-of-charge (SoC) conditions of plug-in electric vehicles (PEVs); (iii) power outputs of renewable energy sources (RESs); and (iv) load demand of other appliances. Benders decomposition, such as column-and-constraint generation (C\&CG) algorithm, is used to solve the reformulated NP-hard min-max problem. Extensive simulation results demonstrate the effectiveness of the proposed optimal energy management strategy for campus-based CBs in both minimizing O\&M costs and maximizing comprehensive comfort levels.

math.OC

Online Coherence Identification Using Dynamic Time Warping for Controlled Islanding

Controlled islanding is considered to be the last countermeasure to prevent system-wide blackouts in case of cascading failures. It splits the system into self-sustained islands to maintain transient stability at the expense of possible loss of load. Generator coherence identification is critical to controlled islanding scheme as it helps identify the optimal cut-set to maintain system transient stability. This paper presents a novel approach for online generator coherency identification using phasor measurement unit (PMU) data and dynamic time warping (DTW). Results from the coherence identification are used to further cluster non-generator buses using spectral clustering with the objective of minimizing power flow disruption. The proposed approach is validated and compared to existing methods on the IEEE 39-bus system, through which its advantages are demonstrated.

eess.SY

PMU Assisted Power System Parameter Calibration at Jiangsu Electric Power Company

An online PMU-assisted Power System Parameter Calibration System (PSPCS) was recently developed and implemented at State Grid Jiangsu Electric Power Company (JEPC). PSPCS leverages high-resolution PMU data and data mining techniques to perform online screening of the EMS and Production Management System (PMS) databases for data cleaning, model validation, and parameter calibration. PSPCS calculates transmission line and generator parameters on a regular real-time basis and compares the results with databases to identify record(s) with significant discrepancy, if any. Once consistent discrepancy is observed, the system will raise a flag and further investigation will be initiated, including a novel density-based spatial clustering procedure for parameter/data calibration. A novel metric is proposed to quantify the credibility of PMU-based parameter identification. This paper discusses the proposed methodologies, challenges, as well as implementation issues identified during the development and deployment of PSPCS.

eess.SY