SearcharxivSearch

arXiv subjects

Yun Wang

Publications and source records attributed to Yun Wang.

At least 19 recordsLinked to original sources

Forced self-similar solutions to the stationary Navier--Stokes equations in a half-space

We study axisymmetric self-similar solutions to the stationary Navier--Stokes equations in the half-space with the no-slip boundary condition, driven by an axisymmetric (-3)-homogeneous external force. If the tangential curl of the force on the unit sphere is sufficiently small, we prove the existence of a unique small solution; when the force is swirl-free, the solution is automatically swirl-free and unique. For a swirl-free external force $\boldsymbol{F}$, we introduce a scaling parameter $\lambda$ and consider the system with force $\lambda\boldsymbol{F}$; we prove that solutions exist precisely for $\lambda$ in an open interval containing zero. The same approach extends to solid cones with the no-slip boundary condition, where narrower opening angles allow the existence of solutions under larger external forces.

math.AP

Collascope: Supporting Serendipitous Asset Exploration for Collage-Based Storytelling

Collage-based storytelling requires visual elements that support emerging narratives and inspire creative reinterpretation. Existing tools, however, rely largely on keyword- and image-based retrieval, offering limited support for serendipitous exploration beyond existing assets. We introduce Collascope, an interactive system that helps creators (1) concretize story intent with interactive element groups, (2) expand the exploration space based on concepts or cutouts towards conceptual and visual dimensions, and (3) develop grounded, traceable ideas in parallel with collage composition. Collascope's attribute-aware visual retrieval method, instantiated with collage-relevant visual dimensions, enables creators to retrieve cutouts through dimension-specific visual projections rather than holistic similarity. In a within-subject study (N=12) against a conventional search baseline, our participants used unexpected results and even gaps in the asset collection to redirect narratives, shift tone, and enrich compositions. Scene Parts helped organize exploration into manageable subtasks, while participants used association in distinct ways depending on whether exploration was guided by a clear goal, an evolving story, or visual intuition.

cs.HC

Minute-Scale Training for Microrobot Navigation

Microrobots hold significant potential for various applications, where targeted navigation is a basic requirement. Deep reinforcement learning (DRL) has recently emerged as a powerful paradigm for fully autonomous microrobot navigation. Yet, current DRL-based approaches pay limited attention to learning efficiency and effectiveness, requiring hours to days for model training. Consequently, this impedes both rapid practical deployment and parameter optimization. To address these challenges, we present a learning framework that enables effective microrobot navigation policies to be trained within minutes. In the proposed framework, we develop a fully vectorized simulator with more than 10,000 artificial vascular environments, parallelizing dynamics, LiDAR-inspired perception, and feasibility checks across thousands of environments to achieve roughly 190,000 transitions per second. To achieve effectiveness in fast training, we propose a task-shaping-regularization (TSR) reward framework. The TSR framework accelerates convergence, improves final performance, reduces action variation by at least 33.7%, and increases obstacle clearance by at least 2.1% across all evaluated scenarios. Results show that the proposed learning framework reduces training time to under 10 minutes, while supporting zero-shot deployment across distinct microrobot types and navigation scenarios. Collectively, this framework can substantially shorten the design loop and accelerate the deployment of autonomous microrobots.

cs.RO

SkinSpline: A Body-Attached Skeleton-Supported Haptic Interface for Continuous Skin Deformation through Physical Interpolation

We present SkinSpline, a body-attached skeleton-supported haptic interface that renders continuous skin deformation through physical interpolation of sparse mechanical actuation. SkinSpline combines a low-resolution array of rack-and-pinion linear actuators with an elastic interlocking skeleton that transforms discrete actuator motions into smooth surface deformation, enabling continuous cutaneous feedback without dense actuator arrays. The system includes a modular hardware architecture, a configurable control pipeline, and a visual interface supporting real-time configuration and actuation. We demonstrate SkinSpline through multiple scenarios, including wave rendering, video-synchronized rhythmic touch, visually driven water-wave feedback in VR, and sensor-based remote touch reproduction. SkinSpline explores an alternative approach to continuous on-body haptic rendering by leveraging structural coupling between sparse actuation and deformable surfaces.

cs.HC

InnoText: A Unified Model for Visual Text Generation and Editing

Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing remains relatively underexplored. Unlike general image generation, visual text tasks demand precise structural regularity and legibility, which may pose additional challenges for small-scale text and non-Latin scripts such as Chinese. Existing UNet-based models often struggle to produce clear and coherent text, while DiT-based models, though more expressive, are typically limited to a single task, which may lead to redundant training pipelines, inconsistent visual styles, and reduced cross-task generalization. To address these challenges, we propose InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model. We introduce a Font Size-Aware Modulation (FSAM) module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a Task-Specific Region Weighted Loss for adaptive optimization. To support training and evaluation, we also construct a high-quality bilingual (English-Chinese) visual text dataset covering diverse fonts, sizes, and backgrounds. Experimental results demonstrate that our method achieves superior generation accuracy and editing quality, producing visually appealing and realistic text images.

cs.CV

fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction

Reconstructing dynamic human faces from brain activity provides a powerful way to study how the mind perceives identity, expression, and facial motion. However, progress in fMRI-based face decoding has been limited by scarce controlled, high-resolution neural datasets and by methods that struggle to recover both identity-specific appearance and time-varying facial dynamics. We present fMRI-Face, the first fMRI dataset paired with controllable full-HD digital human facial videos rendered at 1920$\times$1080 resolution. During scanning, participants watched photorealistic, background-free facial videos with controlled identity, expression, and head pose, while fMRI activity was recorded. The resulting dataset contains 62,856 paired fMRI-video samples, providing a structured resource for studying dynamic face perception and reconstruction. Building on this dataset, we propose fMRI2Face, a geometry-guided neural video decoding framework for reconstructing facial videos from fMRI signals. fMRI2Face derives two complementary neural controls from brain activity: Brain-derived Appearance Context, which captures global identity-related visual attributes, and Morphable 3D Facial Control, which provides explicit geometry-aware guidance for pose, expression, and non-rigid facial dynamics. These controls are integrated through Neural-Controlled Video Diffusion with auxiliary latent completion, enabling high-fidelity facial video reconstruction directly from brain activity. Experiments show that fMRI2Face consistently improves reconstruction fidelity, identity preservation, facial geometry, and motion consistency over representative neural decoding baselines. Together, fMRI-Face and fMRI2Face establish a controlled platform for studying dynamic face perception and provide a new benchmark for fMRI-based digital human reconstruction.

cs.CV

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.

cs.CV

Hybrid Rigid-Soft Robotic Gripper with Shape Adaptation, Uniform Force Distribution, and Self-Locking Capabilities

Conventional robotic grippers face a significant challenge in agricultural automation: the trade-off between compliant, adaptive grasping, pressure balancing among all joints, and high load capacity, often at the cost of high energy consumption. This paper presents a novel hybrid rigid-soft gripper that integrated low-cost, membrane-based pneumatic actuators with 3D-printed dual ratchet-pawl mechanisms to simultaneously achieve shape adaptation, uniform force distribution, and energy-free self-locking. The dual-ratchet structure assembled in an offset configuration significantly increased the angular resolution of the joint locking mechanism. Key experimental results demonstrated the gripper's superior performance: a remarkable maximum load capacity of 4200 g, far exceeding that of conventional soft grippers (45-210 g); more uniform force distribution across object sizes (1.75-35.29% difference ratio) compared to a rigid gripper (56.77-66.44%), with peak contact forces remaining below surface damage thresholds; and a 50.05% reduction in total energy consumption to 42.6 J per grasp cycle, achieved by eliminating the need for continuous pneumatic pressure through the self-locking mechanism, compared to 85.28 J for a conventional soft gripper. The combination of additive manufacturing for ratchets and commercially available materials for pneumatic chambers ensured a low-cost and easily fabricated design. These findings validated that the proposed gripper successfully bridged the gap between soft compliance and rigid reliability, offering a robust and efficient solution for scalable agricultural harvesting and manipulation tasks.

cs.RO

Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors

Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recent Gaussian Scene Graph (GSG) methods achieve impressive visual quality, they heavily rely on precise priors, such as accurate camera poses and LiDAR depth, or manual annotations. When initialized with noisy priors estimated from in-the-wild videos, existing GSG methods suffer from optimization ambiguity (e.g., entangling camera and agent poses) and topological failures (e.g., missing objects), causing severe rendering artifacts. To enable robust in-the-wild reconstruction, we introduce Adaptive Gaussian Graph (AGG), a self-correcting 4D framework. Our Semantically-Guided Tick-Tock Strategy leverages 2D foundation features to explicitly decouple static background and camera pose updates from dynamic agent learning. Concurrently, our Adaptive Topology Evolution module actively rectifies graph structures by spawning missing agents, reassigning misclassified Gaussians, and pruning false positives. To rigorously evaluate this in-the-wild setting, we introduce Wild-30, a challenging benchmark of internet and generative videos. Extensive experiments on KITTI and Wild-30 validate that AGG consistently outperforms state-of-the-art approaches in visual fidelity and robustness under noisy priors.

cs.CV

HBM Is Not All You Need: Efficient Disaggregated LLM Serving across Memory-heterogeneous Accelerators

LLM inference comprises a compute-bound prefill phase and a memory-bound decode phase, and recent systems disaggregate them onto separate hardware. Yet today's datacenter GPUs rely on costly HBM whose bandwidth sits almost entirely idle during prefill. LLM serving across memory-heterogeneous accelerators (MemHA) pairs GDDR-based accelerators for prefill with HBM-based GPUs for decode, promising lower cost without sacrificing performance. Pushed to its most economical form, MemHA serving is inherently cross-vendor, since the best-suited chip for each phase may come from a different vendor. This breaks two assumptions that single-vendor disaggregation takes for granted -- a KV format both ends consume natively, and a shared software stack. We present \textbf{HMA-Serve}, a MemHA-centric disaggregated serving system pairing GDDR-based accelerators for prefill with HBM-based GPUs for decode efficiently. HMA-Serve achieves this through (1) phase-wise quantization, applying vendor-native low precision for high-throughput prefill while keeping decode in high-precision BF16, (2) a compute-transfer pipeline that overlaps each layer's KV cache transfer with later-layer prefill to reduce time-to-first-token (TTFT), and (3) deferred dequantization, shipping raw quantized bytes and reconstructing them lazily on the decode GPU to reduce network bandwidth and HBM usage. Across four Qwen3 models (4B--32B) and three production traces, HMA-Serve delivers up to $3.2\times$ higher goodput than state-of-the-art memory-homogeneous methods and $4.8\times$ higher goodput-per-dollar, with no measurable loss on generation-quality benchmarks.

cs.AR

GRB 250424A: A Case Study of Energy Injection with Multiwavelength Observations

We present a comprehensive multiwavelength analysis of the long-duration gamma-ray burst (GRB) 250424A. Our dataset spans from the prompt gamma-ray emission to late-time optical monitoring, including spectra obtained with the Keck 10\,m telescope. We find that the afterglow light curves display a prominent, simultaneous shallow decay phase in both X-ray and optical bands, followed by an achromatic transition to a standard decay regime. The broadband spectral energy distributions are well-modeled by a single power-law function, indicating a common synchrotron origin for the emission across frequencies. We interpret the afterglow evolution within the framework of a relativistic forward shock refreshed by continuous energy injection. This scenario successfully reproduces the observed temporal and spectral behavior, yielding an isotropic equivalent kinetic energy of $E_{\rm K,iso} \approx 5.5 \times 10^{52}$ erg and an injection index of $q\approx 0.34$ in a constant-density circumburst environment. The shallow decay phase is consistent with sustained energy injection lasting $\sim$ 9 ks. Despite the relatively low redshift, late-time optical observations reveal no distinct supernova component; however, our derived upper limits do not strictly rule out the presence of a typical GRB-associated supernova.

astro-ph.HE

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence in egocentric video streams. UCS-Bench targets a new problem that emphasizes dynamic spatial reasoning, long-term memory, and their alignment with users' real-time locations. We propose DirectMe, a framework that incrementally constructs and maintains a structured spatial memory from streaming egocentric observations. DirectMe enables robust tracking and recall of object locations, all relative to the user's movement over time. By tightly coupling visual perception with memory updates and spatial reasoning, our approach supports long-horizon queries that require recalling interactions, resolving viewpoint-induced ambiguities, and adapting to dynamic scenes. Our experiments show that DirectMe significantly improves the spatial reasoning of leading multimodal LLMs; it also surpasses many spatially aware and long-form streaming video models. We hope our benchmark and solution will advance spatial intelligence research for egocentric AI assistants. Data and code are available at https://github.com/cocowy1/UCS-Bench.

cs.CV

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, however, requires collaborators to continuously maintain and align mental models of their own reasoning,partners' intentions, and shared goals during the collaborative process. Today's agents rarely develop such capabilities since they are primarily optimized for task completion, and the community lacks authentic human collaboration data with action-level mental model annotations that could guide agents toward process-level collaborative competence. To bridge this gap, we present ALMANAC, a dataset of Action-Level Mental model ANnotations for Agent Collaboration built from the Map Task, a classic dyadic routing task from social science. ALMANAC contains 2,987 collaboration actions, each paired with theory-informed mental model annotations that record the participants' self-reasoning, perceived partner intent, and perceived team goal. We benchmark six LLMs on predicting humans' next-turn behavior and mental models. Our results demonstrate ALMANAC's utility in evaluating models' ability to simulate human collaborative behaviors and infer their underlying mental models.

cs.AI

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate through text-based channels much as human teams do. Yet recent study suggests that MAS often falter not because agents lack individual task-solving ability, but because they lack collaborative competence: the capacity to establish common ground, maintain shared task understanding, balance individual and collective incentives, and repair misalignment as interaction unfolds. Decades of research in Computer-Supported Cooperative Work have characterized these requirements for human teams coordinating under constrained communication, yet existing MAS evaluations focus mainly on task outcomes or single-agent proficiency in reasoning, planning, and tool use. To enable a systematic analysis of agents' collaborative competence in MAS, we introduce CollabSim, a configurable simulation framework that combines a theory-grounded definition of collaborative capabilities, controlled manipulation of interaction conditions, and action-level probing of agents' internal states. Experiments across four LLMs show that CollabSim can capture condition effects, separate model performance patterns, and reveal task-dependent effects of agent design.

cs.CL

Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval

Language agents increasingly rely on reusable skills to improve multi-step web automation across related tasks. A growing line of work studies online skill learning, where agents continually induce skills from previous task trajectories and reuse them in future tasks on the fly. However, existing methods mainly reuse skills at the task-level: a fixed set of skills is retrieved based on the initial task instruction and then held fixed throughout execution. This static strategy is misaligned with web execution, where the appropriate next action depends not only on the task goal but also on the current webpage state, which often transitions into situations that the initial skills fail to cover. To address this gap, we propose State-Grounded Dynamic Retrieval (SGDR), an online skill learning method that enables stepwise skill reuse for web agents. SGDR consists of three components: a sliding-window extraction process that turns completed trajectories into reusable sub-procedures invokable at intermediate execution states, a dual text-code representation that connects skill retrieval with executable action, and a state-grounded dynamic retrieval mechanism that matches skills to both the task goal and the current webpage state. Experiments on WebArena across five domains show that SGDR consistently outperforms strong baselines, achieving average success rates of 37.5% with GPT-4.1 and 24.3% with Qwen3-4B, corresponding to relative gains of 10.6% and 10.0% over the strongest baseline, respectively. The code is available at https://github.com/plusnli/skill-dynamic-retrieval.

cs.AI

Revealing the high redshift host galaxy of the short GRB 061201 with JWST

Using deep near-infrared and optical images from JWST and HST, we identify a new host galaxy candidate for GRB 061201. It lies ~2" from the optical afterglow position. Photometric redshift fitting yields z~1.2. We compare the previously proposed host at z=0.111 with the new candidate. The chance-coincidence probability is $P_{cc}=0.18$, above the classical threshold of 0.1 but consistent with a physical association given the extreme depth of JWST imaging. In contrast, evaluated with corresponding JWST observations, the previously claimed host has a lower $P_{cc}=0.11$, which is driven primarily by bright-tail statistics rather than a more plausible association. A high-z origin is favored by three independent lines of evidence. First, for the z=0.111 scenario, the beaming-corrected energy shows GRB 061201 is an outlier of the Ghirlanda ($E_{p,i}-E_\gamma$) relation for short GRBs, while for the z=1.2 scenario, it is well consistent with the Amati relation. Second, deep near-infrared observations rule out a kilonova similar to AT2017gfo at z=0.111. Third, afterglow modeling yields an AIC criterion of $\Delta$AIC=16.35, providing strong evidence for the high-redshift scenario. Assuming the host candidate is the actual host galaxy of GRB 061201, the physical offset is 16.4-16.9 kpc (substantially reduced from ~42 kpc) and the host stellar age is ~2 Gyr, which are consistent with the host population of short GRBs. A low-redshift origin would lead to a very high binary neutron star merger rate of ~1400 Gpc$^{-3}$ yr$^{-1}$, which is contradictory to the gravitational-wave constraint. We suggest that GRB 061201 originates from a moderately high-redshift (z~1.2) host, significantly alleviating this apparent merger rate discrepancy. This case demonstrates the power of deep JWST exposures in revealing the host galaxies of historically hostless GRBs.

astro-ph.HE

Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this bottleneck through evaluation heuristics developed over extensive practice. We ask whether LLMs can learn similar heuristics directly from scoring experience, and formalize this as the concept of assessment skills: item-independent natural-language procedural knowledge that guides LLMs through specific stages of the scoring workflow. Focusing on rubric construction as a first instantiation, we propose an iterative framework that decomposes a skill into a fixed scaffold and learnable item-agnostic rules, refining the rules through LLM-driven diagnosis of scoring errors and validation-gated selection. The framework requires no expert-written rubric. On all ten ASAP-SAS items, optimized skills substantially improve LLM-based scoring and frequently surpass the dataset-provided expert rubric. Cross-item transfer experiments further reveal that learned skills capture both generalizable and item-specific patterns.

cs.CL

Towards precision cosmology with Void x CMB correlations (II): Impact of mock catalogs on the Void x CMB lensing signal

Gravitational lensing by large-scale structure imprints secondary anisotropies on the Cosmic Microwave Background (CMB) that can be exploited to probe cosmology. In particular, cosmic voids produce a characteristic lensing signature detectable through Void x CMB cross-correlations. This signal has been robustly measured in the past but its cosmological constraining power remains limited by the incomplete knowledge of how methodological choices affect its measurement and by its uncertain dependence on cosmological parameters. Using a set of validated Roman mock catalogs, we first quantify how mock construction impacts the measured signal and then forecast the capabilities of Roman, in combination with current and upcoming CMB surveys such as Planck, SO and CMB-S4-like experiments. We analyze the signal-to-noise ratio (S/N) for different void definitions (2D and 3D), stacking approaches (rescaled versus non-rescaled profiles), CMB map filtering schemes and noise levels. In contrast to galaxy and void statistics, we find that the Void x CMB lensing signal is less sensitive to the choice of mock catalog, indicating that future tensions with data are unlikely to stem from mock inaccuracies alone. The highest S/N is achieved for 2D voids with rescaled profiles. We forecast S/N ~13$\sigma$ (8$\sigma$) for 2D (3D) Roman voids combined with Planck, increasing to 22$\sigma$ (13$\sigma$) for SO and 31$\sigma$ (18$\sigma$) for CMB-S4-like surveys. While the cosmological dependence of this observable remains to be quantified, Roman together with next-generation of LSS and CMB surveys opens a path toward the first direct cosmological constraints from Void x CMB lensing.

astro-ph.CO