SearcharxivSearch

arXiv subjects

Hongxin Zhang

Publications and source records attributed to Hongxin Zhang.

At least 19 recordsLinked to original sources

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic system that emulates how humans traditionally create 4D worlds, yet automates the entire process. We present GS-Agent, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language. Inspired by how humans build 4D worlds, GS-Agent decomposes the task into entity management, covering 3D asset curation, material tuning, placement, and motion control, and rendering configuration, including camera and lighting manipulation. Multiple agents with distinct expertise interact with the physics engine via code, seek multimodal feedback, and collaborate to iteratively construct 4D worlds that align with the given descriptions. Experimental results show that GS-Agent effectively converts natural language into diverse and physically plausible 4D worlds exhibiting rich interactions among liquids, deformable objects, and rigid bodies, while achieving cinematic camera and lighting control. We envision GS-Agent as a foundation for a new paradigm in 4D world generation, empowering creative content creation and physical AI. Project page at https://umass-embodied-agi.github.io/gs-agent/

cs.RO

The Next Generation Virgo Cluster Survey (NGVS). II. A Catalog of Galaxies in the Virgo Cluster

The Next Generation Virgo Cluster Survey (NGVS) is a deep, high resolution imaging campaign that used the 1 deg$^2$ MegaCam instrument on the Canada-France-Hawaii Telescope to carry out a comprehensive optical survey of the Virgo cluster, from its core to its virial radius. The NGVS covers a contiguous area of 104 deg$^2$ (8.63 Mpc$^2$ at the 16.5 Mpc distance of Virgo) in the $u^*$-,$g$-,$i$-, and $z$-band, with additional limited coverage in $r$. In this paper, we present the final catalog of Virgo galaxies across the entire NGVS area. The catalog includes 3680 galaxies considered to be $bona~fide$ members of the cluster, spanning a factor of 2.5 million in luminosity, from $g = 8.42$ mag to $g = 24.41$ mag ($M_g = -22.67$ mag to $M_g = -6.68$ mag). With 2100 previously uncataloged galaxies, the NGVS catalog augments the number of known Virgo members by a factor 2.3. The catalog is complete down to $g = 18.6$ mag ($M_g=-12.5$ mag, corresponding to a stellar mass $M_* \sim 1.6\times10^7~M_{\odot}$ for an old stellar population) and 50% complete at $g = 22.0$ mag ($M_g=-9.1$ mag, $M_* \sim 6.2\times10^5~M_{\odot}$), three magnitudes deeper than the venerable Virgo Cluster Catalog (VCC), which for over 40 years has served as the reference standard for Virgo. Photometric and structural parameters are derived for all NGVS galaxies and presented in a series of tables, alongside nuclear and morphological classification, as well as stellar masses and, when available, radial velocities.

astro-ph.GA

CritLens: Visual Analytics for Criteria Discovery in Review-Based Decision Making

We present CritLens, a visual analytics system that helps users build personalized multi-criteria decision models from review text. In everyday decisions -- choosing equipment, hotels, or restaurants -- evaluation criteria are either preset by platforms or generated by LLMs, leaving users unable to discover, adjust, or verify them against the underlying evidence. This is problematic because many preferences are latent: they surface only upon encountering specific reviews, and any fixed framework risks overlooking low-frequency but decisive details. CritLens addresses this gap by using LLMs to transform reviews into an initial AHP decision model, then supporting iterative, human-in-the-loop refinement. Through coverage gap detection in the embedding space, users discover criteria missed by the initial model; through interactive weight adjustment under AHP consistency constraints, they express personal priorities; and through a multi-level scorecard and exportable decision report, they trace every ranking back to the original review text. Two case studies, an eight-participant user study, and a quantitative consistency-repair experiment demonstrate the system's effectiveness.

cs.HC

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across geometry, material, and assembly-based reasoning. We benchmark popular robot policies, pre-trained VLAs, and oracle-state planners. Our results reveal a significant performance gap: while pre-trained VLAs exhibit preliminary success on seed tasks after single-task fine-tuning, they struggle to perform on mutated tasks, implying their brittleness in manipulation tasks requiring reasoning, strategy adaptation, and robustness to deceptive or constrained environments. Project page is available at https://umass-embodied-agi.github.io/RoboWits.

cs.RO

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmental constraints across city-scale outdoor domains. We introduce Sentinel Challenge, a benchmark where multiple decentralized embodied agents must communicate in natural language to agree on a mutually safe and convenient meeting point within large, city-scale outdoor environments. Each agent must then navigate safely while avoiding dynamic sentinels patrolling the area, using a tool that provides coarse spatial information. To address this, we propose CoSaR (Cooperative Spatial Reasoning and Planning), a framework that bridges the high-level communication and planning abilities of foundation models with the precision of classical spatial navigation algorithms. CoSaR enables agents to exchange situational updates, reason over evolving spatial constraints, and collaboratively replan trajectories. Evaluated across 14 city-level scenes with 3-5 agents, CoSaR consistently leads to faster gathering, shorter path lengths, and improved safety. Our results demonstrate that integrating dynamic communication with spatial reasoning is essential for robust multi-agent cooperation. By formalizing this new setting and providing a scalable benchmark, we aim to build a foundation for advancing cooperative spatial intelligence in embodied multi-agent systems. Code and challenge are available at https://github.com/UMass-Embodied-AGI/Sentinel.

cs.CV

Discovery of low-redshift analogues to "Little Red Dots" in DESI: A later evolutionary stage of compact LRDs?

The James Webb Space Telescope (JWST) has recently discovered a population of compact, red sources at z > 4 known as "Little Red Dots" (LRDs). They are characterized by their V-shaped continuum spectra and prominent broad Balmer emission lines. As their underlying physical nature remains debated and direct study at high-redshift is challenging; therefore, we seek to identify and characterize LRD analogues in the low-redshift universe to constrain their properties and potential evolutionary pathways. We identified five candidates at z = 0.2-0.4 from the Dark Energy Spectroscopic Instrument (DESI) that exhibit spectral energy distributions (SEDs) and broad Balmer emission lines closely resembling their high-redshift counterparts. However, we find significant differences: our low-redshift sample occupies a different region on the Baldwin, Phillips \& Terlevich (BPT) diagram, and their stellar masses are significantly higher, suggesting a more substantial host galaxy contribution. These sources are not necessarily direct local analogues of high-redshift LRDs, but may represent later evolutionary stages of compact, rapidly accreting systems, or systems with related observational properties arising under different physical conditions. This sample provides a valuable laboratory for detailed follow-up studies to elucidate the nature of LRD-like phenomena.

astro-ph.GA

FatigueFormer: Static-Temporal Feature Fusion for Robust sEMG-Based Muscle Fatigue Recognition

We present FatigueFormer, a semi-end-to-end framework that deliberately combines saliency-guided feature separation with deep temporal modeling to learn interpretable and generalizable muscle fatigue dynamics from surface electromyography (sEMG). Unlike prior approaches that struggle to maintain robustness across varying Maximum Voluntary Contraction (MVC) levels due to signal variability and low SNR, FatigueFormer employs parallel Transformer-based sequence encoders to separately capture static and temporal feature dynamics, fusing their complementary representations to improve performance stability across low- and high-MVC conditions. Evaluated on a self-collected dataset spanning 30 participants across four MVC levels (20-80%), it achieves state-of-the-art accuracy and strong generalization under mild-fatigue conditions. Beyond performance, FatigueFormer enables attention-based visualization of fatigue dynamics, revealing how feature groups and time windows contribute differently across varying MVC levels, offering interpretable insight into fatigue progression.

cs.LG

ODD-SEC: Onboard Drone Detection with a Spinning Event Camera

The rapid proliferation of drones requires balancing innovation with regulation. To address security and privacy concerns, techniques for drone detection have attracted significant attention.Passive solutions, such as frame camera-based systems, offer versatility and energy efficiency under typical conditions but are fundamentally constrained by their operational principles in scenarios involving fast-moving targets or adverse illumination.Inspired by biological vision, event cameras asynchronously detect per-pixel brightness changes, offering high dynamic range and microsecond-level responsiveness that make them uniquely suited for drone detection in conditions beyond the reach of conventional frame-based cameras.However, the design of most existing event-based solutions assumes a static camera, greatly limiting their applicability to moving carriers--such as quadrupedal robots or unmanned ground vehicles--during field operations.In this paper, we introduce a real-time drone detection system designed for deployment on moving carriers. The system utilizes a spinning event-based camera, providing a 360{\deg} horizontal field of view and enabling bearing estimation of detected drones. A key contribution is a novel image-like event representation that operates without motion compensation, coupled with a lightweight neural network architecture for efficient spatiotemporal learning. Implemented on an onboard Jetson Orin NX, the system can operate in real time. Outdoor experimental results validate reliable detection with a mean angular error below 2{\deg} under challenging conditions, underscoring its suitability for real-world surveillance applications. We will open-source our complete pipeline to support future research.

cs.CV

Doc2AHP: Inferring Structured Multi-Criteria Decision Models via Semantic Trees with LLMs

While Large Language Models (LLMs) demonstrate remarkable proficiency in semantic understanding, they often struggle to ensure structural consistency and reasoning reliability in complex decision-making tasks that demand rigorous logic. Although classical decision theories, such as the Analytic Hierarchy Process (AHP), offer systematic rational frameworks, their construction relies heavily on labor-intensive domain expertise, creating an "expert bottleneck" that hinders scalability in general scenarios. To bridge the gap between the generalization capabilities of LLMs and the rigor of decision theory, we propose Doc2AHP, a novel structured inference framework guided by AHP principles. Eliminating the need for extensive annotated data or manual intervention, our approach leverages the structural principles of AHP as constraints to direct the LLM in a constrained search within the unstructured document space, thereby enforcing the logical entailment between parent and child nodes. Furthermore, we introduce a multi-agent weighting mechanism coupled with an adaptive consistency optimization strategy to ensure the numerical consistency of weight allocation. Empirical results demonstrate that Doc2AHP not only empowers non-expert users to construct high-quality decision models from scratch but also significantly outperforms direct generative baselines in both logical completeness and downstream task accuracy.

cs.AI

Updated Metallicity Diagnostics for Precision Oxygen Abundance Measurements in High-redshift Galaxies with JWST

Recent work has demonstrated that widely used strong-line oxygen abundance indicators, such as O3N2, $\rm R23$, and $\widehat{\rm R}$, suffer from large uncertainties when applied to high-redshift galaxies. We show that this loss of precision primarily arises because, at fixed \Oabund, galaxies span a wide dynamic range in ionization parameter and nitrogen enrichment. Here we develop updated indicators that explicitly incorporate both effects via the proxies O32 and N2O2. We define ${\rm R}_{\rm u}\equiv \rm R23+\alpha_1 O32+\alpha_2 N2O2$, $\widehat{\rm R}_{\rm u}\equiv \rm \widehat{R}+\beta_1 O32+\beta_2 N2O2$, and ${\rm O}_{\rm u}\equiv \rm O3N2+\gamma_1 O32+\gamma_2 N2O2$, and calibrate \Oabund~as low-order polynomials in each composite indicator. Applied to a JWST sample with $T_{\rm e}$-method abundances, the updated indicators substantially tighten the correlations with \Oabund, boosting adjusted coefficients of determination from $\mathbb{R}^2\lesssim 0$ (classical indicators) to $\mathbb{R}^2\gtrsim 0.5$ for the full sample and to $\sim 0.7$ at $z>2$. The residuals reveal a redshift evolution in the mapping between \Oabund, strong lines, ionization, and nitrogen enrichment, with a pivotal turning point near the cosmic noon ($z\sim 2$). Our calibrations provide a practical, physically grounded path to precise metallicity measurements in the JWST era and a firmer basis for quantifying early chemical enrichment and feedback.

astro-ph.GA

Radio AGN feedback sustains quiescence only in a minority of massive galaxies

Radio active galactic nuclei (AGNs) eject a huge amount of energy into the surrounding medium and are thought to potentially prevent gas cooling and maintain the quiescence of massive galaxies. The short-lived, sporadic, and anisotropic nature of radio activities, coupled with the detection of abundant cold gas around some massive quiescent galaxies, raise questions about the efficiency of radio feedback in massive galaxies. Here we present an innovative method rooted in artificial intelligence to separate galaxies in which radio feedback is effective (RFE), regardless of current radio emission, from those in which radio feedback is ineffective (RFI), according to their optical images. Galaxies categorized as RFE are all dynamically hot, whereas quiescent RFI (RFI-Q) galaxies usually have extended cold-disk components. At given stellar mass, dark matter halos hosting RFE galaxies are between four to ten times more massive than those of RFI-Q galaxies. We find, for the first time, that almost all RFE galaxies have scant cold gas, irrespective of AGN activity. In contrast, many RFI-Q galaxies are surrounded by substantial amounts of condensed atomic gas, indicating a different evolutionary path from RFE galaxies. Our finding provides direct and compelling evidence that a radio AGN has gone through about 300 on-off cycles and that radio feedback can prevent gas cooling over a timescale much longer than that of radio activity. Contrary to general belief, our analysis shows that only a small fraction of massive galaxies are influenced by strong radio AGNs, suggesting that current galaxy formation models need serious revision.

astro-ph.GA

Mass stratification in the globular cluster system revealing the assembly history of the nearest S0 galaxy NGC 3115

Galaxy formation and evolution is hierarchical. The most massive galaxies are thought to form their central regions early through violent dissipational processes, then grow inside-out by accreting smaller satellites. While widely supported, direct observational confirmation of this process in individual galaxies remains lacking, except for the Milky Way. We present a detailed analysis of globular cluster (GC) candidates within a $70^\prime$ ($\sim190$ kpc) radius around the nearest S0 galaxy, NGC 3115, using images in \textit{g,r,z} bands from the DESI Legacy Imaging Surveys and data from Gaia. We report the discovery of mass stratification in the GC system (GCS), evident in two ways: first, the effective radius of the GCS increases monotonically from the bright to faint end, up to the detection limit near the turnover magnitude of the GC luminosity function (GCLF); second, the GCLF shows fainter turnover magnitudes and smaller standard deviations at larger galactocentric radii. This stratification cannot be readily explained by radial migration or tidal dissolution, but most likely reflects the hierarchical assembly of NGC 3115's stellar halo, with later-accreted satellites deposited across broader galactocentric distances. This interpretation is supported by cosmological simulations of subhalos with comparable mass and bulge-to-total mass ratios and is consistent with the negative color gradients observed in the GCS. Additionally, we identify several substructures within the GCS, indicating ongoing assembly of NGC 3115. This work highlights the power of GCS as tracers of galaxy assembly and sets the stage for upcoming space-based wide-field imaging surveys to constrain the assembly of massive galaxies.

astro-ph.GA

MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM

Recent advancements in 3D Gaussian Splatting (3DGS) have made a significant impact on rendering and reconstruction techniques. Current research predominantly focuses on improving rendering performance and reconstruction quality using high-performance desktop GPUs, largely overlooking applications for embedded platforms like micro air vehicles (MAVs). These devices, with their limited computational resources and memory, often face a trade-off between system performance and reconstruction quality. In this paper, we improve existing methods in terms of GPU memory usage while enhancing rendering quality. Specifically, to address redundant 3D Gaussian primitives in SLAM, we propose merging them in voxel space based on geometric similarity. This reduces GPU memory usage without impacting system runtime performance. Furthermore, rendering quality is improved by initializing 3D Gaussian primitives via Patch-Grid (PG) point sampling, enabling more accurate modeling of the entire scene. Quantitative and qualitative evaluations on publicly available datasets demonstrate the effectiveness of our improvements.

cs.CV

Virtual Community: An Open World for Humans, Robots, and Society

The rapid progress in AI and Robotics may lead to a profound societal transformation, as humans and robots begin to coexist within shared communities, introducing both opportunities and challenges. To explore this future, we present Virtual Community-an open-world platform for humans, robots, and society-built on a universal physics engine and grounded in real-world 3D scenes. With Virtual Community, we aim to enable the study of embodied social intelligence at scale. To support these, Virtual Community features: 1) An open-source multi-agent physics simulator that supports robots, humans, and their interactions within a society; 2) A large-scale, real-world aligned community generation pipeline, including vast outdoor space, diverse indoor scenes, and a community of grounded agents with rich characters and appearances. Leveraging Virtual Community, we propose two novel challenges. The Community Planning Challenge evaluates multi-agent reasoning and planning ability in open-world settings, such as cooperating to help agents with daily activities and efficiently connecting other agents. The Community Robot Challenge requires multiple heterogeneous robots to collaborate in solving complex open-world tasks. We evaluate various baselines on these tasks and demonstrate the challenges in both high-level open-world task planning and low-level cooperation controls. We hope that Virtual Community will unlock further study of human-robot coexistence within open-world environments.

cs.CV

Ella: Embodied Social Agents with Lifelong Memory

We introduce Ella, an embodied social agent capable of lifelong learning within a community in a 3D open world, where agents accumulate experiences and acquire knowledge through everyday visual observations and social interactions. At the core of Ella's capabilities is a structured, long-term multimodal memory system that stores, updates, and retrieves information effectively. It consists of a name-centric semantic memory for organizing acquired knowledge and a spatiotemporal episodic memory for capturing multimodal experiences. By integrating this lifelong memory system with foundation models, Ella retrieves relevant information for decision-making, plans daily activities, builds social relationships, and evolves autonomously while coexisting with other intelligent beings in the open world. We conduct capability-oriented evaluations in a dynamic 3D open world where 15 agents engage in social activities for days and are assessed with a suite of unseen controlled evaluations. Experimental results show that Ella can influence, lead, and cooperate with other agents well to achieve goals, showcasing its ability to learn effectively through observation and social interaction. Our findings highlight the transformative potential of combining structured memory systems with foundation models for advancing embodied intelligence. More videos can be found at https://umass-embodied-agi.github.io/Ella/.

cs.CV

A Glimpse of Satellite Galaxies in the Milky Way with the 2.5-meter Wide Field Survey Telescope (WFST): Bootes III and Draco

We carry out deep imaging of the Milky Way satellite galaxies, Bootes III and Draco, with WFST as one pilot observing program to demonstrate the capability of WFST. Combining catalogs with PS1 DR2 and Gaia DR3, we derive proper motions for candidate member stars in these two satellite galaxies over a 12-year time baseline, yielding uncertainties of ~1.8 mas/yr at 21 mag and ~3.0 mas/yr at 22 mag in the r band. The proper motions derived from bright and faint stars are consistent, indicating no significant variation in proper motion across stellar luminosity as these galaxies undergo tidal interactions with the MW. Meanwhile, we suggest that Bootes III represents the bound remnant of the progenitor galaxy that gave rise to the Styx stream, as evidenced by its elongated density profile and overdensity in both spatial and kinematic space. This is the first paper to use WFST to measure the proper motions of faint stars in Milky Way satellite galaxies. More detailed analyses will be presented in forthcoming papers from the wide field survey (WFS) program.

astro-ph.GA

Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge

We introduce Constrained Human-AI Cooperation (CHAIC), an inclusive embodied social intelligence challenge designed to test social perception and cooperation in embodied agents. In CHAIC, the goal is for an embodied agent equipped with egocentric observations to assist a human who may be operating under physical constraints -- e.g., unable to reach high places or confined to a wheelchair -- in performing common household or outdoor tasks as efficiently as possible. To achieve this, a successful helper must: (1) infer the human's intents and constraints by following the human and observing their behaviors (social perception), and (2) make a cooperative plan tailored to the human partner to solve the task as quickly as possible, working together as a team (cooperative planning). To benchmark this challenge, we create four new agents with real physical constraints and eight long-horizon tasks featuring both indoor and outdoor scenes with various constraints, emergency events, and potential risks. We benchmark planning- and learning-based baselines on the challenge and introduce a new method that leverages large language models and behavior modeling. Empirical evaluations demonstrate the effectiveness of our benchmark in enabling systematic assessment of key aspects of machine social intelligence. Our benchmark and code are publicly available at https://github.com/UMass-Embodied-AGI/CHAIC.

cs.AI

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

Intelligent game creation represents a transformative advancement in game development, utilizing generative artificial intelligence to dynamically generate and enhance game content. Despite notable progress in generative models, the comprehensive synthesis of high-quality game assets, including both images and videos, remains a challenging frontier. To create high-fidelity game content that simultaneously aligns with player preferences and significantly boosts designer efficiency, we present Hunyuan-Game, an innovative project designed to revolutionize intelligent game production. Hunyuan-Game encompasses two primary branches: image generation and video generation. The image generation component is built upon a vast dataset comprising billions of game images, leading to the development of a group of customized image generation models tailored for game scenarios: (1) General Text-to-Image Generation. (2) Game Visual Effects Generation, involving text-to-effect and reference image-based game visual effect generation. (3) Transparent Image Generation for characters, scenes, and game visual effects. (4) Game Character Generation based on sketches, black-and-white images, and white models. The video generation component is built upon a comprehensive dataset of millions of game and anime videos, leading to the development of five core algorithmic models, each targeting critical pain points in game development and having robust adaptation to diverse game video scenarios: (1) Image-to-Video Generation. (2) 360 A/T Pose Avatar Video Synthesis. (3) Dynamic Illustration Generation. (4) Generative Video Super-Resolution. (5) Interactive Game Video Generation. These image and video generation models not only exhibit high-level aesthetic expression but also deeply integrate domain-specific knowledge, establishing a systematic understanding of diverse game and anime art styles.

cs.CV