SearcharxivSearch

arXiv subjects

Zhiyuan Song

Publications and source records attributed to Zhiyuan Song.

13 recordsLinked to original sources

NTIRE 2024 Challenge on Image Super-Resolution (x4): Methods and Results

This paper reviews the NTIRE 2024 challenge on image super-resolution ($\times$4), highlighting the solutions proposed and the outcomes obtained. The challenge involves generating corresponding high-resolution (HR) images, magnified by a factor of four, from low-resolution (LR) inputs using prior information. The LR images originate from bicubic downsampling degradation. The aim of the challenge is to obtain designs/solutions with the most advanced SR performance, with no constraints on computational resources (e.g., model size and FLOPs) or training data. The track of this challenge assesses performance with the PSNR metric on the DIV2K testing dataset. The competition attracted 199 registrants, with 20 teams submitting valid entries. This collective endeavour not only pushes the boundaries of performance in single-image SR but also offers a comprehensive overview of current trends in this field.

cs.CV

Learning When to Think While Listening in Large Audio-Language Models

Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: delaying reasoning until the speech endpoint can improve answer quality but moves deliberation into user-visible response delay, while answering too early risks committing before decisive evidence arrives. We introduce a learnable wait-think-answer control formulation for LALMs. Motivated by the incremental nature of human conversation, the controller decides under partial audio evidence when to wait, when to externalize a compact reasoning update, and when to answer. Using Qwen2.5-Omni-7B as the base model, we construct aligned wait-think-answer traces from spoken reasoning data, train the controller with supervised fine-tuning (SFT), and then apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The reward combines answer correctness, action validity, update timing, latency synchronization, reasoning quality, and chain consistency, optimizing the complete wait-think-answer trajectory and not the final answer alone. On a six-task synthetic spoken reasoning question answering (SRQA) benchmark, the six-reward DAPO controller improves the row-weighted accuracy from 67.6% to 70.3% while reducing post-endpoint final-think length by 14% under the same Qwen deployment harness. On a 186-item human-recorded Real Audio Bench, a transfer check beyond text-to-speech (TTS)-rendered speech, the controller family remains functional: SFT achieves the strongest accuracy, while the six-reward DAPO controller is the only learned variant whose final-think length falls below the base. These results suggest that a streaming model should learn when to make intermediate reasoning explicit during the audio stream.

cs.CL

Fast and Safe Trajectory Optimization for Mobile Manipulators With Neural Configuration Space Distance Field

Mobile manipulators promise agile, long-horizon behavior by coordinating base and arm motion, yet whole-body trajectory optimization in cluttered, confined spaces remains difficult due to high-dimensional nonconvexity and the need for fast, accurate collision reasoning. Configuration Space Distance Fields (CDF) enable fixed-base manipulators to model collisions directly in configuration space via smooth, implicit distances. This representation holds strong potential to bypass the nonlinear configuration-to-workspace mapping while preserving accurate whole-body geometry and providing optimization-friendly collision costs. Yet, extending this capability to mobile manipulators is hindered by unbounded workspaces and tighter base-arm coupling. We lift this promise to mobile manipulation with Generalized Configuration Space Distance Fields (GCDF), extending CDF to robots with both translational and rotational joints in unbounded workspaces with tighter base-arm coupling. We prove that GCDF preserves Euclidean-like local distance structure and accurately encodes whole-body geometry in configuration space, and develop a data generation and training pipeline that yields continuous neural GCDFs with accurate values and gradients, supporting efficient GPU-batched queries. Building on this representation, we develop a high-performance sequential convex optimization framework centered on GCDF-based collision reasoning. The solver scales to large numbers of implicit constraints through (i) online specification of neural constraints, (ii) sparsity-aware active-set detection with parallel batched evaluation across thousands of constraints, and (iii) incremental constraint management for rapid replanning under scene changes.

cs.RO

Online Trajectory Optimization for Arbitrary-Shaped Mobile Robots via Polynomial Separating Hypersurfaces

An emerging class of trajectory optimization methods enforces collision avoidance by jointly optimizing the robot's configuration and a separating hyperplane. However, as linear separators only apply to convex sets, these methods require convex approximations of both the robot and obstacles, which becomes an overly conservative assumption in cluttered and narrow environments. In this work, we unequivocally remove this limitation by introducing nonlinear separating hypersurfaces parameterized by polynomial functions. We first generalize the classical separating hyperplane theorem and prove that any two disjoint bounded closed sets in Euclidean space can be separated by a polynomial hypersurface, serving as the theoretical foundation for nonlinear separation of arbitrary geometries. Building on this result, we formulate a nonlinear programming (NLP) problem that jointly optimizes the robot's trajectory and the coefficients of the separating polynomials, enabling geometry-aware collision avoidance without conservative convex simplifications. The optimization remains efficiently solvable using standard NLP solvers. Simulation and real-world experiments with nonconvex robots demonstrate that our method achieves smooth, collision-free, and agile maneuvers in environments where convex-approximation baselines fail.

cs.RO

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation robustly is challenging due to ambiguities inherent in egocentric videos and biases present in training data. Consequently, existing methods often struggle, learning spurious correlations from skewed object-action pairings in datasets and fundamental visual confounding factors of the egocentric perspective, such as rapid motion and frequent occlusions. To address these limitations, we introduce Causal Ego-REferring Segmentation (CERES), a plug-in causal framework that adapts strong, pre-trained RVOS backbones to the egocentric domain. CERES implements dual-modal causal intervention: applying backdoor adjustment principles to counteract language representation biases learned from dataset statistics, and leveraging front-door adjustment concepts to address visual confounding by intelligently integrating semantic visual features with geometric depth information guided by causal principles, creating representations more robust to egocentric distortions. Extensive experiments demonstrate that CERES achieves state-of-the-art performance on Ego-RVOS benchmarks, highlighting the potential of applying causal reasoning to build more reliable models for broader egocentric video understanding.

cs.CV

Visualization and manipulation of four-leaf clover-shaped electronic state in cuprate

High-Tc superconductivity in cuprates arises from carrier doping of an antiferromagnetic Mott insulator. Associated with these changes are spectral-weight transfers from the high-energy to low-energy, giving rise to a variety of intriguing electronic phenomena. In this study, for the first time, we discovered a 2a0 sized four-leaf clover-shaped (FLC) electronic state at low-energy, accompanied with the emergence of a characteristic "kink" around 16meV. With increasing doping, the number of FLC pattern decreases and ultimately vanishes in the overdoped region. Remarkably, we achieved real-time electric-field manipulation of this FLC state, through innovative in-situ scanning tunneling microscopy probe. This novel FLC state may not only redefine our understanding of precursor states of pairing, but also reveals its crucial role as a tunable electronic phase in high-Tc superconductors.

cond-mat.supr-con

Local Reactive Control for Mobile Manipulators with Whole-Body Safety in Complex Environments

Mobile manipulators typically encounter significant challenges in navigating narrow, cluttered environments due to their high-dimensional state spaces and complex kinematics. While reactive methods excel in dynamic settings, they struggle to efficiently incorporate complex, coupled constraints across the entire state space. In this work, we present a novel local reactive controller that reformulates the time-domain single-step problem into a multi-step optimization problem in the spatial domain, leveraging the propagation of a serial kinematic chain. This transformation facilitates the formulation of customized, decoupled link-specific constraints, which is further solved efficiently with augmented Lagrangian differential dynamic programming (AL-DDP). Our approach naturally absorbs spatial kinematic propagation in the forward pass and processes all link-specific constraints simultaneously during the backward pass, enhancing both constraint management and computational efficiency. Notably, in this framework, we formulate collision avoidance constraints for each link using accurate geometric models with extracted free regions, and this improves the maneuverability of the mobile manipulator in narrow, cluttered spaces. Experimental results showcase significant improvements in safety, efficiency, and task completion rates. These findings underscore the robustness of the proposed method, particularly in narrow, cluttered environments where conventional approaches could falter. The open-source project can be found at https://github.com/Chunx1nZHENG/MM-with-Whole-Body-Safety-Release.git.

cs.RO

Ly$α$ Halo Properties and Dust in the Circumgalactic Medium of $z \sim 2$ Star-forming Galaxies

We present Keck Cosmic Web Imager IFU observations around extended Ly$α$ halos of 27 typical star-forming galaxies with redshifts $2.0 < z < 3.2$ drawn from the MOSFIRE Deep Evolution Field survey. We examine the average Ly$α$ surface-brightness profiles in bins of star-formation rate (SFR), stellar mass ($M_*$), age, stellar continuum reddening, SFR surface density ($\rm Σ_{SFR}$), and $\rm Σ_{SFR}$ normalized by stellar mass ($\rm Σ_{sSFR}$). The scale lengths of the halos correlate with stellar mass, age, and stellar continuum reddening; and anti-correlate with star-formation rate, $\rm Σ_{SFR}$, and $\rm Σ_{sSFR}$. These results are consistent with a scenario in which the down-the-barrel fraction of Ly$α$ emission is modulated by the low-column-density channels in the ISM, and that the neutral gas covering fraction is related to the physical properties of the galaxies. Specifically, we find that this covering fraction increases with stellar mass, age, and $E(B-V)$; and decreases with SFR, $\rm Σ_{SFR}$ and $\rm Σ_{sSFR}$. We also find that the resonantly scattered Ly$α$ emission suffers greater attenuation than the (non-resonant) stellar continuum emission, and that the difference in attenuation increases with stellar mass, age, and stellar continuum reddening, and decreases with $\rm Σ_{sSFR}$. These results imply that more reddened galaxies have more dust in their CGM.

astro-ph.GA

The MOSDEF Survey: Properties of Warm Ionised Outflows at $z=$ 1.4-3.8

We use the large spectroscopic data set of the MOSFIRE Deep Evolution Field survey to investigate the kinematics and energetics of ionised gas outflows. Using a sample of 598 star-forming galaxies at redshift 1.4 < $z$ < 3.8, we decompose $\rm{H}α$ and [OIII] emission lines into narrow and broad components, finding significant detections of broad components in 10% of the sample. The ionised outflow velocity from individual galaxies appears independent of galaxy properties, such as stellar mass, star-formation rate (SFR), and star-formation-rate surface density ($Σ_{\rm SFR}$). Adopting a simple outflow model, we estimate the mass-, energy- and momentum-loading factors of the ionised outflows, finding modest values with averages of 0.33, 0.04, and 0.22, respectively. The larger momentum- than energy-loading factors, for the adopted physical parameters, imply that these ionised outflows are primarily momentum-driven. We further find a marginal correlation (2.5$σ$) between the mass-loading factor and stellar mass in agreement with predictions by simulations, scaling as $η_{m}$ $\propto M_{\star}^{-0.45}$. This shallow scaling relation is consistent with these ionised outflows being driven by a combination of mechanical energy generated by supernovae explosions and radiation pressure acting on dusty material. In a majority of galaxies, the outflowing material does not appear to have sufficient velocity to escape the gravitational potential of their host, likely recycling back at later times. Together, these results suggest that the ionised outflows traced by nebular emission lines are negligible, with the bulk of mass and energy carried out in other gaseous phases.

astro-ph.GA

LoG-CAN: local-global Class-aware Network for semantic segmentation of remote sensing images

Remote sensing images are known of having complex backgrounds, high intra-class variance and large variation of scales, which bring challenge to semantic segmentation. We present LoG-CAN, a multi-scale semantic segmentation network with a global class-aware (GCA) module and local class-aware (LCA) modules to remote sensing images. Specifically, the GCA module captures the global representations of class-wise context modeling to circumvent background interference; the LCA modules generate local class representations as intermediate aware elements, indirectly associating pixels with global class representations to reduce variance within a class; and a multi-scale architecture with GCA and LCA modules yields effective segmentation of objects at different scales via cascaded refinement and fusion of features. Through the evaluation on the ISPRS Vaihingen dataset and the ISPRS Potsdam dataset, experimental results indicate that LoG-CAN outperforms the state-of-the-art methods for general semantic segmentation, while significantly reducing network parameters and computation. Code is available at~\href{https://github.com/xwmaxwma/rssegmentation}{https://github.com/xwmaxwma/rssegmentation}.

cs.CV

Adaptive Zeroing-Type Neural Dynamics for Solving Quadratic Minimization and Applied to Target Tracking

The time-varying quadratic miniaturization (TVQM) problem, as a hotspot currently, urgently demands a more reliable and faster--solving model. To this end, a novel adaptive coefficient constructs framework is presented and realized to improve the performance of the solution model, leading to the adaptive zeroing-type neural dynamics (AZTND) model. Then the AZTND model is applied to solve the TVQM problem. The adaptive coefficients can adjust the step size of the model online so that the solution model converges faster. At the same time, the integration term develops to enhance the robustness of the model in a perturbed environment. Experiments demonstrate that the proposed model shows faster convergence and more reliable robustness than existing approaches. Finally, the AZTND model is applied in a target tracking scheme, proving the practicality of our proposed model.

math.OC

GRATIS: Deep Learning Graph Representation with Task-specific Topology and Multi-dimensional Edge Features

Graph is powerful for representing various types of real-world data. The topology (edges' presence) and edges' features of a graph decides the message passing mechanism among vertices within the graph. While most existing approaches only manually define a single-value edge to describe the connectivity or strength of association between a pair of vertices, task-specific and crucial relationship cues may be disregarded by such manually defined topology and single-value edge features. In this paper, we propose the first general graph representation learning framework (called GRATIS) which can generate a strong graph representation with a task-specific topology and task-specific multi-dimensional edge features from any arbitrary input. To learn each edge's presence and multi-dimensional feature, our framework takes both of the corresponding vertices pair and their global contextual information into consideration, enabling the generated graph representation to have a globally optimal message passing mechanism for different down-stream tasks. The principled investigation results achieved for various graph analysis tasks on 11 graph and non-graph datasets show that our GRATIS can not only largely enhance pre-defined graphs but also learns a strong graph representation for non-graph data, with clear performance improvements on all tasks. In particular, the learned topology and multi-dimensional edge features provide complementary task-related cues for graph analysis tasks. Our framework is effective, robust and flexible, and is a plug-and-play module that can be combined with different backbones and Graph Neural Networks (GNNs) to generate a task-specific graph representation from various graph and non-graph data. Our code is made publicly available at https://github.com/SSYSteve/Learning-Graph-Representation-with-Task-specific-Topology-and-Multi-dimensional-Edge-Features.

cs.LG

The Most Predictive Physical Properties for the Stellar Population Radial Profiles of Nearby Galaxies

We present a study on the radial profiles of D4000,luminosity-weighted stellar ages $τ_L$,and luminosity-weighted stellar metallicities $[Z/H]_L$ of 3654 nearby galaxies($0.01<z<0.15$)using the IFU spectroscopic data from the MaNGA survey available in the SDSS DR15,in an effort to explore the connection between median stellar population radial gradients($\nabla$D4000,$\nablaτ_L,\nabla[Z/H]_L$)out to~$1.5R_e$ and various galaxy properties,including stellar mass($M_\star$),specific star formation rate(sSFR),morphologies,and local environment. We find that $M_\star$ is the single most predictive physical property for$\nabla$D4000 and$\nabla[Z/H]_L$. The most predictive properties for $\nablaτ_L$ are sSFR,and to a lesser degree,$M_\star$. The environmental parameters,including local galaxy overdensities and central-satellite division,have virtually no correlation with stellar population radial profiles for the whole sample,but the $\nabla$D4000 of star-forming satellite galaxies with$M_\star\lesssim 10^{10}M_\odot$exhibit a significant positive correlation with galaxy overdensities. Galaxies with lower sSFR have on average steeper negative stellar population gradients,and this sSFR dependence is stronger for more massive star-forming galaxies. The negative correlation between the median stellar population gradients and$M_\star$ are best described largely as segmented relationships, whereby median gradients of galaxies with$\log M_\star\lesssim 10$(with the exact value depending on sSFR)have much weaker mass dependence than galaxies with higher$M_\star$. While the dependence of the radial gradients of ages and metallicities on T-Types and central stellar mass surface densities are generally not significant,galaxies with later T-Types or lower central mass densities tend to have significantly lower D4000,younger$τ_L$ and lower$[Z/H]_L$ across the radial ranges probed in this study.

astro-ph.GA