SearcharxivSearch

arXiv subjects

Hao Zhang

Publications and source records attributed to Hao Zhang.

At least 19 recordsLinked to original sources

Full Inseparability and Genuine Multipartite Entanglement Coincide for Finite-Mode Gaussian States

For general mixed states, entanglement across every bipartition need not imply genuine multipartite entanglement (GME), because a biseparable decomposition may switch the separable cut from term to term. We prove that this convex ambiguity disappears for Gaussian states of finitely many bosonic modes. More generally, for any finite family of partitions, a Gaussian density operator in the trace-norm-closed convex class generated by states separable across those partitions is already separable across one fixed partition in the family. Only the target is Gaussian; a valid decomposition may be continuous and may contain arbitrary non-Gaussian states. Thus full inseparability and GME coincide, Gaussian k-separability and k-producibility reduce to fixed-partition tests, and party-wise tensor powers cannot activate GME from a biseparable Gaussian state. The proof combines a spectral selector with a holomorphic rigidity argument that converts one product vector in the square-root range of a Gaussian state into a block-local covariance certificate. The result shows that partition mixing, a generic mixed-state mechanism, adds no new exact finite-mode Gaussian states.

quant-ph

PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images

Breast cancer survival prediction in computational pathology presents a remarkable challenge due to tumor heterogeneity. For instance, different regions of the same tumor in the pathology image can show distinct morphological and molecular characteristics. This makes it difficult to extract representative features from whole slide images (WSIs) that truly reflect the tumor's aggressive potential and likely survival outcomes. In this paper, we present PathoHR, a novel pipeline for accurate breast cancer survival prediction that enhances any size of pathological images to enable more effective feature learning. Our approach entails (1) the incorporation of a plug-and-play high-resolution Vision Transformer (ViT) to enhance patch-wise WSI representation, enabling more detailed and comprehensive feature extraction, (2) the systematic evaluation of multiple advanced similarity metrics for comparing WSI-extracted features, optimizing the representation learning process to better capture tumor characteristics, (3) the demonstration that smaller image patches enhanced follow the proposed pipeline can achieve equivalent or superior prediction accuracy compared to raw larger patches, while significantly reducing computational overhead. Experimental findings valid that PathoHR provides the potential way of integrating enhanced image resolution with optimized feature learning to advance computational pathology, offering a promising direction for more accurate and efficient breast cancer survival prediction. Code will be available at https://github.com/AIGeeksGroup/PathoHR.

eess.IV

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.

cs.AI

A Better Spur Should Start From Each Objective

Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing traditional linear scalarization to experience severe metric oscillations. To address optimization conflicts among multiple objectives in real-world deployment scenarios, we propose Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization. Specifically, MMPO performs exposure debiasing to mitigate sparse and biased rewards, applies priority-aware orthogonal projection to decouple conflicting gradients, and introduces self-prompted gradient constraints to prevent dominant objectives from overwhelming weaker ones. Experiments on real-world e-commerce datasets show that MMPO improves training stability and consistently achieves better performance across conflicting metrics. Moreover, it generalizes robustly to broader tasks such as ToolRL and code generation, demonstrating its effectiveness as a practical paradigm for multi-objective alignment.

cs.AI

SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation

Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements, and in practice, often fall short. Symmetry is one such requirement: violations, even subtle ones, on symmetry can render a model physically unusable. We present SymTRELLIS, a method that enforces arbitrary finite point group symmetries (rotational, reflectional, and polyhedral) during the flow-based 3D generation of TRELLIS.2, without retraining the underlying VAE or flow model. Our key idea is to approximate the latent-space action of spatial transformations as a learned linear operator on voxel latents, implemented as a lightweight spatial-transform latent mapper trained on generic, non-symmetric 3D data. At generation time, we enforce symmetry by averaging predicted flow velocities across all symmetry-equivalent transformations at each ODE step, a process we call velocity symmetrization. The symmetry specification can be estimated automatically from an initial TRELLIS.2 generation or supplied by the user, enabling deliberate fold manipulation beyond what the input image suggests. On a curated benchmark of 266 strictly symmetric objects spanning 2- to 20-fold rotations and polyhedral symmetry groups, SymTRELLIS substantially reduces all symmetry error metrics compared to TRELLIS.2, Hunyuan3D-2.1, and TripoSG, while maintaining reconstruction accuracy comparable to the base model.

cs.GR

TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models

Large audio-language models (LALMs) can hallucinate audio objects, answering "yes" to absent sound events, thus undermining reliability in audio question answering. We propose Token-Adaptive Decoding (TAD), a training-free strategy for hallucination mitigation that grounds the initial yes/no decision by contrasting logits under real audio with a matched silent reference. TAD introduces a token-adaptive, confidence-guided gate that is decision-critical at the first decoding step and class-conditional on affirmative tokens, using the audio-silent margin to avoid overcorrection when evidence is weak or already sufficient. Experiments on AudioCaps-Hallucination show that, relative to Audio-Aware Decoding (AAD), a contrastive baseline with fixed contrast strength, TAD improves F1 for Qwen2 by 0.059 to 0.117 across Popular, Adversarial, and Random splits, and for Gemma by 0.025 to 0.064, while on Clotho-AQA it raises F1 from 0.810 to 0.816 on Qwen2 and remains comparable to AAD on Gemma.

cs.SD

PTIR-GS: Path-Traced Inverse Rendering with Global Illumination in 3D Gaussian Fields

Ray tracing enables 3D Gaussian fields to serve as a representation for physically based light transport. Faithful inverse rendering requires forward rendering and backward optimization to be defined within a consistent light-transport pipeline. Existing Gaussian inverse-rendering methods typically rely on splatting-derived G-buffers and screen-space optimization. The local affine approximation of perspective projection can introduce inconsistencies with path tracing, while simplified rendering formulations often neglect or approximate indirect illumination and visibility. Therefore, we propose a splatting-free path-traced inverse-rendering framework for 3D Gaussian fields that unifies forward rendering and backward optimization within the same recursive multi-bounce light-transport pipeline. We formulate light transport over overlapping Gaussian primitives directly in path space, enabling Monte Carlo path tracing and pathwise gradient replay over the same recursively sampled transport paths. The framework jointly optimizes materials and environment light under the full rendering equation, with ray-traced visibility and global illumination explicitly evaluated. Extensive experiments demonstrate competitive material estimation and improved path-traced rendering quality, producing more plausible shadows, reflections, and relighting under global illumination.

cs.GR

X2-N: A Transformable Wheel-legged Humanoid Robot with Dual-mode Locomotion and Manipulation

Wheel-legged robots combine the efficiency of wheeled locomotion with the versatility of legged systems, enabling rapid traversal over both continuous and discrete terrains. However, conventional designs typically employ fixed wheels as feet and limited degrees of freedom (DoFs) at the hips, resulting in reduced stability and mobility during legged locomotion compared to humanoids with flat feet. In addition, most existing platforms lack a full upper body with arms, which limits their ability to perform dexterous manipulation tasks. In this letter, we present X2-N, a high-DoF transformable robot with dual-mode locomotion and manipulation. X2-N can operate in both humanoid and wheel-legged forms and transform seamlessly between them through joint reconfiguration. We further propose a reinforcement learning (RL)-based whole-body control framework tailored to this morphology, enabling control across hybrid locomotion, transformation, and manipulation. We validate X2-N in a range of challenging locomotion and manipulation tasks, including dynamic skating-like motion, stair climbing, and package delivery. Results demonstrate high

cs.RO

RW-TTT: Batched Serving for Request-Owned Test-Time Training State

Test-time training (TTT) adapts an LLM during generation by reading and updating request-owned state, such as fast weights, low-rank deltas, or streaming learner state. This breaks batched LLM serving, which assumes shared static weights: serial execution is correct but slow, while naive batching can corrupt request state. We formulate this problem as read-write TTT serving and present RW-TTT , which tags each decode step with its owner, version, and READ/WRITE effect, batches only compatible phases, and commits updates only to the owner. On one GPU with eight fast-weight InPlace-TTT streams, RW-TTT reaches 274.61 aggregate tok/s, 9.31x over sequential serving and 3.44x over per-stream replicas under the same memory budget. It preserves behavior on RULER, a long-context benchmark, and passes owner/version checks.

cs.LG

MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding

Matrix optimizers such as Muon are attractive for large-scale training because they can improve convergence and token efficiency over coordinate-wise optimizers. Muon does this by orthogonalizing momentum-smoothed matrix updates with Newton-Schulz, producing spectrum-balanced updates that require the complete 2D matrix as input. This exposes a systems mismatch: FSDP/ZeRO-3 saves memory by making the optimizer see shards, not whole matrices. Existing systems therefore either reconstruct matrices at every optimizer step, paying weight-sized communication after backward, or make the update local by using ZeRO-1 owner placement with full parameters resident. MatrixFSDP takes a third path: it changes where ZeRO-3 shards live, not the optimizer being computed. For each 2D weight, one data-parallel rank owns the whole matrix and the other ranks hold empty shards; non-matrix tensors are packed into tail owners and stay on AdamW. The ordinary backward reduction then lands the full Muon input on the owner, so Newton-Schulz runs locally with no optimizer-step matrix collective. Forward and backward still materialize and reshard parameters; the runtime challenge is to make that uneven layout efficient and correct. MatrixFSDP does so with MatrixShard metadata, a balance-aware owner planner, deterministic owner-segment P2P collectives, owner-buffer pinning, and owner-shard checkpoint resharding. The resulting update matches full-matrix Muon while preserving ZeRO-3-scale memory: on 64 A100s, MatrixFSDP reduces optimizer-step latency over stock FSDP2-Muon by 4.2x on one node and 54.6x on eight nodes, reaches up to 2.15x end-to-end speedup, and runs model sizes where ZeRO-1 owner placement exceeds an 80 GB GPU.

cs.DC

A New Sufficient Condition for Oriented Graphs Determined by Their Generalized Skew Spectra

Characterizing graphs uniquely determined by their spectra (DS) is a core open problem in spectral graph theory. While this problem has been extensively investigated for simple undirected graphs, it remains relatively underexplored for oriented graphs. For a simple undirected graph $G$ equipped with an orientation $σ$, the corresponding oriented graph $Σ=(G,σ)$ is the digraph obtained by orienting each edge of $G$ according to $σ$. An oriented graph $Σ$ is said to be \emph{determined by its generalized skew spectrum} (DGSS) if every oriented graph sharing the same generalized skew spectrum is isomorphic to $Σ$. This paper develops a new sufficient criterion for recognizing DGSS controllable oriented graphs, which applies to a much broader family of graphs than previously known results. Let $S$ be the skew-adjacency matrix of $Σ$, $W(Σ)=[e,Se,\ldots,S^{n-1}e]$, and $d_n$ the last invariant factor of $W(Σ)$. For each odd prime $p$, we define the polynomial $Φ_p(Σ;x)=\gcd(χ(S;x),χ(S+J;x))$ over the finite field $\mathbb{F}_p$, which is invariant under generalized skew cospectrality. By analyzing the square-free part of $Φ_p(Σ;x)$ and the associated $p$-main polynomial, we establish a DGSS sufficient condition under the square-free assumption on $d_n$. The proposed criterion allows higher $p$-nullity and recovers the square-free determinant criterion of Qiu, Wang and Wang~(2019) as a special case. We further provide illustrative examples to verify the wider applicability of our new condition and to highlight the role of the compatibility constraints on the irreducible factors of $Φ_p(Σ;x)$.

math.CO

The Endpoint Fractional Riesz Estimate on the Hamming Cube

Let $1<p<2$ and $D_j$ be the discrete partial derivative on the Hamming cube $Ω_n = \{ -1, 1\}^n$. Let $Δ=\sum_{j=1}^nD_j$ be the discrete Laplacian.We prove the endpoint inequality \[ \left\|\left(\sum_{j=1}^n|D_jf|^2\right)^{1/2}\right\|_{p} \lesssim_p\|Δ^{1/p}f\|_{p}, \quad \forall f:Ω_n \to \mathbb C. \] This result answers the conjecture proposed by Naor, Eskenazis and Ivanisvili (see [BenEfraimLustPiquard, IvanisviliVolberg] or [Remark 45, EskenazisIvanisvili]). The proof relies heavily on the noncommutative semigroup BMO theory [JungeMei].

math.FA

Twin-photon generation in a silicon nitride microresonator

Photonic chips with silicon nitride ($\mathrm{Si_3N_4}$) microring resonators are well established as heralded single-photon sources, but their operation as frequency-degenerate twin-photon sources has not previously been demonstrated. Here, we realise a twin-photon source at telecommunication wavelengths in a $\mathrm{Si_3N_4}$ ring microresonator via an inverse four-wave mixing (FWM) process, in which two photons from spectrally distinct pumps are converted into a pair of identical twin photons. The measurements show a maximum coincidence-to-accidental ratio (CAR) of $5.4\pm0.6$. In addition, the microresonator functions as a heralded single-photon source through pump-degenerate spontaneous four-wave mixing (SFWM), exhibiting a spectral purity of $P=0.67\pm0.05$ and a heralded anti-bunching of $g^{(2)}_h(0)=0.0042\pm0.0015$. Together, these results demonstrate both photon-generation schemes on a single integrated $\mathrm{Si_3N_4}$ platform, highlighting its potential for scalable, tailored quantum light generation.

physics.optics

Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction

Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanoparticles (NAPs) derived from Chinese herbal formulas and applications such as anti-lung-cancer therapy demonstrate the substantial potential of self-assembly for nanomedicine discovery. Yet discovery still relies on costly wet-lab screening, while existing machine learning approaches lack standardized tasks, effective pairwise compatibility modeling, and public benchmarks with unified evaluation. To address these limitations, we formalize NSA prediction as a binary classification task for predicting self-assembly between molecular pairs and then establish NSA-Bench, the first public benchmark with curated molecular combinations, experimental conditions, self-assembly labels, and standardized evaluation protocols. We further develop NSA-Net, an interaction-aware multimodal framework that integrates complementary molecular evidence from graph topology, sequence semantics, and physicochemical descriptors to learn molecular-pair representations for self-assembly prediction. Extensive experiments on NSA-Bench show that NSA-Net achieves a ROC-AUC of $0.9470\pm0.0112$ (Small) and $0.9492\pm0.0062$ (Large). On the Small track, it surpasses the strongest machine-learning and graph-based baselines by 3.9 and 17.1 percentage points, respectively. Representation analyses reveal interpretable molecular characteristics associated with self-assembly prediction captured by the learned representations. Moreover, an NSA-Agent case study further demonstrates how NSA-Net predictions can support formulation refinement through experimental-condition-aware reasoning. Our code is available at https://github.com/developer-hq/NSA-Net.

q-bio.QM

Error estimate of the nonuniform BDF3-L2 method for subdiffusion equations via multiscale solution decomposition

Numerical experiments reported by Quan and Wu [SIAM J Numer Anal 61 (2023) 2106-2132] show that the observed temporal convergence rates of nonuniform L2 methods for subdiffusion models are not consistent with the theoretically predicted order $3-α$. This discrepancy suggests that a more refined analysis is needed and motivates the development of a nonuniform BDF3-L2 method for the subdiffusion equation. To account for the initial solution singularity, we employ the multiscale solution decomposition to decompose the original solution and approximate a smoother unknown variable that satisfies the subdiffusion model with a smoother source term. The resulting formulation, however, involves restrictive high-order boundary conditions on the source term and initial data. To overcome this difficulty, we introduce a spectral truncation technique that requires only slightly stronger regularity of the data and a controllable truncation error. We establish high-order regularity estimates of the solution to the truncated problem and develop a nonuniform BDF3-L2 method for its numerical approximation, based on which we derive a rigorous error estimate of temporal convergence order $2+α$. Numerical experiments are carried out to substantiate the theoretical findings.

math.NA

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular patterns while suppressing weak yet binding-relevant signals, such as functional groups and residue-context patterns, limiting the modeling of multi-scale biochemical correspondences. To address this issue, we propose ProbeMatchDTI, a pattern-probe-driven framework comprising IterProbe and BindingProbe. IterProbe explicitly retains contextual states across refinement depths and uses learnable probes to select them at each position before cross-entity matching, thereby preserving weak biochemical patterns and strengthening associations among functional groups, local motifs, and molecular scaffolds. BindingProbe then characterizes cross-entity drug-protein complementarity at local biochemical-unit and whole-pair levels, jointly modeling fine-grained interactions and multi-scale correspondences while preserving weaker binding-relevant associations. Extensive experiments demonstrate the superiority of ProbeMatchDTI, achieving 2.0% and 0.5% higher AUC-ROC on BindingDB and DrugBank, respectively. Feature-level pattern analyses further characterize its probe-driven behavior in cross-scale biochemical pattern matching. We further connect ProbeMatchDTI predictions with an evidence-guided downstream drug-discovery workflow, demonstrating their utility for candidate refinement and validation planning. Our code is available at https://github.com/developer-hq/ProbeMatchDTI

cs.LG

ByteX: A Unified AI Search Engine at ByteDance

Since 2016, ByteX has been the foundation of ByteDance's search infrastructure, scaling to more than 7,000 clusters and 300 PB of indexed data. Driven by the demands of AI workloads, ByteX has evolved from a text search engine into a unified AI search system supporting vector retrieval, lexical matching, and predicate filtering. Its largest deployment indexes nearly one trillion high-dimensional vectors. This scale exposes two central bottlenecks in AI-era retrieval: memory-intensive graph-index construction under sustained ingestion, and the prohibitive cost of keeping vector indexes entirely in memory. ByteX addresses these bottlenecks with two techniques. First, it introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors. Second, it provides a hybrid storage engine that supports memory-resident, hybrid, and SSD-resident deployments, with fine-grained record-level caching to trade memory for latency under operational control. On large-scale benchmarks, ByteX improves throughput by up to 3x, reduces indexing memory by 80%, and lowers operating cost by 86% compared with prior systems, while supporting trillion-vector scale, write-heavy or latency-sensitive workloads in production.

cs.DB

CAT-Flow: Curvature-Adaptive sTeps for Flow Matching

Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of its ODE-based sampling process creates a fundamental efficiency bottleneck: the quality of generated samples is highly sensitive to the choice of step-sizes, and current models typically require 20 to 30 steps for good quality. In this work, we propose two lightweight, training-free algorithms, CAT-OV and CAT-OT that adapt step-sizes at inference time based on a novel connection between Flow Matching sampling and gradient flow. Our algorithms are computed efficiently by not requiring additional neural function evaluations. Specifically, CAT-OT estimates curvature over time via a finite-difference approximation of the time-derivative of the vector field, while CAT-OV approximates curvature over the state space via a gradient of the vector field. Under suitable conditions, both methods have truncation error bounds of constant order. Empirically, CAT-OV and CAT-OT outperform existing step-size heuristics in image quality metrics across four text- to-image Flow Matching models, reducing the number of generation steps required to reach comparable quality by up to 40%.

cs.LG