Searcharxiv⌕ Search

arXiv subjects

Zhongju Wang

Publications and source records attributed to Zhongju Wang.

7 recordsLinked to original sources

Social Structure Matters in 3D Human-Human Interaction Generation

Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interaction (HHI) remains non-trivial, as HHI requires modeling the underlying \textbf{social structure} that governs phase progression, actor roles, and inter-actor coordination. In this paper, we formulate HHI generation as a social structure modeling and grounding problem: the model must first infer how an interaction unfolds and how the two actors coordinate their roles, and then realize this structure as continuous, physically plausible, and partner-aware 3D motion. To study how such structure should be modeled, we first examine the capability boundary of large language models (LLMs) for HHI generation. Our analysis shows that LLMs can \textit{think} by recovering phase decompositions and partner-aware roles, but cannot directly \textit{move}, as they fail to generate dynamic, physically plausible, and interaction-aware motion. This motivates our planner-executor paradigm, \textbf{Think with LLM, Move with Motion Skill}. The LLM planner converts implicit interaction semantics into motion-aligned social supervision by decomposing interactions into phases, assigning partner-aware actor roles, and aligning them with motion sequence. The motion executor then grounds the planned social structure into coordinated two-person motion by adapting a pretrained solo motion model with LoRA, previous-phase self-conditioning, and ego-relative partner conditioning. Together, our Solo-to-Social framework bridges social organization and motion realization, producing 3D HHI with improved phase consistency, role alignment, and partner-aware coordination.

cs.CV↗

3DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking Avatars

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emotion, and exhibit lifelike spatial dynamics, collectively defining a broader objective of expressivity. However, achieving this remains challenging due to insufficient training data with limited subject identities, narrow audio representations, and restricted explicit controllability. In this paper, we propose 3DXTalker, an expressive 3D talking avatar through data-curated identity modeling, audio-rich representations, and spatial dynamics controllability. 3DXTalker enables scalable identity modeling via 2D-to-3D data curation pipeline and disentangled representations, alleviating data scarcity and improving identity generalization. Then, we introduce frame-wise amplitude and emotional cues beyond standard speech embeddings, ensuring superior lip synchronization and nuanced expression modulation. These cues are unified by a flow-matching-based transformer for coherent facial dynamics. Moreover, 3DXTalker also enables natural head-pose motion generation while supporting stylized control via prompt-based conditioning. Extensive experiments show that 3DXTalker integrates lip synchronization, emotional expression, and head-pose dynamics within a unified framework, achieves superior performance in 3D talking avatar generation.

cs.CV↗

Physics-Aware Inverse Design for Nanowire Single-Photon Avalanche Detectors via Deep Learning

Single-photon avalanche detectors (SPADs) have enabled various applications in emerging photonic quantum information technologies in recent years. However, despite many efforts to improve SPAD's performance, the design of SPADs remained largely an iterative and time-consuming process where a designer makes educated guesses of a device structure based on empirical reasoning and solves the semiconductor drift-diffusion model for it. In contrast, the inverse problem, i.e., directly inferring a structure needed to achieve desired performance, which is of ultimate interest to designers, remains an unsolved problem. We propose a novel physics-aware inverse design workflow for SPADs using a deep learning model and demonstrate it with an example of finding the key parameters of semiconductor nanowires constituting the unit cell of an SPAD, given target photon detection efficiency. Our inverse design workflow is not restricted to the case demonstrated and can be applied to design conventional planar structure-based SPADs, photodetectors, and solar cells.

physics.app-ph↗

BERT-based Chinese Text Classification for Emergency Domain with a Novel Loss Function

This paper proposes an automatic Chinese text categorization method for solving the emergency event report classification problem. Since bidirectional encoder representations from transformers (BERT) has achieved great success in natural language processing domain, it is employed to derive emergency text features in this study. To overcome the data imbalance problem in the distribution of emergency event categories, a novel loss function is proposed to improve the performance of the BERT-based model. Meanwhile, to avoid the impact of the extreme learning rate, the Adabound optimization algorithm that achieves a gradual smooth transition from Adam to SGD is employed to learn parameters of the model. To verify the feasibility and effectiveness of the proposed method, a Chinese emergency text dataset collected from the Internet is employed. Compared with benchmarking methods, the proposed method has achieved the best performance in terms of accuracy, weighted-precision, weighted-recall, and weighted-F1 values. Therefore, it is promising to employ the proposed method for real applications in smart emergency management systems.

cs.CL↗

Optimal Portfolio Design for Statistical Arbitrage in Finance

In this paper, the optimal mean-reverting portfolio (MRP) design problem is considered, which plays an important role for the statistical arbitrage (a.k.a. pairs trading) strategy in financial markets. The target of the optimal MRP design is to construct a portfolio from the underlying assets that can exhibit a satisfactory mean reversion property and a desirable variance property. A general problem formulation is proposed by considering these two targets and an investment leverage constraint. To solve this problem, a successive convex approximation method is used. The performance of the proposed model and algorithms are verified by numerical simulations.

q-fin.PM↗

Effective Low-Complexity Optimization Methods for Joint Phase Noise and Channel Estimation in OFDM

Phase noise correction is crucial to exploit full advantage of orthogonal frequency division multiplexing (OFDM) in modern high-data-rate communications. OFDM channel estimation with simultaneous phase noise compensation has therefore drawn much attention and stimulated continuing efforts. Existing methods, however, either have not taken into account the fundamental properties of phase noise or are only able to provide estimates of limited applicability owing to considerable computational complexity. In this paper, we have reformulated the joint estimation problem in the time domain as opposed to existing frequency-domain approaches, which enables us to develop much more efficient algorithms using the majorization-minimization technique. In addition, we propose a method based on dimensionality reduction and the Bayesian Information Criterion (BIC) that can adapt to various phase noise levels and accomplish much lower mean squared error than the benchmarks without incurring much additional computational cost. Several numerical examples with phase noise generated by free-running oscillators or phase-locked loops demonstrate that our proposed algorithms outperform existing methods with respect to both computational efficiency and mean squared error within a large range of signal-to-noise ratios.

cs.IT↗

Design of PAR-Constrained Sequences for MIMO Channel Estimation via Majorization-Minimization

PAR-constrained sequences are widely used in communication systems and radars due to various practical needs; specifically, sequences are required to be unimodular or of low peak-to-average power ratio (PAR). For unimodular sequence design, plenty of efforts have been devoted to obtaining good correlation properties. Regarding channel estimation, however, sequences of such properties do not necessarily help produce optimal estimates. Tailored unimodular sequences for the specific criterion concerned are desirable especially when the prior knowledge of the channel is taken into account as well. In this paper, we formulate the problem of optimal unimodular sequence design for minimum mean square error estimation of the channel impulse response and conditional mutual information maximization, respectively. Efficient algorithms based on the majorization-minimization framework are proposed for both problems with guaranteed convergence. As the unimodular constraint is a special case of the low PAR constraint, optimal sequences of low PAR are also considered. Numerical examples are provided to show the performance of the proposed training sequences, with the efficiency of the derived algorithms demonstrated.

math.OC↗