SearcharxivSearch

arXiv subjects

Junhui Liu

Publications and source records attributed to Junhui Liu.

At least 19 recordsLinked to original sources

Position: Embodied AI Requires a Privacy-Utility Trade-off

Embodied AI (EAI) systems are rapidly transitioning from simulations into real-world domestic and other sensitive environments. However, recent EAI solutions have largely demonstrated advancements within isolated stages such as instruction, perception, planning and interaction, without considering their coupled privacy implications in high-frequency deployments where privacy leakage is often irreversible. This position paper argues that optimizing these components independently creates a systemic privacy crisis when deployed in sensitive settings, thereby advancing the position that privacy in EAI is a life cycle-level architectural constraint rather than a stage-local feature. To address these challenges, we propose Secure Privacy Integration in Next-generation Embodied AI (SPINE), a unified privacy-aware framework that treats privacy as a dynamic control signal governing cross-stage coupling throughout the entire EAI life cycle. SPINE decomposes the EAI pipeline into various stages and establishes a multi-criterion privacy classification matrix to orchestrate contextual sensitivity across stage boundaries. We conduct preliminary simulation and real-world case studies to conceptually validate how privacy constraints propagate downstream to reshape system behavior, illustrating the insufficiency of fragmented privacy patches and motivating future research directions into secure yet functional embodied AI systems. We detail the SPINE framework and case studies at https://github.com/rminshen03/EAI_Privacy_Position.

cs.AI

BridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off

Virtual try-off (VTOFF) aims to recover canonical flat-garment representations from images of dressed persons for standardized display and downstream virtual try-on. Prior methods often treat VTOFF as direct image translation driven by local masks or text-only prompts, overlooking the gap between on-body appearances and flat layouts. This gap frequently leads to inconsistent completion in unobserved regions and unstable garment structure. We propose BridgeDiff, a diffusion-based framework that explicitly bridges human-centric observations and flat-garment synthesis through two complementary components. First, the Garment Condition Bridge Module (GCBM) builds a garment-cue representation that captures global appearance and semantic identity, enabling robust inference of continuous details under partial visibility. Second, the Flat Structure Constraint Module (FSCM) injects explicit flat-garment structural priors via Flat-Constraint Attention (FC-Attention) at selected denoising stages, improving structural stability beyond text-only conditioning. Extensive experiments on standard VTOFF benchmarks show that BridgeDiff achieves state-of-the-art performance, producing higher-quality flat-garment reconstructions while preserving fine-grained appearance and structural integrity.

cs.CV

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation

Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a different source language. However, this task remains challenging due to the scarcity of parallel multilingual emotional corpora, the presence of foreign accent artifacts, and the difficulty of separating emotion from language-specific prosodic features. In this paper, we propose XEmoRAG, a novel framework to enable zero-shot emotion transfer from Chinese to Thai using a large language model (LLM)-based model, without relying on parallel emotional data. XEmoRAG extracts language-agnostic emotional embeddings from Chinese speech and retrieves emotionally matched Thai utterances from a curated emotional database, enabling controllable emotion transfer without explicit emotion labels. Additionally, a flow-matching alignment module minimizes pitch and duration mismatches, ensuring natural prosody. It also blends Chinese timbre into the Thai synthesis, enhancing rhythmic accuracy and emotional expression, while preserving speaker characteristics and emotional consistency. Experimental results show that XEmoRAG synthesizes expressive and natural Thai speech using only Chinese reference audio, without requiring explicit emotion labels. These results highlight XEmoRAG's capability to achieve flexible and low-resource emotional transfer across languages. Our demo is available at https://tlzuo-lesley.github.io/Demo-page/ .

eess.AS

Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

Despite remarkable achievements, automatic speech recognition (ASR) in low-resource scenarios still faces two challenges: high-quality data scarcity and high computational demands. This paper proposes EThai-ASR, the first to apply large language models (LLMs) to Thai ASR and create an efficient LLM-based ASR system. EThai-ASR comprises a speech encoder, a connection module and a Thai LLM decoder. To address the data scarcity and obtain a powerful speech encoder, EThai-ASR introduces a self-evolving data refinement strategy to refine weak labels, yielding an enhanced speech encoder. Moreover, we propose a pluggable sequence compression module used in the connection module with three modes designed to reduce the sequence length, thus decreasing computational demands while maintaining decent performance. Extensive experiments demonstrate that EThai-ASR has achieved state-of-the-art accuracy in multiple datasets. We release our refined text transcripts to promote further research.

cs.SD

UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching pre-trained model (UFM) designed to address feature matching challenges across a wide spectrum of modal images. We present Multimodal Image Assistant (MIA) transformers, finely tunable structures adept at handling diverse feature matching problems. UFM exhibits versatility in addressing both feature matching tasks within the same modal and those across different modals. Additionally, we propose a data augmentation algorithm and a staged pre-training strategy to effectively tackle challenges arising from sparse data in specific modals and imbalanced modal datasets. Experimental results demonstrate that UFM excels in generalization and performance across various feature matching tasks. The code will be released at:https://github.com/LiaoYun0x0/UFM.

cs.CV

C3PO IV: co-natal stars depleted in refractories are magnetically more active -- possible imprints of planets

Chemical abundance anomalies in twin stars have recently been considered tell-tale signs of interactions between stars and planets. While such signals are prevalent, their nature remains a subject of debate. On one hand, exoplanet formation may induce chemical depletion in host stars by locking up refractory elements. On the other hand, exoplanet engulfment can result in chemical enrichment, both processes potentially producing similar differential signals. In this study, we aim to observationally disentangle these processes by using the Ca II infrared triplet to measure the magnetic activity of 125 co-moving star pairs with high SNR, high-resolution spectra from the Magellan, Keck, and VLT telescopes. We find that co-natal star pairs in which the two stars exhibit significant chemical abundance differences also show differences in their magnetic activity, with stars depleted in refractories being magnetically more active. Furthermore, the strength of this correlation between differential chemical abundances and differential magnetic activity increases with condensation temperature. One possible explanation is that the chemical anomaly signature may be linked to planet formation, wherein refractory elements are locked into planets, and the host stars become more active due to more efficient contraction during the pre-main-sequence phase or star-planet tidal and magnetic interactions.

astro-ph.EP

Applications of Large Models in Medicine

This paper explores the advancements and applications of large-scale models in the medical field, with a particular focus on Medical Large Models (MedLMs). These models, encompassing Large Language Models (LLMs), Vision Models, 3D Large Models, and Multimodal Models, are revolutionizing healthcare by enhancing disease prediction, diagnostic assistance, personalized treatment planning, and drug discovery. The integration of graph neural networks in medical knowledge graphs and drug discovery highlights the potential of Large Graph Models (LGMs) in understanding complex biomedical relationships. The study also emphasizes the transformative role of Vision-Language Models (VLMs) and 3D Large Models in medical image analysis, anatomical modeling, and prosthetic design. Despite the challenges, these technologies are setting new benchmarks in medical innovation, improving diagnostic accuracy, and paving the way for personalized healthcare solutions. This paper aims to provide a comprehensive overview of the current state and future directions of large models in medicine, underscoring their significance in advancing global health.

cs.AI

Double-lined Spectroscopic Binaries from the LAMOST Low-Resolution Survey

We report on a data-driven spectral model that we have developed for the identification of double-lined spectroscopic binary stars (SB2s) in the LAMOST low-resolution survey (R$\sim$1800). Employing simultaneous fitting with both single-star and binary-star models, we detected over 4800 SB2 candidates, where both components are detectably contributing to the spectrum, from an initial pool of 2.6 million objects. Tests show that our model favors FGK-type main-sequence binaries with high mass ratio ($q\geq$ 0.7) and large radial velocity separation ($\Delta \rm RV \geq$ 100~km$\,$s$^{-1}$). Almost all of these candidates are positioned above the main sequence in the color-magnitude diagram, indicating their binary nature. Additionally, we utilized various observational data, including spectroscopy, photometry, parallax, and extinction, to determine multiple physical parameters such as the effective temperature, age, metallicity, radial velocity, mass, mass ratio, stellar radius, along with their associated uncertainties for these SB2 candidates. For the 44 candidates with seven or more observational epochs, we provide complete orbital solutions. We make available catalogs containing various stellar parameters for identified SB2 systems.

astro-ph.SR

FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter

Vocoders reconstruct speech waveforms from acoustic features and play a pivotal role in modern TTS systems. Frequent-domain GAN vocoders like Vocos and APNet2 have recently seen rapid advancements, outperforming time-domain models in inference speed while achieving comparable audio quality. However, these frequency-domain vocoders suffer from large parameter sizes, thus introducing extra memory burden. Inspired by PriorGrad and SpecGrad, we employ pseudo-inverse to estimate the amplitude spectrum as the initialization roughly. This simple initialization significantly mitigates the parameter demand for vocoder. Based on APNet2 and our streamlined Amplitude prediction branch, we propose our FreeV, compared with its counterpart APNet2, our FreeV achieves 1.8 times inference speed improvement with nearly half parameters. Meanwhile, our FreeV outperforms APNet2 in resynthesis quality, marking a step forward in pursuing real-time, high-fidelity speech synthesis. Code and checkpoints is available at: https://github.com/BakerBunker/FreeV

cs.SD

The X-ray Emission Reveals the Coronal Activities of Semi-detached Binaries

X-ray emission is an important tracer of stellar magnetic activity. We carried out a systematic correlation analysis for the X-ray luminosity $\log L_{\rm X}$, bolometric luminosity $\log L_{\rm bol}$, and X-ray activity level $\log(L_{\textrm{X}}$/$L_{\textrm{bol}})$ versus the binary parameters including orbital period $P$, Rossby number $R_{\rm O}$, effective temperature $T_{\rm eff}$, metallicity [Fe/H] and the surface gravity $\log g$, and the stellar mass $M$ \& radius $R$, by assembling a large sample of semi-detached (EB-type) binaries with X-ray emission (EBXs). The fact that both $\log L_{\rm X}$ and $\log L_{\rm bol}$ change in accordance with $\log P$ indicates that X-ray emission originates from the convection zone, while $\log L_{\rm X}$ is proportional to the convection zone area. We found that EBXs with main-sequence components exhibit an upward and then a downward trend in both the $\log T_{\rm eff}$-$\log L_{\textrm{X}}$ and $M$-$\log L_{\textrm{X}}$ relations, which is different from the monotonically decreasing trend shown by EBXs containing sub-giant and giant components. The magnetic activity level is negatively correlated with $\log T_{\rm eff}$ and stellar mass. Based on the magnetic dynamo model, the variations in the size and thickness of the surface convection zones can explain the observed relations. EBXs with main-sequence components have similar $R_{\rm O}$-$\log(L_{\textrm{X}}/L_{\textrm{bol}})$ relationship to that of the binaries in the clusters as Praesepe and Hyade. We compared the X-ray radiation properties of EBXs with those of the X-ray-emitting contact binaries and found that EBXs have broader value ranges for $\log L_{\rm X}$ and $\log(L_{\textrm{X}}$/$L_{\textrm{bol}})$.

astro-ph.SR

Wide binaries with white dwarf or neutron star companions discovered from Gaia DR3 and LAMOST

Gaia DR3 mission has identified and provided about 440,000 binary systems with orbital solutions, offering a valuable resource for searching binaries including a compact component. By combining the Gaia DR3 data with radial velocities (RVs) from the LAMOST spectroscopic survey, we identify three wide binaries possibly containing a compact object. For two of these sources with a main-sequence companion, no obvious excess is observed in the blue/red band of the Gaia DR3 XP spectra, and the LAMOST medium-resolution spectra exhibit clear single-lined features. The absence of an additional component from spectral disentangling analysis further suggests the presence of compact objects within these systems. On the other hand, the visible star of the third source is a stripped giant star. In contrast to most binaries including stripped stars, no emission line is detected in the optical spectra. The unseen star could potentially be a massive white dwarf or neutron star, but the possibility of an F-type dwarf star scenario cannot be ruled out. An examination of about ten binaries containing white dwarfs or neutron stars using both kinematic and chemical methods suggest most of these systems are located in the thin disk of the Milky Way.

astro-ph.SR

Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of unnatural foreign accents and difficulty in modeling the shared emotional expressions of different languages. Building on the DelightfulTTS neural architecture, this paper addresses these challenges by introducing specifically-designed modules to model the language-specific prosody features and language-shared emotional expressions separately. Specifically, the language-specific speech prosody is learned by a non-autoregressive predictive coding (NPC) module to improve the naturalness of the synthetic cross-lingual speech. The shared emotional expression between different languages is extracted from a pre-trained self-supervised model HuBERT with strong generalization capabilities. We further use hierarchical emotion modeling to capture more comprehensive emotions across different languages. Experimental results demonstrate the proposed framework's effectiveness in synthesizing bi-lingual emotional speech for the monolingual target speaker without emotional training data.

cs.SD

TKwinFormer: Top k Window Attention in Vision Transformers for Feature Matching

Local feature matching remains a challenging task, primarily due to difficulties in matching sparse keypoints and low-texture regions. The key to solving this problem lies in effectively and accurately integrating global and local information. To achieve this goal, we introduce an innovative local feature matching method called TKwinFormer. Our approach employs a multi-stage matching strategy to optimize the efficiency of information interaction. Furthermore, we propose a novel attention mechanism called Top K Window Attention, which facilitates global information interaction through window tokens prior to patch-level matching, resulting in improved matching accuracy. Additionally, we design an attention block to enhance attention between channels. Experimental results demonstrate that TKwinFormer outperforms state-of-the-art methods on various benchmarks. Code is available at: https://github.com/LiaoYun0x0/TKwinFormer.

eess.IV

Preserving background sound in noise-robust voice conversion via multi-task learning

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare attention to VC with background sound. The critical problem for preserving background sound in VC is inevitable speech distortion by the neural separation model and the cascade mismatch between the source separation model and the VC model. In this paper, we propose an end-to-end framework via multi-task learning which sequentially cascades a source separation (SS) module, a bottleneck feature extraction module and a VC module. Specifically, the source separation task explicitly considers critical phase information and confines the distortion caused by the imperfect separation process. The source separation task, the typical VC task and the unified task shares a uniform reconstruction loss constrained by joint training to reduce the mismatch between the SS and VC modules. Experimental results demonstrate that our proposed framework significantly outperforms the baseline systems while achieving comparable quality and speaker similarity to the VC models trained with clean data.

eess.AS

X-ray emission of contact binary variables within 1 kpc

By assembling the largest sample to date of X-ray emitting EW-type binaries (EWXs), we carried out correlation analyses for the X-ray luminosity log$L_{\textrm{X}}$, and X-ray activity level log($L_{\textrm{X}}$/$L_{\textrm{bol}}$) versus the orbital period $P$ and effective temperature $T_{\rm eff}$. We find strong $P$-log$L_{\textrm{X}}$ and $P$-log($L_{\textrm{X}}$/$L_{\textrm{bol}}$) correlations for EWXs with $P$ < 0.44 days and we provide the linear parametrizations for these relations, on the basis of which the orbital period can be treated as a good predictor for log$L_{\textrm{X}}$ and log($L_{\textrm{X}}$/$L_{\textrm{bol}}$). The aforementioned binary stellar parameters are all correlated with log$L_{\textrm{X}}$, while only $T_{\rm eff}$ exhibits a strong correlation with log($L_{\textrm{X}}$/$L_{\textrm{bol}}$). Then, EWXs with higher temperature show lower X-ray activity level, which could indicate the thinning of the convective area related to the magnetic dynamo mechanism. The total X-ray luminosity of an EWX is essentially consistent with that of an X-ray saturated main sequence star with the same mass as its primary, which may imply that the primary star dominates the X-ray emission. The monotonically decreasing $P$-log($L_{\textrm{X}}$/$L_{\textrm{bol}}$) relation and the short orbital periods indicate that EWXs could all be in the X-ray saturated state, and they may inherit the changing trend of the saturated X-ray luminosities along with the mass shown by single stars. For EWXs, the orbital period, mass, and effective temperature increase in concordance. We demonstrate that the period $P=0.44$ days corresponds to the primary mass of $\sim1.1 \rm M_\odot$, beyond which the saturated X-ray luminosity of single stars will not continue to increase with mass. This explains the break in the positive $P$-log$L_{\textrm{X}}$ relation for EWXs with $P>0.44$ days.

astro-ph.HE

ClothFormer:Taming Video Virtual Try-on in All Module

The task of video virtual try-on aims to fit the target clothes to a person in the video with spatio-temporal consistency. Despite tremendous progress of image virtual try-on, they lead to inconsistency between frames when applied to videos. Limited work also explored the task of video-based virtual try-on but failed to produce visually pleasing and temporally coherent results. Moreover, there are two other key challenges: 1) how to generate accurate warping when occlusions appear in the clothing region; 2) how to generate clothes and non-target body parts (e.g. arms, neck) in harmony with the complicated background; To address them, we propose a novel video virtual try-on framework, ClothFormer, which successfully synthesizes realistic, harmonious, and spatio-temporal consistent results in complicated environment. In particular, ClothFormer involves three major modules. First, a two-stage anti-occlusion warping module that predicts an accurate dense flow mapping between the body regions and the clothing regions. Second, an appearance-flow tracking module utilizes ridge regression and optical flow correction to smooth the dense flow sequence and generate a temporally smooth warped clothing sequence. Third, a dual-stream transformer extracts and fuses clothing textures, person features, and environment information to generate realistic try-on videos. Through rigorous experiments, we demonstrate that our method highly surpasses the baselines in terms of synthesized video quality both qualitatively and quantitatively.

cs.CV

Migrating Face Swap to Mobile Devices: A lightweight Framework and A Supervised Training Solution

Existing face swap methods rely heavily on large-scale networks for adequate capacity to generate visually plausible results, which inhibits its applications on resource-constraint platforms. In this work, we propose MobileFSGAN, a novel lightweight GAN for face swap that can run on mobile devices with much fewer parameters while achieving competitive performance. A lightweight encoder-decoder structure is designed especially for image synthesis tasks, which is only 10.2MB and can run on mobile devices at a real-time speed. To tackle the unstability of training such a small network, we construct the FSTriplets dataset utilizing facial attribute editing techniques. FSTriplets provides source-target-result training triplets, yielding pixel-level labels thus for the first time making the training process supervised. We also designed multi-scale gradient losses for efficient back-propagation, resulting in faster and better convergence. Experimental results show that our model reaches comparable performance towards state-of-the-art methods, while significantly reducing the number of network parameters. Codes and the dataset have been released.

cs.CV

The Disk Veiling Effect of the Black Hole Low-Mass X-ray Binary A0620-00

The optical light curves of quiescent black hole low-mass X-ray binaries often exhibit significant non-ellipsoidal variabilities, showing the photospheric radiation of the companion star is veiled by other source of optical emission. Assessing this "veiling" effect is critical to the black hole mass measurement. Here in this work, we carry out a strictly simultaneous spectroscopic and photometric campaign on the prototype of black hole low-mass X-ray binary A0620-00. We find that for each observation epoch, the extra optical flux beyond a pure ellipsoidal modulation is positively correlated with the fraction of veiling emission, indicating the accretion disk contributes most of the non-ellipsoidal variations. Meanwhile, we also obtain a K2V spectral classification of the companion, as well as the measurements of the companion's rotational velocity $v \sin i = 83.8\pm1.9$ km s$^{-1}$ and the mass ratio between the companion and the black hole $q=0.063\pm0.004$.

astro-ph.HE