SearcharxivSearch

arXiv subjects

Yuxuan Wu

Publications and source records attributed to Yuxuan Wu.

At least 19 recordsLinked to original sources

WM-VS: Progress-Aligned World Models for Closed-Loop Visual Servoing

Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We call this gap the prediction-control mismatch and introduce WM-VS, a target-centric progress-aligned world-model framework for closed-loop visual servoing. Offline target-region DINOv2 correspondences define a signed four-dimensional servo coordinate for translation, scale, and in-plane rotation. Stage 1 aligns action-conditioned latent transitions with this coordinate; Stage 2 freezes the world model and trains a reactive joint-velocity policy with action imitation, consequence supervision, and short imagined rollouts that favor error contraction. Deployment is RGB-only and reactive, without online trajectory optimization. On a real 7-DoF eye-to-hand system, WM-VS reaches a corner RMSE no larger than 10 percent of its initial value in 30/30 trials and retains this criterion at the final valid frame in 25/30 (83.33 percent). Removing future-error alignment reduces retention to 26.67 percent. The learned progress signal agrees with an external AprilTag corner error not used for training or control (mean Spearman rho = 0.8778). Without retraining, two unseen 3D targets achieve translation-error reductions of 86.48 percent and 90.27 percent and rotation-error reductions of 70.01 percent and 65.70 percent. These results link progress-aligned action consequences to repeated closed-loop correction and transfer. Code and data will be released as open source.

cs.RO

UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given background image. The variety of real-world applications makes it highly challenging to develop a single model capable of unifying the diverse range of input-constrained generation sub-tasks, such as those conditioned by element types, sizes, or their relationships. Current methods either address only a subset of these tasks or necessitate separate model parameters for different conditions, failing to offer a truly unified solution. In this paper, we propose UniLayDiff: a Unified Diffusion Transformer, that for the first time, addresses various content-aware layout generation tasks with a single, end-to-end trainable model. Specifically, we treat layout constraints as a distinct modality and employ Multi-Modal Diffusion Transformer framework to capture the complex interplay between the background image, layout elements, and diverse constraints. Moreover, we integrate relation constraints through fine-tuning the model with LoRA after pretraining the model on other tasks. Such a schema not only achieves unified conditional generation but also enhances overall layout quality. Extensive experiments demonstrate that UniLayDiff achieves state-of-the-art performance across from unconditional to various conditional generation tasks and, to the best of our knowledge, is the first model to unify the full range of content-aware layout generation tasks.

cs.CV

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and weak task guidance, followed by supervised fine-tuning. Across VIG-48 and BCI Competition IV-2a/IV-2b, Stage 1 improves mean accuracy in 41 of 42 matched backbone-dataset-protocol comparisons, including all 12 leave-one-subject-out settings, with a maximum gain of 16.22 percentage points. On the 48-region cross-day VIG-48 task, EEG-VID achieves 6.52% Top-1 and 30.50% Top-5 accuracy. In a separate six-participant offline robot-scene study, candidate-constrained target selection reaches 40.24% versus a 25% chance level after subject-specific calibration. These results support task-guided latent prediction as a transferable pretraining strategy for EEG decoding and scene-constrained assistive target selection.

cs.LG

AdoDAS: A Privacy-Preserving Multimodal Challenge for Adolescent Depression, Anxiety, and Stress Assessment

Adolescent depression, anxiety, and stress (D/A/S) call for scalable tools that complement, rather than replace, professional evaluation. Under a privacy-preserving policy, the AdoDAS Grand Challenge withholds minors' raw recordings and distributes anonymized audio-visual representations and ASR-derived text. Its 6,000 participants provide 24,000 segments across one scripted-reading and three open-response sessions. Two tracks assess multi-task binary D/A/S screening and ordinal prediction of 21 DASS-21 item responses. From 191 registrations, the final leaderboards included 95 eligible screening teams and 64 item-prediction teams. Audio-visual baselines achieved 0.4604 mean F1 and 0.2675 mean Quadratic Weighted Kappa; leading submissions reached 0.5921 and 0.2776. Representative systems emphasize cross-session modelling, temporal multimodal fusion, psychometric structure, and task-aware calibration.

cs.MM

CogEvol: Towards Efficient and Reliable Learning Environment Generation

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.

cs.CL

Extracting the pairing gap from van Hove singularities in rf spectra of the Fermi Hubbard model

We show that van Hove singularities in rf spectra of the 3D attractive Fermi Hubbard model provide a robust route to extracting the pairing gap. Four types of singularities are classified, and their spectral positions are shown to depend solely on the pairing gap $Δ$ and chemical potential $μ$ through simple algebraic relations. Measuring two well-resolved singularities therefore determines both parameters without requiring full spectral fitting. Numerical simulations incorporating phenomenological lifetime and scattering broadenings confirm that these features remain visible in both momentum-integrated and $k_z$-integrated spectra, and become more pronounced at stronger coupling where conventional back-bending methods lose sensitivity. At half filling, particle-hole symmetry fixes $μ$, reducing the extraction to a single singularity measurement. These results establish vHS analysis as a practical spectroscopic diagnostic for pairing in quantum-simulated 3D Fermi Hubbard systems.

cond-mat.quant-gas

Quantum Geometry, Anomalous Scaling, and Strong Pseudogap Superfluidity in a Flat-Band Lieb Lattice

Flat-band systems such as magic-angle twisted bilayer graphene host strong-correlation superconductivity at vanishingly weak coupling, yet how quantum geometry and pairing fluctuations conspire to drive this phenomenon remains an open question. We investigate finite-temperature superfluidity in a quasi-two-dimensional Lieb lattice using a pairing fluctuation theory with a band-uniform attractive interaction $g<0$ that isolates the intrinsic quantum geometric contributions. Quantum geometry significantly amplifies superfluidity; the geometric pair hopping integral surpasses its conventional counterpart, and the geometric superfluid density becomes the dominant in-plane transport component. When the Fermi level enters the flat band, the BCS paradigm breaks down entirely; the pairing gap and $T_\text{c}$ shift from exponential to anomalous power-law scaling $Δ, T_\text{c} \propto |g|^ν$ ($ν>1$), and the superfluid density inherits an unconventional power-law temperature dependence at low temperatures. In the 2D limit ($t_z=0$), the pseudogap at $T_\text{c}$ nearly saturates the zero-temperature gap even at $|g|/t=0.001$, placing the system in a strong-pseudogap regime that would otherwise require unitary or BEC-scale interactions. These findings establish a microscopic mechanism for flat-band enhanced superfluidity and offer testable predictions for ultracold atom experiments.

cond-mat.quant-gas

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.

cs.SD

Computing Tools for Translation-Invariant Total Orders

We introduce TITO_Explore, a software package for representing and computing with Translation-Invariant Total Orders (TITOs). We define a canonical window notation for TITOs and design and implement algorithms for several computational tasks involving them. The package normalizes the window notation of a given TITO into its canonical form, computes its inversion set, compares the weak order between two TITOs, and computes the join of two specified TITOs. Our weak order comparison algorithm operates by partitioning the inversion sets into disjoint subsets, thereby breaking down the comparison problem into evaluations of paired subsets. The join algorithm uses an edge-weighted directed graph to represent inversions and converts the problem of finding the join into a weighted path problem in the graph.

math.CO

Eulerian-spanning set and coboundary operator: An investigation of maxcut beyond planar graphs

Using the concepts of Eulerian-spanning set and coboundary operator, we generalize Hadlock's conversion of the maxcut problem on planar graphs to one on general graphs with non-negative weights. Using our conversion, we can explore algorithms for maxcut beyond the class of planar graphs. We obtain a Fixed-Parameter Tractable algorithm for $k$-contraction apex graphs. Specifically, our algorithm can be applied to graphs with crossing number $k$, giving an $O(2^k(n+k)^{3/2}\log (n+k))$-time algorithm that matches the best known results when restricted to non-negative weights.

cs.DS

Continually Evolving Skill Knowledge in Vision Language Action Model

Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adaptation. Existing continual imitation learning (CIL) methods often rely on additional parameters or external modules, limiting scalability for large VLA models. We propose Stellar VLA, a knowledge-driven CIL framework without increasing network parameters. Two progressively extended variants are designed: T-Stellar for flat task-centric modeling and TS-Stellar for hierarchical task-skill structure. Stellar VLA enables self-evolving knowledge learning by jointly optimizing task representations and a learned knowledge space. We propose a knowledge-guided expert routing mechanism conditioned on knowledge relation and Top-K semantic embeddings, enabling task specialization without increasing model size. Experiments on the LIBERO benchmark show that Stellar VLAs achieve strong performance among both VLA and CIL baselines, using only 1 % data replay. Real-world evaluation on a dual-arm platform with distinct embodiment and scene configurations validates effective knowledge transfer. TS-Stellar excels in hierarchical manipulation, and visualizations reveal robust knowledge retention and task discovery. Project Website: https://stellarvla.github.io/

cs.RO

Resolving the bias-precision paradox with stochastic causal representation learning for personalized medicine

Estimating individualized treatment effects from longitudinal observational data is central to data-driven medicine, yet existing methods face a fundamental limitation: reducing confounding bias often suppresses clinically informative heterogeneity, degrading patient-specific predictions. Here, we identify this tension as a bias-precision paradox in causal representation learning and introduce sampling-based maximum mean discrepancy (sMMD), a stochastic alignment strategy that replaces global adversarial balancing with subset-level matching. We instantiate this approach in a framework for counterfactual outcome prediction with attribution-grounded interpretability. Across two large-scale ICU cohorts (n = 27,783), our framework improves accuracy under distribution shift, reducing error by up to 11.5% and substantially increasing recall in high-risk tasks. Mechanistic analyses show that sMMD selectively preserves clinically decisive variables. In human-AI evaluation, our method outperforms clinicians-in-training and large language models, and improves clinician accuracy by 14.7% while reducing decision time, enabling interpretable, real-time clinical decision support.

cs.AI

Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation

Retrieval-augmented generation (RAG) has become a widely recognized paradigm to combine parametric memory with non-parametric memories. An RAG model consists of two serial connecting components (retriever and generator). A major challenge in end-to-end optimization of the RAG model is that marginalization over relevant passages (modeled as discrete latent variables) from a knowledge base is required. Traditional top-K marginalization and variational RAG (VRAG) suffer from biased or high-variance gradient estimates. In this paper, we propose and develop joint stochastic approximation (JSA) based end-to-end training of RAG, which is referred to as JSA-RAG. The JSA algorithm is a stochastic extension of the EM (expectation-maximization) algorithm and is particularly powerful in estimating discrete latent variable models. Extensive experiments are conducted on five datasets for two tasks (open-domain question answering, knowledge-grounded dialogs) and show that JSA-RAG significantly outperforms both vanilla RAG and VRAG. Further analysis shows the efficacy of JSA-RAG from the perspectives of generation, retrieval, and low-variance gradient estimate.

cs.CL

Spectral study of the pseudogap in unitary Fermi gases

The existence of a pseudogap in unitary Fermi gases has recently been established and measured experimentally [Li et al., Nature 626, 288 (2024)]. This lends strong support for the pairing origin as the mechanism of the pseudogap in Fermi superfluids. Here we present a spectral study of unitary Fermi gases, and show how the data can be understood quantitatively, when compared with theoretically calculated momentum-resolved rf or microwave spectra, and the pseudogap extracted from the spectra. We use an iterative treatment of the fermion self energy and hence the spectral function, beyond previous pseudogap approximation, based on a pairing fluctuation theory that incorporates both particle-particle and particle-hole T matrices, with self-consistent self energy feedback. Our results not only provide a microscopic explanation of the experimental data but also strengthen the support for both the pairing-induced pseudogap physics and the pairing fluctuation theory of Fermi superfluidity.

cond-mat.quant-gas

Rf spectra and pseudogap in ultracold Fermi gases across the BCS-BEC crossover from pairing fluctuation theory

The pseudogap phenomenon is a hallmark of strongly interacting Fermi systems, from high-temperature superconductors to ultracold atomic gases, yet its precise origin remains debated. Here we calculate the spectral function and rf spectra of ultracold atomic gases across the BCS-BEC crossover to quantitatively investigate the pairing mechanism of the pseudogap. We advance our pairing fluctuation theory by incorporating particle-hole fluctuations, which renormalize the effective interaction in the particle-particle channel. To achieve quantitative accuracy, we employ a full numerical convolution for the pair susceptibility and self-energy, moving beyond previous analytic pseudogap approximations. This convolution approach automatically captures two critical effects: (i) the full spectral broadening of fermions due to finite pair lifetime, and (ii) the previously neglected pair-hole scattering effect, which manifests as a substantial Hartree energy. We calculate the spectral function, and use rf spectral intensity maps and energy distribution curves to determine the quasiparticle dispersion. From these, we extract the pseudogap $Δ$, Hartree energy, and chemical potential, mapping their evolution across the crossover. Our results show that the pseudogap emerges continuously as the system moves from the BCS regime toward BEC. Furthermore, the pair spectral function reveals that pairs become diffusive at energies above 2$Δ$, indicating that the pair lifetime is governed by virtual binding and unbinding processes. Our calculations achieve quantitative agreement with recent experiments across the BCS-BEC crossover, including at unitarity, providing strong support for a pairing-based origin of the pseudogap as described by our pairing fluctuation theory.

cond-mat.quant-gas

A census of quiescent galaxies across $0.5 < z < 8$ with JWST/MIRI: Mass-dependent number density evolution of quiescent galaxies in the early Universe

Recent JWST observations have revealed a large population of quiescent galaxies (QGs) at high redshift ($z \sim 4-8$), challenging current models of early galaxy formation and quenching. Accurate number density estimates are crucial but remain uncertain. We present a systematic study of QGs at $0.5 < z < 8$ using a mass-complete sample from the JWST/PRIMER survey with deep NIRCam and MIRI imaging. We demonstrate that MIRI photometry is essential for refining the QG sample: it helps to mitigate contamination from dusty star-forming galaxies in the high-mass regime at $z \sim 3-5$ and aids in recovering lower-mass QG candidates at $z > 5$ that are often missed without including MIRI data. We find that the evolution of the QG number density is strongly mass-dependent. The density of massive QGs ($\log (M_{\star}/M_{\odot}) > 10.6$) declines rapidly, falling from $n \approx 1.32\times10^{-5}~~\mathrm{Mpc^{-3}}$ at $z \sim 3-4$ to $n < 1 \times10^{-6}~~\mathrm{Mpc^{-3}}$ at $z \sim 6$, and becomes negligible at $z > 6$. In contrast, low-mass QGs ($9.5 < \log (M_{\star}/M_{\odot}) < 10.6$) exhibit a remarkably constant number density of $n \sim 3\times10^{-6}~\mathrm{Mpc^{-3}}$ across the redshift range $z = 4-8$. This plateau suggests that these high-redshift, low-mass QGs may be galaxies undergoing temporary quenching episodes, likely subject to rejuvenation upon future gas accretion. Comparisons with leading galaxy formation models reveal significant tensions: most models underestimate the abundance of massive QGs at $z > 4$ and fail to reproduce the flat density evolution observed for the low-mass population.

astro-ph.GA

TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning

Recent Vision-Language-Action models show potential to generalize across embodiments but struggle to quickly align with a new robot's action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages abundant wheeled-humanoid data to boost bipedal VLA. Our key idea is to use end-effector trajectories as a morphology-agnostic interface. TrajBooster (i) extracts 6D dual-arm end-effector trajectories from real-world wheeled humanoids, (ii) retargets them in simulation to Unitree G1 with a whole-body controller trained via a heuristic-enhanced harmonized online DAgger to lift low-dimensional trajectory references into feasible high-dimensional whole-body actions, and (iii) forms heterogeneous triplets that couple source vision/language with target humanoid-compatible actions to post-pre-train a VLA, followed by only 10 minutes of teleoperation data collection on the target humanoid domain. Deployed on Unitree G1, our policy achieves beyond-tabletop household tasks, enabling squatting, cross-height manipulation, and coordinated whole-body motion with markedly improved robustness and generalization. Results show that TrajBooster allows existing wheeled-humanoid data to efficiently strengthen bipedal humanoid VLA performance, reducing reliance on costly same-embodiment data while enhancing action space understanding and zero-shot skill transfer capabilities. For more details, For more details, please refer to our \href{https://jiachengliu3.github.io/TrajBooster/}.

cs.RO

Effects of particle-hole fluctuations on the superfluid transition in two-dimensional atomic Fermi gases

Proper treatment of the many-body interactions is of paramount importance in our understanding of strongly correlated systems. Here we investigate the effects of particle-hole fluctuations on the Berezinskii-Kosterlitz-Thouless (BKT) transition in two-dimensional Fermi gases throughout the entire BCS-BEC crossover. We include self-consistently in the self energy treatment the entire particle-hole $T$ matrix, which constitutes a renormalization of the bare interaction that appears in the particle-particle scattering $T$ matrix, leading to a screening of the pairing interaction and hence a dramatic reduction of the pairing gap and the transition temperature. The BKT transition temperature $T_\text{BKT}$ is determined by the critical phase space density, for which the pair density and pair mass are determined using a pairing fluctuation theory, which accommodates self-consistently the important self-energy feedback in the treatment of finite-momentum pairing fluctuations. The screening strength varies continuously from its maximum in the BCS limit to essentially zero in BEC limit. In the unitary regime, it leads to an interaction-dependent shift of $T_\text{BKT}$ towards the BEC regime. This shift is crucial in an attempt to explain experimental data quantitatively, which often depends on the interaction strength. Our findings are consistent with available experimental results in the unitary and BEC regimes and with quantum Monte Carlo simulations in the BCS and unitary regimes.

cond-mat.quant-gas