SearcharxivSearch

arXiv subjects

Hao Guo

Publications and source records attributed to Hao Guo.

At least 19 recordsLinked to original sources

Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection

Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challenging due to instance-dependent modality contributions and misleading semantic consistency, where surface-level alignment masks underlying contradictory intent. Existing methods often rely on fixed fusion strategies and treat sarcasm as generic cross-modal mismatch, limiting their ability to capture subtle sarcasm cues and instance-specific modality interactions. To address these challenges, we propose a novel MSD framework that integrates Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization (SaCR). Specifically, a bidirectional gated interaction module performs cross-modal feature filtering and adaptively calibrates textual and visual contributions at the instance level. A dynamic fusion gate further balances modality importance to generate more robust multimodal representations. Furthermore, SaCR is introduced as a label-aware contrastive regularization objective that encourages semantic consistency for non-sarcastic samples while suppressing misleading consistency in sarcastic cases. The proposed framework is trained end-to-end with a multi-objective learning strategy that jointly optimizes multimodal classification and auxiliary unimodal supervision. Extensive experiments on MMSD and MMSD2.0 demonstrate that the proposed method consistently outperforms strong baselines.

cs.CL

Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.

cs.CL

A Frequency-Space Terahertz Transceiver Chip for Multi-Agent Communications and Spatial Awareness

Future indoor embodied-intelligence systems require scalable hardware platforms that support both high-capacity multi-agent connectivity and mutual spatial awareness. The terahertz (THz) spectrum offers abundant bandwidth and inherent spatial selectivity for integrated sensing and communication (ISAC); however, conventional phased arrays and programmable metasurfaces rely on dense beamforming networks, element-level control, or external THz illumination, making scalable multibeam operation challenging. Here, we report a fully integrated 208-258GHz 65-nm CMOS THz transceiver chip that monolithically integrates broadband front ends with heterogeneous leaky-wave metasurface (HLM) apertures within a 1.5mm by 4.9mm area. The HLM generates strongly dispersive leaky modes, enabling 75 degree frequency-controlled beam scanning with only four meta-atoms. Co-design of frequency-domain and spatial-domain mixing achieves spectrally clean frequency-to-space mapping for spatial-frequency division multiple access (SFDMA) communication. The THz chip demonstrates multi-agent simultaneous transmission and reception, two-dimensional localization, and sensing-enhanced communication, providing a scalable hardware platform for future THz embodied-intelligence networks.

physics.app-ph

Freeform super-oscillatory optics for CMOS-integrated THz super-resolution imaging

The diffraction limit fundamentally constrains the spatial resolution of far-field imaging systems. While near-field techniques can circumvent this limit, their inherently short working distances (WD) severely restrict practical applications. Super-oscillatory lenses (SOLs) offer a far-field alternative; however, conventional SOLs are plagued by discrete operating wavelengths, low efficiencies (below 5%), and formidable trade-offs among numerical aperture, chromatic aberration, and depth of focus (DOF). Here, we introduce a nonlocal, nonlinear-curvature mechanism to design a freeform SOL that achieves ultrabroadband (0.3 to 1 THz), achromatic super-resolution focusing with an unprecedented efficiency of 44%. Operating at a 9 mm WD, the lens maintains a consistent sub-diffraction full-width at half-maximum (FWHM) of around 0.45 wavelength alongside an extended DOF of around 10 wavelengths. By integrating a compact 65-nm CMOS oscillator-radiator array, we establish an advanced imaging platform capable of resolving complex 2D and 3D sub-millimeter features (down to 0.15 mm). Readily scalable to the optical regime via two-photon lithography, this freeform SOL paradigm paves the way for next-generation, high-performance integrated photonics.

physics.optics

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.

cs.AI

An integrated super resolution THz 3D imaging system based on a linear nonlocal achromatic freeform Bessel beam lens and high power oscillator radiator array

High performance terahertz (THz) 3D imaging is critical for non-destructive evaluation. However, conventional architectures are fundamentally limited by severe chromatic aberrations, modest spatial resolution, restricted depths of focus (DOF), and the bulky nature of commercial transceivers. While metasurfaces offer a compact alternative, achieving broadband achromatic super-resolution with an extended DOF remains a formidable challenge. Here, we present a highly integrated 3D THz imaging platform that synergizes a 3D printed nonlocal freeform Bessel-beam lens with a high power, 65nm CMOS oscillator radiator array. Harnessing nonlocal interactions within the lens, we generate an achromatic super resolution Bessel beam (0.3 to 1 THz) with a subdiffraction full width at half maximum (FWHM) of 0.65{\lambda} and a robust 4.7-mm DOF. Crucially, the system overcomes conventional sidelobe limitations, enabling high-fidelity 2D imaging of intricate sub-millimeter targets (e.g., USAF 1951 charts and QR codes) alongside robust 3D volumetric imaging through highly scattering media, such as printed circuit boards. By converging standard CMOS technology with additive manufacturing, this work establishes a versatile, cost-effective paradigm for next-generation integrated THz photonics

physics.optics

Thermal Suppression of Dynamical Quantum Phase Transitions in Finite-Dimensional Systems A Quasi-Hermitian Framework

We investigate dynamical quantum phase transitions (DQPTs) in finite-dimensional systems prepared in thermal equilibrium states and subjected to a sudden quench. A mixed-state Loschmidt amplitude is constructed from first principles within a metric-stationary pseudo-Hermitian framework, providing a self-contained derivation of the finite-temperature quench dynamics. Applying this framework to an $N$-level model consisting of a two-level sector coupled to $N-2$ spectator states, we find that temperature controls the DQPTs through the redistribution of thermal weights among the eigenstates. This mechanism leads to a dimensionality-dependent threshold temperature that becomes finite when the Hilbert-space dimension reaches five, above which the Loschmidt amplitude loses all real zeros and the DQPTs are fully suppressed. The thermal suppression mechanism suggests a general principle for controlling dynamical criticality through thermal occupation, while the quasi-Hermitian framework provides the self-consistent foundation for its rigorous derivation.

quant-ph

Wilczek-Zee Realization of Uhlmann Parallel Transport

The Uhlmann phase extends geometric phases to mixed quantum states via a parallel-transport condition on purification amplitudes, yet its direct implementation under standard Hamiltonian dynamics is obstructed by the non-Hermitian nature of the purification. We establish that for any smooth one-dimensional closed loop of full-rank qubit density matrices, there exists a four-level Hermitian parent Hamiltonian whose doubly degenerate ground-state subspace carries a Wilczek--Zee connection exactly equal to the Uhlmann connection. Consequently, the Uhlmann holonomy is faithfully reproduced by adiabatic evolution in the enlarged system. We further prove that this auxiliary-field construction is obstructed in generic two-dimensional parameter spaces by a Frobenius integrability condition, which we derive explicitly. The one-dimensional Uhlmann phase is thus placed on the same footing as the non-Abelian Berry phase, offering a purely Hermitian, Hamiltonian-based route to simulating mixed-state geometric phases. Numerical integration of the adiabatic dynamics confirms the exact correspondence and validates the convergence to the Uhlmann holonomy in the large-gap limit.

quant-ph

Electrical-Circuit Simulation of the Uhlmann Phase

The Uhlmann phase extends the concept of geometric phases to mixed quantum states through a parallel-transport condition on purification amplitudes, but its experimental realization has so far required sophisticated quantum platforms with carefully engineered auxiliary degrees of freedom. In this work, we reformulate the Uhlmann parallel-transport condition as a linear matrix differential equation and vectorize it to obtain an effective dynamical generator. This generator can be directly mapped onto the admittance matrix of a classical RC circuit, thereby translating the Uhlmann dynamics into the evolution of circuit node voltages. We illustrate the mapping using the equatorial-loop model and, via a rotating-frame transformation followed by a real decomposition, derive a time-independent, real-valued dynamical system suitable for analog implementation. LTspice simulations of the resulting active RC network faithfully reproduce the Uhlmann geometric phase and its topological transition at the critical purity, demonstrating that classical electrical circuits offer a simple and accessible platform for probing mixed-state geometric phases.

quant-ph

Fusion-E2Pulse: A Multimodal Event-RGB Fusion Network for Non-contact Pulse Wave Reconstruction

Non-contact pulse wave reconstruction hinges on the precise recovery of waveform morphology, including the dicrotic notch. Conventional Red-Green-Blue (RGB)-based methods, which extract physiological signals from recorded facial videos, are constrained by the integral imaging mechanism of standard cameras, where the exposure process induces a smoothing effect that attenuates subtle vascular pulsation details. Conversely, neuromorphic event cameras, while offering exceptional sensitivity to intensity fluctuations, are inherently susceptible to noise and artifacts induced by minor motion. To exploit the synergy between frame-based integration and event-based differential sensing, we propose a novel multimodal network named Fusion-E2Pulse. This framework utilizes filtered RGB signals as structural priors to suppress motion artifacts, while leveraging the high-sensitivity of event streams to recover fine-grained morphological details. Experimental results demonstrate that Fusion-E2Pulse achieves state-of-the-art performance, effectively balancing noise suppression and morphological fidelity, achieving a mean absolute error of 0.78 bpm for heart rate estimation, a waveform correlation of 0.89, and a systolic phase duration error of 16.74 ms, validating its efficacy in reconstructing fine-grained pathological features.

cs.CV

SkillChain: Closing the Loop on Skill Evolution for Image-Based E-Commerce AI Assistants

Image-based AI assistants are now deployed at production scale on e-commerce platforms, where a single uploaded image can trigger fundamentally different user intents: product search, style recommendation, visual encyclopedia, or utility tool calls, each demanding its own response format, tool invocation, and domain knowledge. Without per-intent behavioral constraints, LLM-based systems conflate these heterogeneous modes and fall short of domain quality standards, while the breadth and dynamism of the intent space render manual engineering infeasible. To address this, we present SkillChain, which closes the production feedback loop on Skill evolution, automating the lifecycle of Skills through three stages: Skill Creator for bootstrapping from task specs and trajectories, Route Optimizer for routing alignment, and Body Refiner for iterative Skill Body refinement via dual-path LLM-Judge evaluation. Deployed on a production-scale e-commerce image assistant, SkillChain substantially improves aggregate response quality, with the strongest gains on structural compliance and content quality; a one-week online A/B experiment further confirms significant gains in user engagement, content consumption, and long-term retention.

cs.CL

From Uniform to Learned Graph Priors: Diffusion for Structure Discovery

Neural relational inference (NRI) methods discover interaction graphs from trajectories through variational reasoning on discrete potential edges. However, these methods typically rely on oversimplified, factorized graph priors. Such priors, typically nearing uniform distributions, treat edges as independent entities. This systemic misalignment does not match the real-world systems and yields diffuse and indecisive edge posteriors limiting the reliability of structural discovery. To address this, we propose \textit{Diff-prior}, a diffusion-parameterized adaptive prior used to calibrate latent graph distribution rather than generate graphs. Our core insight is to reframe prior integration as a learnable denoising-style calibration that organizes scattered, uncertain edge posteriors into a more reliable overall structure which can be trained by the diffusion model. Diff-prior learns an adaptive structure prior that performs structured calibration on the edge posteriors during inference, guiding it towards a distribution closer to the underlying structure. The diff-prior operates before structural sampling and acts as a denoising calibrator directly on the encoder edge distribution, which provides a generic training paradigm over structured variables. Experiments on standard benchmarks validated our framework, and the results indicate that Diff-prior improves the performance of structure inference and generates more decisive edge posteriors across multiple NRI-family architectures. The code is available on https://github.com/Hardy158118/Diffprior.

cs.LG

Torsion-induced gauge structure in curved quantum waveguides

We investigate the effective dynamics of a particle confined near a space curve. In the strict thin-layer reduction of the nondegenerate transverse ground state, torsion does not enter the local effective Hamiltonian, which contains only the curvature-induced scalar geometric potential. In contrast, for a thin guide with finite transverse width, a leading-order adiabatic projection onto the twofold-degenerate first-excited transverse band renders the rotation of the Frenet normal frame dynamically relevant and generates a matrix-valued Abelian gauge potential. Using a projection-based derivation in a co-rotating Frenet-frame basis, we show that this effective gauge potential is directly determined by the local torsion of the curve. The resulting effective Hamiltonian takes a gauge-covariant form and produces two transverse-mode branches whose parabolic dispersions are shifted in opposite directions in momentum space. For closed curves, the associated holonomy is controlled by the integrated torsion and leads to geometric interference. These results provide a direct realization of a Wilczek--Zee-type connection induced purely by spatial geometry in curved quantum waveguides. We further construct a classical-wave analogue using the degenerate bending modes of an isotropic elastic rod, demonstrating that the same torsion-induced gauge structure appears in continuum wave physics.

quant-ph

Defect Holonomy Near Rank-Deficient Mixed States

We investigate the geometry of mixed quantum states near rank-changing points, showing that these singularities function as effective geometric defects. The Uhlmann connection is well-defined on the full-rank sector of the density-matrix manifold, while rank-deficient states form singular boundary strata where the bundle structure degenerates. By restricting to a punctured state manifold that excludes the singular set, we obtain a well-defined gauge structure and identify an asymptotically robust invariant: the Uhlmann holonomy around noncontractible loops encircling the defect on a restricted two-dimensional punctured submanifold. In an exactly solvable qutrit model, a restricted submanifold emerges on which the connection is locally flat yet carries nontrivial monodromy, analogous to flat connections with Aharonov--Bohm-type transport. The holonomy depends only on the ratios of the vanishing eigenvalues under frozen radial dependence of the eigenbasis geometry and a fixed angular loop. In contrast, the Uhlmann curvature may diverge path-dependently when eigenvalues shrink with distinct powers, with a leading spectral-prefactor scaling law, establishing that the holonomy survives as a universal asymptotic invariant while the curvature remains non-universal. Within the effective SU(2) defect sector, the conjugacy class of the holonomy, equivalently the Wilson loop variable, provides a continuous, non-quantized classification of the asymptotic monodromy surrounding the rank-deficient defect. This non-quantization does not imply a lack of robustness: the asymptotic holonomy is an invariant of the restricted punctured submanifold and is insensitive to smooth deformations of the loop or the radial profile within the fixed spectral-ratio sector.

quant-ph

Geometry near rank-changing points on the mixed-state manifold: Bures metric, conical singularities, and Lindblad dynamics

We elucidate the Bures metric in quantum state space near a rank-changing point of the density matrix and show contrasting behavior for two-level ($N=2$) systems versus higher-level systems. Due to the smooth pure-state boundary for $N=2$, we prove the apparent metric divergences to be merely coordinate artifacts and present three Lindblad processes exhibiting qualitatively different evolution near rank-changing points, showing geodesic approach, power-law scaling, and pure-state escape law. For higher-dimensional ($N\ge 3$) systems, the geometry near a rank-changing point differs fundamentally. Under suitable restrictions of the density matrix and its approach towards a pure state, the Bures metric reduces to a conical metric with the pure state at the cone tip. Such a conic geometry leads to genuine curvature singularities: A two-dimensional cone exhibits a Dirac delta-function curvature near the tip while a higher-dimensional cone shows a power-law divergence of the curvature towards the cone tip. A construction of Lindblad evolution for $N=3$ systems with conic singularities is presented, along with possible implications for future experimental and theoretical research.

quant-ph

Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard fine-tuning at ~40% less peak training memory, while leaving off-domain benchmarks within seed-noise of baseline, a property that full-network fine-tuning does not reliably reproduce. FPO rests on a single empirical observation: at late layers of a transformer, the output-layer prediction error approximates the true gradient with cosine similarity 0.47--0.59 across six public models we survey. We introduce a two-minute diagnostic that quantifies this approximation per layer for any model, identifying where late-layer adaptation is viable. Informed by the diagnostic, FPO computes a single error signal at the output and applies it to each target layer. No signal is propagated between layers, and no autograd graph is constructed at any point. We evaluate FPO on three model families (OLMo-2-7B, Qwen3-8B, Falcon3-7B). Across all three, FPO produces in-domain perplexity improvement and leaves MMLU, ARC-Challenge, HellaSwag, and Winogrande within seed-noise of baseline. Localizing SFT to FPO's target layers to enter this regime is also feasible, but at 2.2x the wall-clock cost of FPO.

cs.LG

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedly re-teach preferences, corrections, and project context, and context-based workarounds consume context-window space and degrade under cascading compaction. We evaluate an alternative: nightly consolidation of interaction knowledge into model weights via reflection, synthesis, and Low-Rank Adaptation (LoRA) fine-tuning on a single consumer GPU. Across ten realistic software development conversations (n = 10, 1,146 test questions across three memory types), three cycles of cascading compaction retain 36.8 +/- 3.0% of knowledge (between an 11.8% no-context floor and a 90.1% full-context ceiling), while consolidation retains 80.4 +/- 1.3% -- a 43.6 pp gain (paired t(9) = 14.8, p < 0.001) that more than doubles what compaction preserves, with the largest gains on procedural corrections (36.3% -> 74.6%) and episodic project facts (31.5% -> 78.2%). As a methodological aside, mean per-token validation cross-entropy is negatively correlated with LLM-judged accuracy (r = -0.51) while median per-token validation cross-entropy tracks accuracy almost exactly (r = +0.99): under evaluators that tolerate surface-form variation, the mean is misleading and a heavy-tail-robust statistic is the faithful signal. Persistent personalization requires moving beyond inference-only deployment toward architectures that consolidate knowledge into weights.

cs.AI

When Mean CE Fails: Median CE Can Better Track Language Model Quality

Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two common scenarios. First, in Qwen2.5-1.5B SFT on synthetic fact-learning, we find that mean CE rises substantially after the initial learning phase while held-out fact-recall accuracy remains near its peak. Second, we find that in top-K distillation on TinyStories, decreasing K improves median CE while worsening mean CE; the Top-5 student attains the highest LLM-judge score and crosses below its teacher on median CE, despite having the worst mean CE. In both cases, median CE correlates much more closely with task performance than does mean CE. Analyzing how bulk and tail percentile CE move during training reveals that training reshapes the empirical per-token CE distribution. In top-K distillation, smaller K yields a distribution with more mass at both extremes, decreasing the median and increasing the mean. In Qwen SFT, the bulk saturates quickly while the tail extends in the latter half of training. In both, the task-evaluation metric appears more sensitive to the bulk than to the tail. Practically, we recommend reporting a small set of percentile CE summaries alongside the mean, and using concordance among them as a tool to keep track of distribution reshaping, as well as a low-cost diagnostic for when mean and median CE disagree on model selection.

cs.AI