SearcharxivSearch

arXiv subjects

Chao Kong

Publications and source records attributed to Chao Kong.

12 recordsLinked to original sources

Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm Detection

Multimodal sarcasm detection aims to identify sarcastic intent from multimodal content, where inconsistencies between literal meaning and contextual cues often signal irony. This task has attracted increasing research attention. However, accurate detection remains challenging due to instance-dependent modality contributions and misleading semantic consistency, where surface-level alignment masks underlying contradictory intent. Existing methods often rely on fixed fusion strategies and treat sarcasm as generic cross-modal mismatch, limiting their ability to capture subtle sarcasm cues and instance-specific modality interactions. To address these challenges, we propose a novel MSD framework that integrates Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization (SaCR). Specifically, a bidirectional gated interaction module performs cross-modal feature filtering and adaptively calibrates textual and visual contributions at the instance level. A dynamic fusion gate further balances modality importance to generate more robust multimodal representations. Furthermore, SaCR is introduced as a label-aware contrastive regularization objective that encourages semantic consistency for non-sarcastic samples while suppressing misleading consistency in sarcastic cases. The proposed framework is trained end-to-end with a multi-objective learning strategy that jointly optimizes multimodal classification and auxiliary unimodal supervision. Extensive experiments on MMSD and MMSD2.0 demonstrate that the proposed method consistently outperforms strong baselines.

cs.CL

Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.

cs.CL

Stable three-dimensional solitons in spin-orbit-coupled atomic-molecular condensates

We elaborate a mechanism for the creation of stable three-dimensional (3D) solitons in spin-orbit-coupled (SOC) atomic-molecular Bose-Einstein condensate, modeled by the mean-field equations with the quadratic three-wave interaction, characterized by mismatch $\alpha $. The planar (effectively two-dimensional) SOC is applied to the soliton's atomic component, structuring it as a mixed mode (MM) or semi-vortex (SV). The molecular component of the SV soliton is shaped as a 3D vortex, while the molecular component in the MM soliton is an MM too. The solitons exist up to a critical value of $\alpha $. The system demonstrates a relatively large norm share of the vortex components, exceeding $50\%$ of the total norm, which is an essential feature of SOC-supported solitons. This is scheme for realizing stable vortex solitons in free space with the quadratic nonlinearity.

quant-ph

o1-Coder: an o1 Replication for Coding

The technical report introduces O1-CODER, an attempt to replicate OpenAI's o1 model with a focus on coding tasks. It integrates reinforcement learning (RL) and Monte Carlo Tree Search (MCTS) to enhance the model's System-2 thinking capabilities. The framework includes training a Test Case Generator (TCG) for standardized code testing, using MCTS to generate code data with reasoning processes, and iteratively fine-tuning the policy model to initially produce pseudocode and then generate the full code. The report also addresses the opportunities and challenges in deploying o1-like models in real-world applications, suggesting transitioning to the System-2 paradigm and highlighting the imperative for world model construction. Updated model progress and experimental results will be reported in subsequent versions. All source code, curated datasets, as well as the derived models are disclosed at https://github.com/ADaM-BJTU/O1-CODER .

cs.SE

Composite solitary vortices of three-wave mixing in quasi-phase-matched photonic crystals

We report the composite vortex solitons of three-wave mixing propagate stably in a three-dimensional (3D) quasi-phase-matched photonic crystals (QPM-PhC). The modulation of QPM-PhC is designed as a checkerboard pattern. The vortex solitons, composed by three waves ($\omega_{1,2,3}$) propagating through the lattices, exhibit a four-spotted discrete type, which gives rise to four distinct modes: zero-vorticity, vortex, anti-vortex, and quadrupole. The composite vortex solitons result from combinations of these modes and lead to four cases: vortex doubling, hidden vortices, vortex up-conversion, and anti-vortex up-conversion. Our findings indicate that all solitons can propagate stably through the crystals for 10 centimeters; however, only the vortex-doubling case remains stable over longer distances. This work enhances the understanding of vortex beam manipulation within 3D QPM-PhCs.

physics.optics

Semi-vortex solitons and their excited states in spin-orbit-coupled binary bosonic condensates

It is known that two-dimensional two-component fundamental solitons of the semi-vortex (SV) type, with vorticities $(s_{+},s_{-})=(0,1)$ in their components, are stable ground states (GSs) in the spin-orbit-coupled (SOC) binary Bose-Einstein condensate with the contact self-attraction acting in both components, in spite of the possibility of the critical collapse in the system. However, excited states(ESs) of the SV solitons, with the vorticity set $(s_{+},s_{-})=( S_{+},S_{+}+1)$ and $S_{+}=1,2,3,...$, are unstable in the same system. We construct ESs of SV solitons in the SOC system with opposite signs of the self-interaction in the two components. The main finding is stability of the ES-SV solitons, with the extra vorticity (at least) up to $S_{+}=6$. The threshold value of the norm for the onset of the critical collapse, $N_{\mathrm{thr}}$, in these excited states is higher than the commonly known critical value, $N_{c}\approx 5.85$,associated with the single-component Townes solitons, $N_{\mathrm{thr}}$ increasing with the growth of $S_{+}$. A velocity interval for stable motion of the GS-SV solitons is found too. The results suggest a solution for the challenging problem of the creation of stable vortex solitons with high topological charges.

cond-mat.quant-gas

Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning

This paper presents a follow-up study to OpenAI's recent superalignment work on Weak-to-Strong Generalization (W2SG). Superalignment focuses on ensuring that high-level AI systems remain consistent with human values and intentions when dealing with complex, high-risk tasks. The W2SG framework has opened new possibilities for empirical research in this evolving field. Our study simulates two phases of superalignment under the W2SG framework: the development of general superhuman models and the progression towards superintelligence. In the first phase, based on human supervision, the quality of weak supervision is enhanced through a combination of scalable oversight and ensemble learning, reducing the capability gap between weak teachers and strong students. In the second phase, an automatic alignment evaluator is employed as the weak supervisor. By recursively updating this auto aligner, the capabilities of the weak teacher models are synchronously enhanced, achieving weak-to-strong supervision over stronger student models.We also provide an initial validation of the proposed approach for the first phase. Using the SciQ task as example, we explore ensemble learning for weak teacher models through bagging and boosting. Scalable oversight is explored through two auxiliary settings: human-AI interaction and AI-AI debate. Additionally, the paper discusses the impact of improved weak supervision on enhancing weak-to-strong generalization based on in-context learning. Experiment code and dataset will be released at https://github.com/ADaM-BJTU/W2SG.

cs.CL

CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models

As the scaling of Large Language Models (LLMs) has dramatically enhanced their capabilities, there has been a growing focus on the alignment problem to ensure their responsible and ethical use. While existing alignment efforts predominantly concentrate on universal values such as the HHH principle, the aspect of culture, which is inherently pluralistic and diverse, has not received adequate attention. This work introduces a new benchmark, CDEval, aimed at evaluating the cultural dimensions of LLMs. CDEval is constructed by incorporating both GPT-4's automated generation and human verification, covering six cultural dimensions across seven domains. Our comprehensive experiments provide intriguing insights into the culture of mainstream LLMs, highlighting both consistencies and variations across different dimensions and domains. The findings underscore the importance of integrating cultural considerations in LLM development, particularly for applications in diverse cultural settings. Through CDEval, we aim to broaden the horizon of LLM alignment research by including cultural dimensions, thus providing a more holistic framework for the future development and evaluation of LLMs. This benchmark serves as a valuable resource for cultural studies in LLMs, paving the way for more culturally aware and sensitive models.

cs.CL

Towards Alleviating the Object Bias in Prompt Tuning-based Factual Knowledge Extraction

Many works employed prompt tuning methods to automatically optimize prompt queries and extract the factual knowledge stored in Pretrained Language Models. In this paper, we observe that the optimized prompts, including discrete prompts and continuous prompts, exhibit undesirable object bias. To handle this problem, we propose a novel prompt tuning method called MeCoD. consisting of three modules: Prompt Encoder, Object Equalization and Biased Object Obstruction. Experimental results show that MeCoD can significantly reduce the object bias and at the same time improve accuracy of factual knowledge extraction.

cs.IR

Controlling directed atomic motion and second-order tunneling of a spin-orbit-coupled atom in optical lattices

We theoretically explore the tunneling dynamics for the tight-binding (TB) model of a single spin-orbit-coupled atom trapped in an optical lattice subjected to lattice shaking and to time-periodic Zeeman field. By means of analytical and numerical methods, we demonstrate that the spin-orbit (SO) coupling adds some new results to the tunneling dynamics in both multiphoton resonance and far-off-resonance parameter regimes. When the driving frequency is resonant with the static Zeeman field (multi-photon resonances), we obtain an unexpected new dynamical localization (DL) phenomenon where the single SO-coupled atom is restricted to making perfect two-site Rabi oscillation accompanied by spin flipping.By using the unconventional DL phenomenon, we are able to generate a ratchetlike effect which enables directed atomic motion towards different directions and accompanies periodic spin-flipping under the action of SO coupling. For the far-off-resonance case, we show that by suppressing the usual inter-site tunneling alone, it is possible to realize a type of spin-conserving second-order tunneling between next-nearest-neighboring sites, which is not accessible in the conventional lattice system without SO coupling. We also show that simultaneous controls of the usual inter-site tunneling and the SO-coupling-related second-order-tunneling are necessary for quasienergies flatness (collapse) and completely frozen dynamics to exist. These results may be relevant to potential applications such as spin-based quantum information processing and design of novel spintronics devices.

cond-mat.quant-gas

Spatiotemporal Bloch states of a spin-orbit coupled Bose-Einstein condensate in an optical lattice

We study the spatiotemporal Bloch states of a high-frequency driven two-component Bose-Einstein condensate (BEC) with spin-orbit coupling (SOC) in an optical lattice. By adopting the rotating-wave approximation (RWA) and applying an exact trial-solution to the corresponding quasistationary system, we establish a different method for tuning SOC via external field such that the existence conditions of the exact particular solutions are fitted. Several novel features related to the exact states are demonstrated, such as SOC leads to spin-motion entanglement for the spatiotemporal Bloch states, SOC increases the population imbalance of the two-component BEC and SOC can be applied to manipulate the stable atomic flow which is conducive to control quantum transport of the BEC for different application purposes.

quant-ph

A novel exact solution to transmission problem of electron wave in a nonlinear Kronig-Penney superlattice

Nonlinear Kronig-Penney model has been frequently employed to study transmission problem of electron wave in a nonlinear electrified chain or in a doped semiconductor superlattice. Here from an integral equation we derive a novel exact solution of the problem, which contains a simple nonlinear map connecting transmission coefficient with system parameters. Consequently, we suggest a scheme for manipulating electronic distribution and transmission by adjusting the system parameters. A new effect of quantum coherence is evidenced in the strict expression of transmission coefficient by which for some different system parameters we obtain the similar aperiodic distributions and arbitrary transmission coefficients including the approximate zero transmission and total transmission, and the multiple transmissions. The method based on the concise exact solution can be applied directly to some nonlinear cold atomic systems and a lot of linear Kronig-Penney systems, and also can be extended to investigate electron transport in different discrete nonlinear systems.

quant-ph