SearcharxivSearch

arXiv subjects

Zijing Chen

Publications and source records attributed to Zijing Chen.

9 recordsLinked to original sources

Extreme Terahertz Nonlinear Phononics by Coherence-Imprinted Control of Hybrid Order

Coherent control of quantum materials has progressed along two major fronts: nonlinear phononics, which reshapes lattices to induce emergent states, and Floquet engineering, which tailors electronic band reconstruction via time-periodic driving. Both mechanisms face fundamental limitations at terahertz (THz) frequencies: phononic nonlinearities are intrinsically weak in standard lattices, while electronic Floquet states are often constrained by rapid decoherence upon light-off and by a scarcity of coherence-resolved, multi-correlation probes beyond (quasi-)stationary band structures. Here we report an extreme THz nonlinear-phononics mechanism in $\text{Ta}_\text{2}\text{NiSe}_\text{5}$, where a highly susceptible non-equilibrium electronic correlation bath dramatically amplifies lattice nonlinearities under coherent driving. Utilizing THz two-dimensional spectroscopy as a coherence-tomography tool, we resolve an exceptionally rich landscape of approximately 30 distinct multi-order quantum pathways, including high-harmonic phonon generation, multi-quantum coherences, and multi-wave anharmonic cross-mode mixing. The density and complexity of this extreme manifold establishes a new benchmark for THz nonlinear phononics, as the multi-order quantum pathways surpass the limits of conventional lattice responses. These high-order signals collapse above ~100~K, defining an electronic correlation scale of a coherence-imprinted hybrid electronic-phonon order that governs the sustainability of high-order quantum correlations and nonlinear pathways beyond linear and equilibrium responses. Our results establish a route for correlation-boosted, phonon-anchored periodic Hamiltonian engineering and for certifying such periodically-driven states via multi-correlation coherence tomography.

cond-mat.str-el

Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors

Emotion Recognition (ER) is the process of analyzing and identifying human emotions from sensing data. Currently, the field heavily relies on facial expression recognition (FER) because visual channel conveys rich emotional cues. However, facial expressions are often used as social tools rather than manifestations of genuine inner emotions. To understand and bridge this gap between FER and ER, we introduce eye behaviors as an important emotional cue and construct an Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset. To collect data with genuine emotions, spontaneous emotion induction paradigm is exploited with stimulus material, during which non-invasive eye behavior data, like eye movement sequences and eye fixation maps, is captured together with facial expression videos. To better illustrate the gap between ER and FER, multi-view emotion labels for mutimodal ER and FER are separately annotated. Furthermore, based on the new dataset, we design a simple yet effective Eye-behavior-aided MER Transformer (EMERT) that enhances ER by bridging the emotion gap. EMERT leverages modality-adversarial feature decoupling and a multitask Transformer to model eye behaviors as a strong complement to facial expressions. In the experiment, we introduce seven multimodal benchmark protocols for a variety of comprehensive evaluations of the EMER dataset. The results show that the EMERT outperforms other state-of-the-art multimodal methods by a great margin, revealing the importance of modeling eye behaviors for robust ER. To sum up, we provide a comprehensive analysis of the importance of eye behaviors in ER, advancing the study on addressing the gap between FER and ER for more robust ER performance. Our EMER dataset and the trained EMERT models will be publicly available at https://github.com/kejun1/EMER.

cs.CV

Structural contribution to light-induced gap suppression in Ta$_2$NiSe$_5$

An excitonic insulator is a material that hosts an exotic ground state, where an energy gap opens due to spontaneous condensation of bound electron-hole pairs. Ta$_2$NiSe$_5$ is a promising candidate for this type of material, but the coexistence of a structural phase transition with the gap opening has led to a long-standing debate regarding the origin of the insulating gap. Here we employ MeV ultrafast electron diffraction to obtain quantitative insights into the atomic displacements in Ta$_2$NiSe$_5$ following photoexcitation, which has been overlooked in previous time-resolved spectroscopy studies. In conjunction with first-principles calculations using the measured atomic displacements, we find that the structural change can largely account for the photoinduced reduction in the energy gap without considering excitonic effects. Our work illustrates the importance of a quantitative reconstruction of individual atomic pathways during nonequilibrium phase transitions, paving the way for a mechanistic understanding of a diverse array of phase transitions in correlated materials where lattice dynamics can play a pivotal role.

cond-mat.mtrl-sci

FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis

Object Referring Analysis (ORA), commonly known as referring expression comprehension, requires the identification and localization of specific objects in an image based on natural descriptions. Unlike generic object detection, ORA requires both accurate language understanding and precise visual localization, making it inherently more complex. Although recent pre-trained large visual grounding detectors have achieved significant progress, they heavily rely on extensively labeled data and time-consuming learning. To address these, we introduce a novel, training-free framework for zero-shot ORA, termed FLORA (Formal Language for Object Referring and Analysis). FLORA harnesses the inherent reasoning capabilities of large language models (LLMs) and integrates a formal language model - a logical framework that regulates language within structured, rule-based descriptions - to provide effective zero-shot ORA. More specifically, our formal language model (FLM) enables an effective, logic-driven interpretation of object descriptions without necessitating any training processes. Built upon FLM-regulated LLM outputs, we further devise a Bayesian inference framework and employ appropriate off-the-shelf interpretive models to finalize the reasoning, delivering favorable robustness against LLM hallucinations and compelling ORA performance in a training-free manner. In practice, our FLORA boosts the zero-shot performance of existing pretrained grounding detectors by up to around 45%. Our comprehensive evaluation across different challenging datasets also confirms that FLORA consistently surpasses current state-of-the-art zero-shot methods in both detection and segmentation tasks associated with zero-shot ORA. We believe our probabilistic parsing and reasoning of the LLM outputs elevate the reliability and interpretability of zero-shot ORA. We shall release codes upon publication.

cs.CV

Smile upon the Face but Sadness in the Eyes: Emotion Recognition based on Facial Expressions and Eye Behaviors

Emotion Recognition (ER) is the process of identifying human emotions from given data. Currently, the field heavily relies on facial expression recognition (FER) because facial expressions contain rich emotional cues. However, it is important to note that facial expressions may not always precisely reflect genuine emotions and FER-based results may yield misleading ER. To understand and bridge this gap between FER and ER, we introduce eye behaviors as an important emotional cues for the creation of a new Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset. Different from existing multimodal ER datasets, the EMER dataset employs a stimulus material-induced spontaneous emotion generation method to integrate non-invasive eye behavior data, like eye movements and eye fixation maps, with facial videos, aiming to obtain natural and accurate human emotions. Notably, for the first time, we provide annotations for both ER and FER in the EMER, enabling a comprehensive analysis to better illustrate the gap between both tasks. Furthermore, we specifically design a new EMERT architecture to concurrently enhance performance in both ER and FER by efficiently identifying and bridging the emotion gap between the two.Specifically, our EMERT employs modality-adversarial feature decoupling and multi-task Transformer to augment the modeling of eye behaviors, thus providing an effective complement to facial expressions. In the experiment, we introduce seven multimodal benchmark protocols for a variety of comprehensive evaluations of the EMER dataset. The results show that the EMERT outperforms other state-of-the-art multimodal methods by a great margin, revealing the importance of modeling eye behaviors for robust ER. To sum up, we provide a comprehensive analysis of the importance of eye behaviors in ER, advancing the study on addressing the gap between FER and ER for more robust ER performance.

cs.CV

Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting

In Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these models struggle with unknown classes common in real-world scenarios. In this paper, we introduce a challenging Open-set Video-based Facial Expression Recognition (OV-FER) task, aiming to identify both known and new, unseen facial expressions. While existing approaches use large-scale vision-language models like CLIP to identify unseen classes, we argue that these methods may not adequately capture the subtle human expressions needed for OV-FER. To address this limitation, we propose a novel Human Expression-Sensitive Prompting (HESP) mechanism to significantly enhance CLIP's ability to model video-based facial expression details effectively. Our proposed HESP comprises three components: 1) a textual prompting module with learnable prompts to enhance CLIP's textual representation of both known and unknown emotions, 2) a visual prompting module that encodes temporal emotional information from video frames using expression-sensitive attention, equipping CLIP with a new visual modeling ability to extract emotion-rich information, and 3) an open-set multi-task learning scheme that promotes interaction between the textual and visual modules, improving the understanding of novel human emotions in video sequences. Extensive experiments conducted on four OV-FER task settings demonstrate that HESP can significantly boost CLIP's performance (a relative improvement of 17.93% on AUROC and 106.18% on OSCR) and outperform other state-of-the-art open-set video understanding methods by a large margin. Code is available at https://github.com/cosinehuang/HESP.

cs.CV

Femtosecond electron diffraction reveals local disorder and local anharmonicity in thermoelectric SnSe

The microscopic arrangement of atoms and molecules is the determining factor in how materials behave and perform. Beyond the long-range periodicity, the local disorder with local structures deviating from the average lattice structure plays a vital role in determining the physical properties of the phonon, electron and spin subsystems in crystalline functional materials. Experimentally characterizing the 3D atomic configuration of such local disorder and correlating it with the advanced functions remain a big challenge. Time-domain evolution of the local disorder, either static or dynamical, is lost due to the characterization at equilibrium state with conventional probing techniques. With the combination of femtosecond electron diffraction, structure factor calculation and TDDFT-MD simulation, we exclusively identify the static local disorder and the local anharmonicity of it in thermoelectric SnSe. The ultrafast structural dynamics in time domain reveal a dominant static off-symmetry displacement of Sn (~0.4 angstrom) and the anharmonicity of this local disorder induces an ultrafast atomic displacement within 100 fs after photoexcitation. The microscopic picture of the local anharmonicity indicates a direct and first signature of the THz Einstein oscillators in real space. Therefore, a glass-like thermal transport channel with the local disorder, the Einstein oscillators and the local anharmonicity, updates the fundamental insight into the long-debated ultralow thermal conductivity in SnSe. The local disorder over one to a few unit cells is pervasive and indispensable in thermoelectric materials, multiferroic materials and correlated electronic materials. Our method of revealing the 3D local disorder and the local correlated interactions by ultrafast structural dynamics will inspire broad interest in construction of the structure-property relationship in material science.

cond-mat.mtrl-sci

Noise-Resistant Multimodal Transformer for Emotion Recognition

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful semantics. To this end, we present a novel paradigm that attempts to extract noise-resistant features in its pipeline and introduces a noise-aware learning scheme to effectively improve the robustness of multimodal emotion understanding. Our new pipeline, namely Noise-Resistant Multimodal Transformer (NORM-TR), mainly introduces a Noise-Resistant Generic Feature (NRGF) extractor and a Transformer for the multimodal emotion recognition task. In particular, we make the NRGF extractor learn a generic and disturbance-insensitive representation so that consistent and meaningful semantics can be obtained. Furthermore, we apply a Transformer to incorporate Multimodal Features (MFs) of multimodal inputs based on their relations to the NRGF. Therefore, the possible insensitive but useful information of NRGF could be complemented by MFs that contain more details. To train the NORM-TR properly, our proposed noise-aware learning scheme complements normal emotion recognition losses by enhancing the learning against noises. Our learning scheme explicitly adds noises to either all the modalities or a specific modality at random locations of a multimodal input sequence. We correspondingly introduce two adversarial losses to encourage the NRGF extractor to learn to extract the NRGFs invariant to the added noises, thus facilitating the NORM-TR to achieve more favorable multimodal emotion recognition performance. In practice, on several popular multimodal datasets, our NORM-TR achieves state-of-the-art performance and outperforms existing methods by a large margin, which demonstrates that the ability to resist noisy information is important for effective emotion recognition.

cs.MM

Transient dynamics of the phase transition in VO2 revealed by mega electron-volt ultrafast electron diffraction

Vanadium dioxide (VO2) exhibits an insulator-to-metal transition accompanied by a structural transition near room temperature. This transition can be triggered by an ultrafast laser pulse. Exotic transient states, such as a metallic state without structural transition, were also proposed. These unique characteristics let VO2 have great potential in thermal switchable devices and photonic applications. Although great efforts have been made, the atomic pathway during the photoinduced phase transition is still not clear. Here, we synthesized freestanding quasi-single-crystal VO2 films and examined their photoinduced structural phase transition with mega-electron-volt ultrafast electron diffraction. Leveraging the high signal-to-noise ratio and high temporal resolution, we observe that the disappearance of vanadium dimers and zigzag chains does not coincide with the transformation of crystal symmetry. After photoexcitation, the initial structure is strongly modified within 200 femtoseconds, resulting in a transient monoclinic structure without vanadium dimers and zigzag chains. Then, it continues to evolve to the final tetragonal structure in approximately 5 picoseconds. In addition, only one laser fluence threshold instead of two thresholds suggested in polycrystalline samples was observed in our quasi-single-crystal samples. Our findings provide new essential information for a comprehensive understanding of the photoinduced ultrafast phase transition in VO2.

cond-mat.mtrl-sci