Searcharxiv⌕ Search

arXiv subjects

Meng Luo

Publications and source records attributed to Meng Luo.

29 records · Page 2Linked to original sources

PAD: Personalized Alignment of LLMs at Decoding-Time

Aligning with personalized preferences, which vary significantly across cultural, educational, and political differences, poses a significant challenge due to the computational costs and data demands of traditional alignment methods. In response, this paper presents Personalized Alignment at Decoding-time (PAD), a novel framework designed to align LLM outputs with diverse personalized preferences during the inference phase, eliminating the need for additional training. By introducing a unique personalized reward modeling strategy, this framework decouples the text generation process from personalized preferences, facilitating the generation of generalizable token-level personalized rewards. The PAD algorithm leverages these rewards to guide the decoding process, dynamically tailoring the base model's predictions to personalized preferences. Extensive experimental results demonstrate that PAD not only outperforms existing training-based alignment methods in terms of aligning with diverse preferences but also shows significant generalizability to preferences unseen during training and scalability across different base models. This work advances the capability of LLMs to meet user needs in real-time applications, presenting a substantial step forward in personalized LLM alignment.

cs.CL↗

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined to the singleton text modality, limiting its effectiveness since human emotions are inherently conveyed through multiple modalities. To combat this, we introduce an avatar-based Multimodal ERG (MERG) task, entailing rich text, speech, and facial vision information. We first present a large-scale high-quality benchmark dataset, \textbf{AvaMERG}, which extends traditional text ERG by incorporating authentic human speech audio and dynamic talking-face avatar videos, encompassing a diverse range of avatar profiles and broadly covering various topics of real-world scenarios. Further, we deliberately tailor a system, named \textbf{Empatheia}, for MERG. Built upon a Multimodal Large Language Model (MLLM) with multimodal encoder, speech and avatar generators, Empatheia performs end-to-end MERG, with Chain-of-Empathetic reasoning mechanism integrated for enhanced empathy understanding and reasoning. Finally, we devise a list of empathetic-enhanced tuning strategies, strengthening the capabilities of emotional accuracy and content, avatar-profile consistency across modalities. Experimental results on AvaMERG data demonstrate that Empatheia consistently shows superior performance than baseline methods on both textual ERG and MERG. Overall, this work is expected to pioneer the MERG research by introducing a novel benchmark and an end-to-end model, laying a solid foundation for future advancements in multimodal empathetic response generation.

cs.MM↗

PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis

While existing Aspect-based Sentiment Analysis (ABSA) has received extensive effort and advancement, there are still gaps in defining a more holistic research target seamlessly integrating multimodality, conversation context, fine-granularity, and also covering the changing sentiment dynamics as well as cognitive causal rationales. This paper bridges the gaps by introducing a multimodal conversational ABSA, where two novel subtasks are proposed: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, rationale from multi-turn multi-party multimodal dialogue. 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation with the causal reasons. To benchmark the tasks, we construct PanoSent, a dataset annotated both manually and automatically, featuring high quality, large scale, multimodality, multilingualism, multi-scenarios, and covering both implicit and explicit sentiment elements. To effectively address the tasks, we devise a novel Chain-of-Sentiment reasoning framework, together with a novel multimodal large language model (namely Sentica) and a paraphrase-based verification mechanism. Extensive evaluations demonstrate the superiority of our methods over strong baselines, validating the efficacy of all our proposed methods. The work is expected to open up a new era for the ABSA community, and thus all our codes and data are open at https://PanoSent.github.io/

cs.CL↗

A Survey on Benchmarks of Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and reasoning. Over the past few years, significant efforts have been made to examine MLLMs from multiple perspectives. This paper presents a comprehensive review of 200 benchmarks and evaluations for MLLMs, focusing on (1)perception and understanding, (2)cognition and reasoning, (3)specific domains, (4)key capabilities, and (5)other modalities. Finally, we discuss the limitations of the current evaluation methods for MLLMs and explore promising future directions. Our key argument is that evaluation should be regarded as a crucial discipline to support the development of MLLMs better. For more details, please visit our GitHub repository: https://github.com/swordlidev/Evaluation-Multimodal-LLMs-Survey.

cs.CL↗

NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations

This paper describes the architecture of our system developed for Task 3 of SemEval-2024: Multimodal Emotion-Cause Analysis in Conversations. Our project targets the challenges of subtask 2, dedicated to Multimodal Emotion-Cause Pair Extraction with Emotion Category (MECPE-Cat), and constructs a dual-component system tailored to the unique challenges of this task. We divide the task into two subtasks: emotion recognition in conversation (ERC) and emotion-cause pair extraction (ECPE). To address these subtasks, we capitalize on the abilities of Large Language Models (LLMs), which have consistently demonstrated state-of-the-art performance across various natural language processing tasks and domains. Most importantly, we design an approach of emotion-cause-aware instruction-tuning for LLMs, to enhance the perception of the emotions with their corresponding causal rationales. Our method enables us to adeptly navigate the complexities of MECPE-Cat, achieving a weighted average 34.71% F1 score of the task, and securing the 2nd rank on the leaderboard. The code and metadata to reproduce our experiments are all made publicly available.

cs.CL↗

AIM: Automatic Interrupt Modeling for Dynamic Firmware Analysis

The security of microcontrollers, which drive modern IoT and embedded devices, continues to raise major concerns. Within a microcontroller (MCU), the firmware is a monolithic piece of software that contains the whole software stack, whereas a variety of peripherals represent the hardware. As MCU firmware contains vulnerabilities, it is ideal to test firmware with off-the-shelf software testing techniques, such as dynamic symbolic execution and fuzzing. Nevertheless, no emulator can emulate the diverse MCU peripherals or execute/test the firmware. Specifically, the interrupt interface, among all I/O interfaces used by MCU peripherals, is extremely challenging to emulate. In this paper, we present AIM -- a generic, scalable, and hardware-independent dynamic firmware analysis framework that supports unemulated MCU peripherals by a novel interrupt modeling mechanism. AIM effectively and efficiently covers interrupt-dependent code in firmware by a novel, firmware-guided, Just-in-Time Interrupt Firing technique. We implemented our framework in angr and performed dynamic symbolic execution for eight real-world MCU firmware. According to testing results, our framework covered up to 11.2 times more interrupt-dependent code than state-of-the-art approaches while accomplishing several challenging goals not feasible previously. Finally, a comparison with a state-of-the-art firmware fuzzer demonstrates dynamic symbolic execution and fuzzing together can achieve better firmware testing coverage.

cs.CR↗

FakeTagger: Robust Safeguards against DeepFake Dissemination via Provenance Tracking

In recent years, DeepFake is becoming a common threat to our society, due to the remarkable progress of generative adversarial networks (GAN) in image synthesis. Unfortunately, existing studies that propose various approaches, in fighting against DeepFake and determining if the facial image is real or fake, is still at an early stage. Obviously, the current DeepFake detection method struggles to catch the rapid progress of GANs, especially in the adversarial scenarios where attackers can evade the detection intentionally, such as adding perturbations to fool the DNN-based detectors. While passive detection simply tells whether the image is fake or real, DeepFake provenance, on the other hand, provides clues for tracking the sources in DeepFake forensics. Thus, the tracked fake images could be blocked immediately by administrators and avoid further spread in social networks. In this paper, we investigate the potentials of image tagging in serving the DeepFake provenance tracking. Specifically, we devise a deep learning-based approach, named FakeTagger, with a simple yet effective encoder and decoder design along with channel coding to embed message to the facial image, which is to recover the embedded message after various drastic GAN-based DeepFake transformation with high confidence. The embedded message could be employed to represent the identity of facial images, which further contributed to DeepFake detection and provenance. Experimental results demonstrate that our proposed approach could recover the embedded message with an average accuracy of more than 95% over the four common types of DeepFakes. Our research finding confirms effective privacy-preserving techniques for protecting personal photos from being DeepFaked.

cs.CR↗

Spectral Photon Sorting For Large-Scale Cherenkov and Scintillation Detectors

We describe here measurements with a new device, the "dichroicon," a Winston-style light concentrator built out of dichroic reflectors, which could allow large-scale neutrino detectors to sort photons by wavelength with small overall light loss. Photon sorting would benefit large-scale water or ice Cherenkov detectors such as Hyper-Kamiokande or ICECUBE by providing a measure of dispersion, which in turn could allow improved position reconstruction and timing. For scintillator detectors like JUNO, upgrades to SNO+ or KamLAND-ZEN, or to water-based liquid scintillator detectors like Theia, dichroicons would provide effective discrimination between Cherenkov and scintillation light, allowing them to operate as true hybrid detectors. We include measurements with a prototype dichroicon using first a Cherenkov source to show spectral photon sorting works as expected. We then present measurements of two different LAB-based liquid scintillator sources, and demonstrate discrimination between Cherenkov and scintillation light. On the benchtop we can identify Cherenkov light with better than 90% purity while maintaining a high collection efficiency for the scintillation light. First results from simulations of a large-scale detector are also presented.

physics.ins-det↗

Cherenkov and Scintillation Light Separation Using Wavelength in LAB Based Liquid Scintillator

Linear alkyl benzene (LAB) has in recent years been used as a solvent for PPO in large-scale scintillation detectors, like Daya Bay and SNO+. The combination has several nice properties, including high light yield, good materials compatibility, and excellent pulse shape discrimination. As charged particles move through the LAB+PPO, both Cherenkov and scintillation light are created. Separating Cherenkov from scintillation light would allow a broad range of physics in future large-scale detectors like THEIA, by allowing direction reconstruction with Cherenkov light while retaining the high light yield and good particle ID of a scintillator detector. In this paper, we examine the discrimination of Cherenkov and scintillation light using a set of bandpass and dichroic filters. In principle, Cherenkov light emission extends longer in wavelength than the PPO scintillation spectrum, allowing for exclusive identification. We find that by selecting wavelengths above 450 nm the Cherenkov light can be clearly separated from the scintillation light.

physics.ins-det↗

Non-Markovian shot noise spectrum of quantum transport through quantum dots

The generalized quantum master equation with transport particle number resolution, like its conventional unconditioned counterpart, has also the time-local and time-nonlocal prescriptions.The latter is found to be more suitable for the effect of electrodes bandwidth on quantum transport and noise spectrum for weak system-reservoir coupling, as calibrated with the exact results in the absence of Coulomb interaction. We further analyze the effect of Coulomb interaction on the noise spectrum of transport current through quantum dot systems, and show that the realistic finite Coulomb interaction and finite bandwidth are manifested only with non-Markovian treatment. We demonstrate a number of non-Markovian characteristics of shot noise spectrum, including that due to finite bandwidth and that sensitive to and enhanced by the magnitude of Coulomb interaction.

cond-mat.str-el↗

Hierarchical theory of quantum dissipation: Partial fraction decomposition scheme

We propose a partial fraction decomposition scheme to the construction of hierarchical equations of motion theory for bosonic quantum dissipation systems. The expansion of Bose--Einstein function in this scheme shows similar properties as it applies for Fermi function. The performance of the resulting quantum dissipation theory is exemplified with spin--boson systems. In all cases we have tested the new theory performs much better, about an order of magnitude faster, than the best available conventional theory based on Matsubara spectral decomposition scheme.

cond-mat.stat-mech↗