SearcharxivSearch

arXiv subjects

Yuli Lin

Publications and source records attributed to Yuli Lin.

3 recordsLinked to original sources

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-increasing video frames, but also in the unpredictability of future video content and input instructions. In this paper, we study this task from the perspective of constructing a dynamic but fixed-budget memory bank, and propose a novel and training-free approach termed \emph{\textbf{CausalMem}}. CausalMem is dedicated to constructing a dynamic visual memory update mechanism, thereby maximizing the amount of information in streaming video within a limited memory space, much like the human brain. In practice, CausalMem estimates the redundancy of visual tokens and updates the memory bank via an online semantic basis, which models the principal semantics of the observed video stream. To validate CausalMem, we apply it to two representative MLLMs, namely LLaVA-OneVision and Qwen2.5-VL respectively, and conduct extensive experiments on both streaming and offline video understanding benchmarks. The experimental results not only show the great advantages than existing methods under both streaming and offline settings, \emph{e.g.}, $+3.2\%$ and $+3.0\%$ average accuracy gains respectively, but also witness the superior semantic preservation for streaming videos, \emph{e.g.}, using 12$k$ token budgets to memorize hour-long streaming videos, which achieves more than \textbf{20$\times$} visual token compression ratio and only occupies about \textbf{82 MB} storage. \textbf{Our code} is given in \href{https://github.com/hktk07/CausalMem}{CausalMem}.

cs.CV

I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench

Do large language models (LLMs) exhibit any forms of awareness similar to humans? In this paper, we introduce AwareBench, a benchmark designed to evaluate awareness in LLMs. Drawing from theories in psychology and philosophy, we define awareness in LLMs as the ability to understand themselves as AI models and to exhibit social intelligence. Subsequently, we categorize awareness in LLMs into five dimensions, including capability, mission, emotion, culture, and perspective. Based on this taxonomy, we create a dataset called AwareEval, which contains binary, multiple-choice, and open-ended questions to assess LLMs' understandings of specific awareness dimensions. Our experiments, conducted on 13 LLMs, reveal that the majority of them struggle to fully recognize their capabilities and missions while demonstrating decent social intelligence. We conclude by connecting awareness of LLMs with AI alignment and safety, emphasizing its significance to the trustworthy and ethical development of LLMs. Our dataset and code are available at https://github.com/HowieHwong/Awareness-in-LLM.

cs.CL

Electric-field-induced modulation of thermal conductivity in poly(vinylidene fluoride)

Phonon engineering focuses on heat transport modulation on atomic-scale. Different from reported methods, it is shown that electric field can also modulate heat transport in ferroelectric polymers, poly(vinylidene fluoride), by both simulation and measurement. Interestingly, thermal conductivities of poly(vinylidene fluoride) array can be enhanced by a factor of 3.25 along the polarization direction by simulation. The semi-crystalline poly(vinylidene fluoride) film can be also enhanced by a factor of 1.5 which is found by both simulation and measurement. The morphology and phonon property analysis reveal that the enhancement arises from the higher inter-chain lattice order, stronger inter-chain interaction, higher phonon group velocity and suppressed phonon scattering. This study offers a new modulation strategy with quick response and without fillers.

cond-mat.mes-hall