SearcharxivSearch

arXiv subjects

Chao Chang

Publications and source records attributed to Chao Chang.

12 recordsLinked to original sources

Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism

Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and proposed a novel and training-free approach, termed \emph{Flexible Memory} (\textbf{FlexMem}). In principle, FlexMem aims to mimic human behavior of video watching, \emph{i.e.}, continually watching video content and recalling the most relevant memory fragments to answer the question. In this way, FlexMem can help MLLMs achieve video understanding of infinite lengths, unlike previous methods that process all video information at once and have input upper-limit. Concretely, FlexMem first consider the visual KV caches as the memory sources, and realize the effective memory transfer and writing via a dual-pathway compression design. Afterwards, FlexMem also explores different memory reading strategies for the diverse video understanding tasks, including the popular streaming one. To validate FlexMem, we apply it to two popular video-MLLMs, and conduct extensive experiments on five long video and one streaming video task. The experimental results show that on \textbf{a single 3090 GPU}, our FlexMem can achieve obvious improvements than existing efficient video understanding methods and process more than \textbf{1k frames}, which also helps the base MLLMs achieve comparable or even better performance than SOTA MLLMs on some benchmarks, \emph{e.g.} , GPT-4o and Gemini-1.5 Pro.

cs.CV

Improving Sketching Algorithms for Low-Rank Matrix Approximation via Sketch-Power Iterations

Power iteration can improve the accuracy of randomized SVD, but requires multiple data passes, making it impractical in streaming or memory-constrained settings. We introduce a lightweight yet effective sketch-power iteration, allowing power-like iterations with only a single pass of the data, which can be incorporated into one-pass algorithms for low-rank approximation. As an example, we integrate the sketch-power iteration into a one-pass algorithm proposed by Tropp et al., and introduce strategies to reduce its storage cost. We establish meaningful error bounds: given a fixed storage budget, the sketch sizes derived from the bounds closely match the optimal ones observed in reality. This allows one to preselect reasonable parameters. Numerical experiments on both synthetic and real-world datasets indicate that, under the same storage constraints, applying one or two sketch-power iterations can substantially improve the approximation accuracy of the considered one-pass algorithms. In particular, experiments on real data with flat spectrum show that the method can approximate the dominant singular vectors well.

math.NA

Large Language Models for Software Testing Education: an Experience Report

The rapid integration of Large Language Models (LLMs) into software engineering practice is reshaping how software testing activities are performed. LLMs are increasingly used to support software testing. Consequently, software testing education must evolve to prepare students for this new paradigm. However, while students have already begun to use LLMs in an ad hoc manner for testing tasks, there is limited empirical understanding of how such usage influences their testing behaviors, judgment, and learning outcomes. It is necessary to conduct a systematic investigation into how students learn to evaluate, control, and refine LLM-assisted testing results. This paper presents a mixed-methods, two-phase exploratory study on human-LLM collaboration in software testing education. In Phase I, we analyze classroom learning artifacts and interaction records from 15 students, together with a large-scale survey conducted in a national software testing competition (337 valid responses), to identify recurring prompt-related difficulties across testing tasks. The results reveal systematic interaction breakdowns, including missing contextual information, insufficient constraints, rigid one-shot prompting, and limited strategy-driven iteration, with automated test script generation emerging as a particularly heterogeneous and effort-intensive interaction context. Building on these findings, Phase II conducts an illustrative classroom practice that operationalizes the observed breakdowns into a lightweight, stage-aware prompt scaffold for test script generation, guiding students to explicitly articulate execution-relevant information such as environmental assumptions, interaction grounding, synchronization, and validation intent, and reporting descriptive shifts in students' testing-related articulation when interacting with LLMs.

cs.SE

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling

Due to the great saving of computation and memory overhead, token compression has become a research hot-spot for MLLMs and achieved remarkable progress in image-language tasks. However, for the video, existing methods still fall short of high-ratio token compression. We attribute this shortcoming to the insufficient modeling of temporal and continual video content, and propose a novel and training-free token pruning method for video MLLMs, termed ForestPrune, which achieves effective and high-ratio pruning via Spatial-temporal Forest Modeling. In practice, ForestPrune construct token forests across video frames based on the semantic, spatial and temporal constraints, making an overall comprehension of videos. Afterwards, ForestPrune evaluates the importance of token trees and nodes based on tree depth and node roles, thereby obtaining a globally optimal pruning decision. To validate ForestPrune, we apply it to two representative video MLLMs, namely LLaVA-Video and LLaVA-OneVision, and conduct extensive experiments on a bunch of video benchmarks. The experimental results not only show the great effectiveness for video MLLMs, e.g., retaining 95.8% average accuracy while reducing 90% tokens for LLaVA-OneVision, but also show its superior performance and efficiency than the compared token compression methods, e.g., +10.1% accuracy on MLVU and -81.4% pruning time than FrameFusion on LLaVA-Video.

cs.CV

Quantum Tunneling Enables High-Flux Transport in Ion Channels

Classical molecular dynamics and electro-diffusion theories have achieved profound success in elucidating ion selectivity and gating mechanisms. However, reconciling strict selectivity with high flux permeation in Angstrom-scaled biological ion channels poses a universal challenge in nanoscale physics, as classical models consistently underestimate single-channel conductance. Using a non perturbative quantum transport framework, we calculate the ion permeation dynamics through the selectivity filter within a transfer matrix formalism. We demonstrate that quantum tunneling allows ions to bypass classical Arrhenius suppression, quantitatively recovering the experimental conductance of Na+ and K+ channels. Crucially, our findings reveal that the exploitation of quantum mechanics is a fundamental prerequisite for achieving macroscopic physiological efficiency. By reframing ion channels as mesoscopic quantum conductors, this work establishes a transformative paradigm in quantum biology and predicts distinct transport resonances in the terahertz regime.

physics.optics

DeepInv: A Novel Self-supervised Learning Approach for Fast and Accurate Diffusion Inversion

Diffusion inversion is a task of recovering the noise of an image in a diffusion model, which is vital for controllable diffusion image editing. At present, diffusion inversion still remains a challenging task due to the lack of viable supervision signals. Thus, most existing methods resort to approximation-based solutions, which however are often at the cost of performance or efficiency. To remedy these shortcomings, we propose a novel self-supervised diffusion inversion approach in this paper, termed Deep Inversion (DeepInv). Instead of requiring ground-truth noise annotations, we introduce a self-supervised objective as well as a data augmentation strategy to generate high-quality pseudo noises from real images without manual intervention. Based on these two innovative designs, DeepInv is also equipped with an iterative and multi-scale training regime to train a parameterized inversion solver, thereby achieving the fast and accurate image-to-noise mapping. To the best of our knowledge, this is the first attempt of presenting a trainable solver to predict inversion noise step by step. The extensive experiments show that our DeepInv can achieve much better performance and inference speed than the compared methods, e.g., +40.435% SSIM than EasyInv and +9887.5% speed than ReNoise on COCO dataset. Moreover, our careful designs of trainable solvers can also provide insights to the community. Codes and model parameters will be released in https://github.com/potato-kitty/DeepInv.

cs.CV

Terahertz frequency conversion at plasma-induced time boundary

We report on the frequency conversions of terahertz (THz) waves at ultrafast time boundaries created via femtosecond laser-induced air-to-plasma phase transitions. Our combined experimental and theoretical approach reveals that the abrupt change in refractive index at the ultrafast time boundaries drives both the red and blue shifts over the broadband THz spectrum due to the dispersive plasma, with distinctive amplitude variations. The present study contrasts these effects with those from spatial boundaries, highlighting the superior efficacy of temporal manipulations for spectral engineering. These findings not only deepen the understanding of light-matter interactions in time-varying media but also pave the way for innovative applications in THz technology and lay the groundwork for the observation of temporal reflection effects, photonic time crystals, and spatio-temporally modulated matter.

physics.optics

Randomized Large-Scale Quaternion Matrix Approximation: Practical Rangefinders and One-Pass Algorithm

Recently, randomized algorithms for low-rank approximation of quaternion matrices have received increasing attention. However, for large-scale problems, existing quaternion orthonormalizations are inefficient, leading to slow rangefinders. To address this, by appropriately leveraging efficient scientific computing libraries in the complex arithmetic, this work devises two practical quaternion rangefinders, one of which is non-orthonormal yet well-conditioned. They are then integrated into the quaternion version of a one-pass algorithm, which originally takes orthonormal rangefinders only. We establish the error bounds and demonstrate that the error is proportional to the condition number of the rangefinder. The probabilistic bounds are exhibited for both quaternion Gaussian and sub-Gaussian embeddings. Numerical experiments demonstrate that the one-pass algorithm with the proposed rangefinders significantly outperforms previous techniques in efficiency. Additionally, we tested the algorithm in a 3D Navier-Stokes equation ($5.22$GB) and a 4D Lorenz-type chaotic system ($5.74$GB) data compression, as well as a $31365\times 27125$ image compression to demonstrate its capability for handling large-scale applications.

math.NA

Ultra-small topological spin textures with size of 1.3nm at above room temperature in Fe78Si9B13 amorphous alloy

Topologically protected spin textures, such as skyrmions1,2 and vortices3,4, are robust against perturbations, serving as the building blocks for a range of topological devices5-9. In order to implement these topological devices, it is necessary to find ultra-small topological spin textures at room temperature, because small size implies the higher topological charge density, stronger signal of topological transport10,11 and the higher memory density or integration for topological quantum devices5-9. However, finding ultra-small topological spin textures at high temperatures is still a great challenge up to now. Here we find ultra-small topological spin textures in Fe78Si9B13 amorphous alloy. We measured a large topological Hall effect (THE) up to above room temperature, indicating the existence of highly densed and ultra-small topological spin textures in the samples. Further measurements by small-angle neutron scattering (SANS) reveal that the average size of ultra-small magnetic texture is around 1.3nm. Our Monte Carlo simulations show that such ultra-small spin texture is topologically equivalent to skyrmions, which originate from competing frustration and Dzyaloshinskii-Moriya interaction12,13 coming from amorphous structure14-17. Taking a single topological spin texture as one bit and ignoring the distance between them, we evaluated the ideal memory density of Fe78Si9B13, which reaches up to 4.44*104 gigabits (43.4 TB) per in2 and is 2 times of the value of GdRu2Si218 at 5K. More important, such high memory density can be obtained at above room temperature, which is 4 orders of magnitude larger than the value of other materials at the same temperature. These findings provide a unique candidate for magnetic memory devices with ultra-high density.

cond-mat.mtrl-sci

NGAT4Rec: Neighbor-Aware Graph Attention Network For Recommendation

Learning informative representations (aka. embeddings) of users and items is the core of modern recommender systems. Previous works exploit user-item relationships of one-hop neighbors in the user-item interaction graph to improve the quality of representation. Recently, the research of Graph Neural Network (GNN) for recommendation considers the implicit collaborative information of multi-hop neighbors to enrich the representation. However, most works of GNN for recommendation systems do not consider the relational information which implies the expression differences of different neighbors in the neighborhood explicitly. The influence of each neighboring item to the representation of the user's preference can be represented by the correlation between the item and neighboring items of the user. Symmetrically, for a given item, the correlation between one neighboring user and neighboring users can reflect the strength of signal about the item's characteristic. To modeling the implicit correlations of neighbors in graph embedding aggregating, we propose a Neighbor-Aware Graph Attention Network for recommendation task, termed NGAT4Rec. It employs a novel neighbor-aware graph attention layer that assigns different neighbor-aware attention coefficients to different neighbors of a given node by computing the attention among these neighbors pairwisely. Then NGAT4Rec aggregates the embeddings of neighbors according to the corresponding neighbor-aware attention coefficients to generate next layer embedding for every node. Furthermore, we combine more neighbor-aware graph attention layer to gather the influential signals from multi-hop neighbors. We remove feature transformation and nonlinear activation that proved to be useless on collaborative filtering. Extensive experiments on three benchmark datasets show that our model outperforms various state-of-the-art models consistently.

cs.IR

Graph Attention Collaborative Similarity Embedding for Recommender System

We present Graph Attention Collaborative Similarity Embedding (GACSE), a new recommendation framework that exploits collaborative information in the user-item bipartite graph for representation learning. Our framework consists of two parts: the first part is to learn explicit graph collaborative filtering information such as user-item association through embedding propagation with attention mechanism, and the second part is to learn implicit graph collaborative information such as user-user similarities and item-item similarities through auxiliary loss. We design a new loss function that combines BPR loss with adaptive margin and similarity loss for the similarities learning. Extensive experiments on three benchmarks show that our model is consistently better than the latest state-of-the-art models.

cs.IR

Anomalous behavior of membrane fluidity caused by copper-copper bond coupled phospholipids

Membrane fluidity, well-known to be essential for cell functions, is obviously affected by copper. However, the underlying mechanism is still far from being understood, especially on the atomic level. Here, we unexpectedly observed that a decrease in phospholipid (PL) bilayer fluidity caused by Cu2+ was much more significant than those induced by Zn2+ and Ca2+, while a comparable reduction occurred in the last two ions. This finding clearly disagrees with the placement in the periodic table of Cu just next to Zn and far from Ca. The physical nature was revealed to be a special attraction between Cu+ cations, which can induce a motif forming of two phospholipids coupled by Cu-Cu bond (PL-diCu-PL). Namely, upon Cu2+ ion binding to a negatively charged phosphate group of lipid, Cu2+ was reduced to Cu+. The special attraction of the cations then caused one Cu+ ion simultaneously binding to two lipids and another Cu+, resulting in the formation of PL-diCu-PL structure. In contrast, this attraction cannot occur in the cases of Zn and Ca ions due to their electron structure. Remarkably, besides lipids, the phosphate group widely exists in other biological molecules, including DNA, RNA, ADP and ATP, which would also induce the similar structure of Cu ions with the molecules. Our findings thus provide a new view for understanding the biological functions of copper and the mechanism underlying copper-related diseases.

physics.bio-ph