Searcharxiv⌕ Search

arXiv subjects

Lin Chen

Publications and source records attributed to Lin Chen.

At least 199 records · Page 11Linked to original sources

MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for egocentric videos ranging from 30 seconds to one hour long in Ego4D based on human-annotated data. This is one of the largest egocentric QA datasets. Second, we contribute a challenging egocentric QA benchmark with 629 videos and 7,026 questions to evaluate the models' ability in recognizing and memorizing visual details across videos of varying lengths. We introduce a new de-biasing evaluation method to help mitigate the unavoidable language bias present in the models being evaluated. Third, we propose a specialized multimodal architecture featuring a novel "Memory Pointer Prompting" mechanism. This design includes a \textit{global glimpse} step to gain an overarching understanding of the entire video and identify key visual information, followed by a fallback step that utilizes the key visual information to generate responses. This enables the model to more effectively comprehend extended video content. With the data, benchmark, and model, we build MM-Ego, an egocentric multimodal LLM that shows powerful performance on egocentric video understanding.

cs.CV↗

Commutators with multiple unitary symmetry

Commutators are essential in quantum information theory, influencing quantum state symmetries and information storage robustness. This paper systematically investigates the characteristics of bipartite and multipartite quantum states invariant under local unitary group actions. The results demonstrate that any quantum states commuting with $U \otimes U^{\dagger}$ and $U \otimes V$ can be expressed as $\frac{1}{n}I_n$, where $U$ and $V$ are arbitary $n\times n$ unitary matrices. Furthermore, in tripartite systems, any quantum states commuting with $U \otimes U \otimes U^{\dagger}$ must necessarily adopt the form: $W = xI_{n^3} + y\left(\sum_{i,j=1}^n (|i\rangle \langle j|) \otimes (|j\rangle \langle i|)\right) \otimes I_n$, where $F_n$ represents the canonical swap operator. These results provide theoretical tools for characterizing multipartite entanglement constraints and designing symmetry-protected quantum protocols.

quant-ph↗

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video CoT reasoning remains absent. Current video benchmarks fail to adequately assess the reasoning process and expose whether failures stem from deficiencies in perception or reasoning capabilities. Therefore, we introduce VCR-Bench, a novel benchmark designed to comprehensively evaluate LVLMs' Video Chain-of-Thought Reasoning capabilities. VCR-Bench comprises 859 videos spanning a variety of video content and durations, along with 1,034 high-quality question-answer pairs. Each pair is manually annotated with a stepwise CoT rationale, where every step is tagged to indicate its association with the perception or reasoning capabilities. Furthermore, we design seven distinct task dimensions and propose the CoT score to assess the entire CoT process based on the stepwise tagged CoT rationals. Extensive experiments on VCR-Bench highlight substantial limitations in current LVLMs. Even the top-performing model, o1, only achieves a 62.8% CoT score and an 56.7% accuracy, while most models score below 40%. Experiments show most models score lower on perception than reasoning steps, revealing LVLMs' key bottleneck in temporal-spatial information processing for complex video reasoning. A robust positive correlation between the CoT score and accuracy confirms the validity of our evaluation framework and underscores the critical role of CoT reasoning in solving complex video reasoning tasks. We hope VCR-Bench to serve as a standardized evaluation framework and expose the actual drawbacks in complex video reasoning task.

cs.CV↗

Long Arithmetic Progressions in Sumsets and Subset Sums: Constructive Proofs and Efficient Witnesses

Existence of long arithmetic progression in sumsets and subset sums has been studied extensively in the field of additive combinatorics. These additive combinatorics results play a central role in the recent progress of fundamental problems in theoretical computer science including Knapsack and Subset Sum. The non-constructiveness of relevant additive combinatorics results affects their application in algorithms. In particular, several additive combinatorics-based algorithms for Subset Sum work only for the decision version of the problem, but not for the search version. We provide constructive proofs for finite addition theorems [Sárkőzy'89 '94], which are fundamental results in additive combinatorics concerning the existence of long arithmetic progression in sumsets and subset sums. Our constructive proofs yield a near-linear time algorithm that returns an arithmetic progression explicitly, and moreover, for each term in the arithmetic progression, it also returns its representation as the sum of elements in the base set. As an application, we obtain an $\tilde{O}(n)$-time algorithm for the search version of dense subset sum now. Another application of our result is Unbounded Subset Sum, where each input integer can be used an infinite number of times. A classic result on the Frobenius problem [Erdős and Graham '72] implies that for all $t \geq 2a^2_{\max}/n$, the decision version can be solved trivially in linear time. It remains unknown whether the search version can be solved in the same time. Our result implies that for all $t \geq ca^2_{\max}/n$ for some constant $c$, a solution for Unbounded Subset Sum can be obtained in $O(n \log a_{\max})$ time.

cs.DS↗

Temporal-contextual Event Learning for Pedestrian Crossing Intent Prediction

Ensuring the safety of vulnerable road users through accurate prediction of pedestrian crossing intention (PCI) plays a crucial role in the context of autonomous and assisted driving. Analyzing the set of observation video frames in ego-view has been widely used in most PCI prediction methods to forecast the cross intent. However, they struggle to capture the critical events related to pedestrian behaviour along the temporal dimension due to the high redundancy of the video frames, which results in the sub-optimal performance of PCI prediction. Our research addresses the challenge by introducing a novel approach called \underline{T}emporal-\underline{c}ontextual Event \underline{L}earning (TCL). The TCL is composed of the Temporal Merging Module (TMM), which aims to manage the redundancy by clustering the observed video frames into multiple key temporal events. Then, the Contextual Attention Block (CAB) is employed to adaptively aggregate multiple event features along with visual and non-visual data. By synthesizing the temporal feature extraction and contextual attention on the key information across the critical events, TCL can learn expressive representation for the PCI prediction. Extensive experiments are carried out on three widely adopted datasets, including PIE, JAAD-beh, and JAAD-all. The results show that TCL substantially surpasses the state-of-the-art methods. Our code can be accessed at https://github.com/dadaguailhb/TCL.

cs.CV↗

Decidabilities of local unitary equivalence for entanglement witnesses and states

The problem of determining whether two states are equivalent by local unitary (LU) operations is important for quantum information processing. In this paper we propose an alternative perspective to study this problem by comparing the decidabilities of LU equivalence (also known as LU decidabilities for short) between entanglement witnesses and states. We introduce a relation between sets of Hermitian operators in terms of the LU decidability. Then we compare the LU decidability for the set of entanglement witness to those decidabilities for several sets of states, and establish a hierarchy on LU decidabilities for these sets. Moreover, we realize that the simultaneous LU (SLU) equivalence between tuples of mutually orthogonal projectors is crucial to LU equivalent operators. We reveal by examples that for two tuples of projectors, the partial SLU equivalence cannot ensure the overall SLU equivalence. Generally, we present a necessary and sufficient condition such that two tuples of states are SLU equivalent.

quant-ph↗

Temporal Gaussian Copula For Clinical Multivariate Time Series Data Imputation

The imputation of the Multivariate time series (MTS) is particularly challenging since the MTS typically contains irregular patterns of missing values due to various factors such as instrument failures, interference from irrelevant data, and privacy regulations. Existing statistical methods and deep learning methods have shown promising results in time series imputation. In this paper, we propose a Temporal Gaussian Copula Model (TGC) for three-order MTS imputation. The key idea is to leverage the Gaussian Copula to explore the cross-variable and temporal relationships based on the latent Gaussian representation. Subsequently, we employ an Expectation-Maximization (EM) algorithm to improve robustness in managing data with varying missing rates. Comprehensive experiments were conducted on three real-world MTS datasets. The results demonstrate that our TGC substantially outperforms the state-of-the-art imputation methods. Additionally, the TGC model exhibits stronger robustness to the varying missing ratios in the test dataset. Our code is available at https://github.com/MVL-Lab/TGC-MTS.

cs.LG↗

Standing waves with prescribed mass for NLS equations with Hardy potential in the half-space under Neumman boundary condition

Consider the Neumann problem: \begin{eqnarray*} \begin{cases} &-Δu-\fracμ{|x|^2}u +λu =|u|^{q-2}u+|u|^{p-2}u ~~~\mbox{in}~~\mathbb{R}_+^N,~N\ge3, &\frac{\partial u}{\partial ν}=0 ~~ \mbox{on}~~ \partial\mathbb{R}_+^N \end{cases} \end{eqnarray*} with the prescribed mass: \begin{equation*} \int_{\mathbb{R}_+^N}|u|^2 dx=a>0, \end{equation*} where $\mathbb{R}_+^N$ denotes the upper half-space in $\mathbb{R}^N$, $\frac{1}{|x|^2}$ is the Hardy potential, $2 0$, $ν$ stands for the outward unit normal vector to $\partial \mathbb{R}_+^N$, and $λ$ appears as a Lagrange multiplier. Firstly, by applying Ekeland's variational principle, we establish the existence of normalized solutions that correspond to local minima of the associated energy functional. Furthermore, we find a second normalized solution of mountain pass type by employing a parameterized minimax principle that incorporates Morse index information. Our analysis relies on a Hardy inequality in $H^1(\mathbb{R}_+^N)$, as well as a Pohozaev identity involving the Hardy potential on $\mathbb{R}_+^N$. This work provides a variational framework for investigating the existence of normalized solutions to the Hardy type system within a half-space, and our approach is flexible, allowing it to be adapted to handle more general nonlinearities.

math.AP↗

Precision reconstruction of rational CFT from exact fixed point tensor network

The novel concept of entanglement renormalization and its corresponding tensor network renormalization technique have been highly successful in developing a controlled real space renormalization group (RG) scheme. Numerically approximate fixed-point (FP) tensors are widely used to extract the conformal data of the underlying conformal field theory (CFT) describing critical phenomena. In this paper, we present an explicit analytical construction of the FP tensor for 2D rational CFT. We define it as a correlation function between the "boundary-changing operators" (BCO) on triangles. Our construction fully captures all the real-space RG conditions. We also provide concrete examples, such as Ising, Yang-Lee and Tri-critical Ising models to compute the scaling dimensions explicitly based on the corresponding FP tensor. The BCO descendants turn out to be an optimal basis such that truncation in bond dimensions naturally produces comparable accuracies with the leading existing FP algorithms. Interestingly, our construction of FP tensors is closely related to a strange correlator, where the holographic picture naturally emerges. Our results also open a new door towards understanding CFT in higher dimensions.

cond-mat.str-el↗

A Method for Evaluating the Interpretability of Machine Learning Models in Predicting Bond Default Risk Based on LIME and SHAP

Interpretability analysis methods for artificial intelligence models, such as LIME and SHAP, are widely used, though they primarily serve as post-model for analyzing model outputs. While it is commonly believed that the transparency and interpretability of AI models diminish as their complexity increases, currently there is no standardized method for assessing the inherent interpretability of the models themselves. This paper uses bond market default prediction as a case study, applying commonly used machine learning algorithms within AI models. First, the classification performance of these algorithms in default prediction is evaluated. Then, leveraging LIME and SHAP to assess the contribution of sample features to prediction outcomes, the paper proposes a novel method for evaluating the interpretability of the models themselves. The results of this analysis are consistent with the intuitive understanding and logical expectations regarding the interpretability of these models.

q-fin.GN↗

Improving the Stability of GNN Force Field Models by Reducing Feature Correlation

Recently, Graph Neural Network based Force Field (GNNFF) models are widely used in Molecular Dynamics (MD) simulation, which is one of the most cost-effective means in semiconductor material research. However, even such models provide high accuracy in energy and force Mean Absolute Error (MAE) over trained (in-distribution) datasets, they often become unstable during long-time MD simulation when used for out-of-distribution datasets. In this paper, we propose a feature correlation based method for GNNFF models to enhance the stability of MD simulation. We reveal the negative relationship between feature correlation and the stability of GNNFF models, and design a loss function with a dynamic loss coefficient scheduler to reduce edge feature correlation that can be applied in general GNNFF training. We also propose an empirical metric to evaluate the stability in MD simulation. Experiments show our method can significantly improve stability for GNNFF models especially in out-of-distribution data with less than 3% computational overhead. For example, we can ensure the stable MD simulation time from 0.03ps to 10ps for Allegro model.

cs.LG↗

Lighting up the Photon Wigner Distribution via Dilepton Productions

We present a systematic investigation of lepton pair production through photon-photon fusion processes in heavy-ion collisions. It is demonstrated that the dilepton production at a given impact parameter ($b_\perp$) with a fixed transverse momentum imbalance ($q_\perp$) can be factorized into a unified formula in terms of the Wigner photon distribution of heavy nuclei. We show that this framework provides a comprehensive description of all the relevant data from RHIC to the LHC, with a strong evidence that the quasi-real photon can be radiated not only from the nucleus as a whole, standing for the coherent contribution, but also from the sub-structures inside the nucleus, representing the incoherent contribution. Further predictions are made for the anisotropies in the correlations between $q_\perp$, $b_\perp$, and the dilepton transverse momentum ($P_\perp$). This will help us to constrain the photon Wigner distribution which plays a crucial role to study the gluonic matter of nucleus at small-$x$ through the diffractive photoproduction processes in heavy ion collision.

hep-ph↗

A Nearly Quadratic-Time FPTAS for Knapsack

We investigate the classic Knapsack problem and propose a fully polynomial-time approximation scheme (FPTAS) that runs in $\widetilde{O}(n + (1/\varepsilon)^2)$ time. This improves upon the $\widetilde{O}(n + (1/\varepsilon)^{11/5})$-time algorithm by Deng, Jin, and Mao [\textit{Proceedings of the 2023 Annual ACM-SIAM Symposium on Discrete Algorithms, 2023}]. Our algorithm is the best possible (up to a polylogarithmic factor) conditioned on the conjecture that $(\min, +)$-convolution has no truly subquadratic-time algorithm, since this conjecture implies that Knapsack has no $O((n + 1/\varepsilon)^{2-δ})$-time FPTAS for any constant $δ> 0$.

cs.DS↗

Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild

Complex visual reasoning remains a key challenge today. Typically, the challenge is tackled using methodologies such as Chain of Thought (COT) and visual instruction tuning. However, how to organically combine these two methodologies for greater success remains unexplored. Also, issues like hallucinations and high training cost still need to be addressed. In this work, we devise an innovative multi-round training and reasoning framework suitable for lightweight Multimodal Large Language Models (MLLMs). Our self-questioning approach heuristically guides MLLMs to focus on visual clues relevant to the target problem, reducing hallucinations and enhancing the model's ability to describe fine-grained image details. This ultimately enables the model to perform well in complex visual reasoning and question-answering tasks. We have named this framework Socratic Questioning(SQ). To facilitate future research, we create a multimodal mini-dataset named CapQA, which includes 1k images of fine-grained activities, for visual instruction tuning and evaluation, our proposed SQ method leads to a 31.2% improvement in the hallucination score. Our extensive experiments on various benchmarks demonstrate SQ's remarkable capabilities in heuristic self-questioning, zero-shot visual reasoning and hallucination mitigation. Our model and code will be publicly available.

cs.CV↗

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal large language models (MLLMs) have made significant strides in open-world understanding. However, the challenge of continuous and simultaneous streaming perception, memory, and reasoning remains largely unexplored. Current MLLMs are constrained by their sequence-to-sequence architecture, which limits their ability to process inputs and generate responses simultaneously, akin to being unable to think while perceiving. Furthermore, relying on long contexts to store historical data is impractical for long-term interactions, as retaining all information becomes costly and inefficient. Therefore, rather than relying on a single foundation model to perform all functions, this project draws inspiration from the concept of the Specialized Generalist AI and introduces disentangled streaming perception, reasoning, and memory mechanisms, enabling real-time interaction with streaming video and audio input. The proposed framework InternLM-XComposer2.5-OmniLive (IXC2.5-OL) consists of three key modules: (1) Streaming Perception Module: Processes multimodal information in real-time, storing key details in memory and triggering reasoning in response to user queries. (2) Multi-modal Long Memory Module: Integrates short-term and long-term memory, compressing short-term memories into long-term ones for efficient retrieval and improved accuracy. (3) Reasoning Module: Responds to queries and executes reasoning tasks, coordinating with the perception and memory modules. This project simulates human-like cognition, enabling multimodal large language models to provide continuous and adaptive service over time.

cs.CV↗

Simplex tensor network renormalization group for boundary theory of 3+1D symTFT

Following the construction in arXiv:2210.12127, we develop a symmetry-preserving renormalization group (RG) flow for 3D symmetric theories. These theories are expressed as boundary conditions of a symTFT, which in our case is a 3+1D Dijkgraaf-Witten topological theory in the bulk. The boundary is geometrically organized into tetrahedra and represented as a tensor network, which we refer to as the "simplex tensor network" state. Each simplex tensor is assigned indices corresponding to its vertices, edges, and faces. We propose a numerical algorithm to implement RG flows for these boundary conditions, and explicitly demonstrate its application to a $\mathbb{Z}_2$ symmetric theory. By linearly interpolating between three topological fixed-point boundaries, we map the phase transitions characterized by local and non-local order parameters, which respectively detects the breaking of a 0-form and a 2-form symmetry. This formalism is readily extendable to other discrete symmetry groups and, in principle, can be generalized to describe 3D symmetric topological orders.

cond-mat.str-el↗

LHAASO detection of very-high-energy gamma-ray emission surrounding PSR J0248+6021

We report the detection of an extended very-high-energy (VHE) gamma-ray source coincident with the location of middle-aged (62.4~\rm kyr) pulsar PSR J0248+6021, by using the LHAASO-WCDA data of live 796 days and LHAASO-KM2A data of live 1216 days. A significant excess of \gray induced showers is observed both by WCDA in energy bands of 1-25~\rm TeV and KM2A in energy bands of $>$ 25~\rm TeV with 7.3 $σ$ and 13.5 $σ$, respectively. The best-fit position derived through WCDA data is R.A. = 42.06$^\circ \pm$ 0.12$^\circ$ and Dec. = 60.24$^\circ \pm $ 0.13$^\circ$ with an extension of 0.69$^\circ\pm$0.15$^\circ$ and that of the KM2A data is R.A.= 42.29$^\circ \pm $ 0.13$^\circ$ and Dec. = 60.38$^\circ \pm$ 0.07$^\circ$ with an extension of 0.37$^\circ\pm$0.07$^\circ$. No clear extended multiwavelength counterpart of this LHAASO source has been found from the radio band to the GeV band. The most plausible explanation of the VHE \gray emission is the inverse Compton process of highly relativistic electrons and positrons injected by the pulsar. These electrons/positrons are hypothesized to be either confined within the pulsar wind nebula or to have already escaped into the interstellar medium, forming a pulsar halo.

astro-ph.HE↗

Order-six CHMs containing exactly three distinct elements

Complex Hadamard matrices (CHMs) are intimately related to the number of distinct matrix elements. We investigate CHMs containing exactly three distinct elements, which is also the least number of distinct elements. In this paper, we show that such CHMs can only be complex equivalent to two kind of matrices, one is $H_2$-reducible and the other is the Tao matrix. Using our result one can further narrow the range of MUB trio (a set of four MUBs in $\mathbb{C}^6$ consists of an MUB trio and the identity) since we find that the two CHMs neither belong to MUB trios. Our results may lead to the more complete classification of $6\times 6$ CHMs whose elements in the first row are all 1.

quant-ph↗