SearcharxivSearch

arXiv subjects

Seunghwan Ji

Publications and source records attributed to Seunghwan Ji.

2 recordsLinked to original sources

CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model

The integration of Retrieval-Augmented Generation (RAG) with Multimodal Large Language Models (MLLMs) has revolutionized information retrieval and expanded the practical applications of AI. However, current systems struggle in accurately interpreting user intent, employing diverse retrieval strategies, and effectively filtering unintended or inappropriate responses, limiting their effectiveness. This paper introduces Contextual Understanding and Enhanced Search with MLLM (CUE-M), a novel multimodal search framework that addresses these challenges through a multi-stage pipeline comprising image context enrichment, intent refinement, contextual query generation, external API integration, and relevance-based filtering. CUE-M incorporates a robust filtering pipeline combining image-based, text-based, and multimodal classifiers, dynamically adapting to instance- and category-specific concern defined by organizational policies. Extensive experiments on real-word datasets and public benchmarks on knowledge-based VQA and safety demonstrated that CUE-M outperforms baselines and establishes new state-of-the-art results, advancing the capabilities of multimodal retrieval systems.

cs.CL

Fixed-complexity vector perturbation with Block diagonalization for MU-MIMO systems

Block diagonalization (BD) is an attractive technique that transforms the multi-user multiple-input multiple-output (MU-MIMO) channel into parallel single-user MIMO (SU-MIMO) channels with zero inter-user interference (IUI). In this paper, we combine the BD technique with two deterministic vector perturbation (VP) algorithms that reduce the transmit power in MU-MIMO systems with linear precoding. These techniques are the fixed-complexity sphere encoder (FSE) and the QR-decomposition with M-algorithm encoder (QRDM-E). In contrast to the conventional BD VP technique, which is based on the sphere encoder (SE), the proposed techniques have fixed complexity and a tradeoff between performance and complexity can be achieved by controlling the size of the set of candidates for the perturbation vector. Simulation results and analysis demonstrate the properness of the proposed techniques for the next generation mobile communications systems which are latency and computational complexity limited. In MU-MIMO system with 4 users each equipped with 2 receive antennas, simulation results show that the proposed BD-FSE and BD-QRDM-E outperforms the conventional BD-THP (Tomlinson Harashima precoding) by 5.5 and 7.4dB, respectively, at a target BER of 10^{-4}.

cs.IT