SearcharxivSearch

arXiv subjects

Han-Gyu Kim

Publications and source records attributed to Han-Gyu Kim.

14 recordsLinked to original sources

Who Spoke When in Multi-Conversation: Target Speaker Tagging Task and Benchmark

We present target speaker tagging (TST), a task that integrates speaker diarization, verification, and identification into a unified workflow for multi-speaker conversations. Given long recordings and pre-enrolled speakers, TST detects and labels speech segments of known speakers while rejecting unknown ones. Despite its practical importance, research has been limited by the absence of suitable evaluation resources. To address this, we introduce TST-Bench, a large-scale synthetic benchmark with over 150 enrolled speakers, 300 sessions of 20-60 minutes, and reference annotations with global speaker labels. We define an evaluation protocol encompassing diarization and full-pipeline scenarios. Experiments on both real and synthetic data show that TST poses challenges not captured by conventional benchmarks, and that dedicated system design yields significant gains over naive integration of existing solutions. The benchmark dataset and evaluation protocols are publicly released.

eess.AS

MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models

Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvBench) only partially capture the breadth and depth of conversational scenarios encountered by users. In this work, we introduce MultiVerse, a novel multi-turn conversation benchmark featuring 647 dialogues - each averaging four turns - derived from a diverse set of 12 popular VLM evaluation benchmarks. With 484 tasks and 484 interaction goals, MultiVerse covers a wide range of topics, from factual knowledge and perception to advanced reasoning tasks such as mathematics and coding. To facilitate robust assessment, we propose a checklist-based evaluation method that leverages GPT-4o as the automated evaluator, measuring performance across 37 key aspects, including perceptual accuracy, linguistic clarity, and factual correctness. We evaluate 18 VLMs on MultiVerse, revealing that even the strongest models (e.g., GPT-4o) achieve only a 50% success rate in complex multi-turn conversations, highlighting the dataset's challenging nature. Notably, we find that providing full dialogue context significantly enhances performance for smaller or weaker models, emphasizing the importance of in-context learning. We believe MultiVerse is a landscape of evaluating multi-turn interaction abilities for VLMs.

cs.CV

Intriguing Properties of Large Language and Vision Models

Recently, large language and vision models (LLVMs) have received significant attention and development efforts due to their remarkable generalization performance across a wide range of tasks requiring perception and cognitive abilities. A key factor behind their success is their simple architecture, which consists of a vision encoder, a projector, and a large language model (LLM). Despite their achievements in advanced reasoning tasks, their performance on fundamental perception-related tasks (e.g., MMVP) remains surprisingly low. This discrepancy raises the question of how LLVMs truly perceive images and exploit the advantages of the vision encoder. To address this, we systematically investigate this question regarding several aspects: permutation invariance, robustness, math reasoning, alignment preserving and importance, by evaluating the most common LLVM's families (i.e., LLaVA) across 10 evaluation benchmarks. Our extensive experiments reveal several intriguing properties of current LLVMs: (1) they internally process the image in a global manner, even when the order of visual patch sequences is randomly permuted; (2) they are sometimes able to solve math problems without fully perceiving detailed numerical information; (3) the cross-modal alignment is overfitted to complex reasoning tasks, thereby, causing them to lose some of the original perceptual capabilities of their vision encoder; (4) the representation space in the lower layers (<25%) plays a crucial role in determining performance and enhancing visual understanding. Lastly, based on the above observations, we suggest potential future directions for building better LLVMs and constructing more challenging evaluation benchmarks.

cs.CV

Multi-Scale Experimental Characterization for LS-DYNA MAT213 Modeling of Composite Structures under High Strain Rate

Aerospace structures often experience high strain rate events such as ballistic impact, crash, or crush. A material model has been developed that enhances the capability to simulate the dynamic response of composite materials under these loading conditions. The material model has been implemented into the commercially available transient dynamic finite element code LS-DYNA as MAT213. The model can simulate the nonlinear deformation, damage, and failure that takes place in a composite under dynamic loading conditions. The specific goal of this work is to characterize the MAT213 input for the representative material. The specific composite material being examined consists of T700G unidirectional carbon fibers and a low-melt PolyArylEtherKetone (LMPAEK) thermoplastic resin system. It is formally referred to as Toray TC1225 LMPAEK T700G. As the initial part of this work, this paper is focused on characterizing the material parameters for the MAT213 deformation model based on results obtained from multi-scale experimentation. The effort concentrated on characterizing the in-plane material response suitable for use with thin shell elements. For shell elements within MAT213, tabulated stress-strain results from tension and compression tests in the longitudinal and transverse directions and in-plane shear tests are required. Due to the difficulty of measuring small strains in the transverse direction, a multi-scale testing method was developed. Macro-scale testing is performed per the typical ASTM methods while micro-scale testing uses a microscope along with smaller coupon sizes to obtain the smaller strains in the transverse direction of each test. For both testing methods, a VIC-2D camera and software for digital image correlation analysis are used. Using the DIC combined with each test fixture, reliable stress and strain data are collected.

physics.app-ph

Experimental characterization of cohesive laws for mode-II interlaminar fracture in geometrically scaled composites using through-thickness deformation analysis

This work proposes an experimental framework to characterize a cohesive law for mode-II interlaminar fracture and demonstrates its implementation. For a size effect study, geometrically scaled end-notched flexure specimens were tested using microscopic and macroscopic digital image correlation (DIC) systems. The fracture energy was characterized using a compliance calibration method and Ba\v{z}ant's type-II size effect law for comparison. In the proposed experimental framework, the DIC data were post-processed using three steps: coordinate transformation, curve fitting, and through-thickness deformation analysis. Different magnitudes of separation values were measured from different sizes at fracture loads, implying size effect and partial development of cohesive laws. Modeling and simulations were intended to validate the proposed method and demonstrate the utilization of the experimental data. Additionally, challenges related to finding a single cohesive law for geometrically scaled specimens of a single material were exposed. A single cohesive law for the scaled specimens was developed and proposed as a material property of the specimen material. The fracture energy of the single law was smaller than the energy obtained from the size effect analysis, while the sizes of fracture process zones at fracture loads were smaller than the experimental measurements. However, the global fracture behaviors of the models showed good agreement with the experimental data of the mid-size specimen while showing reasonable agreement with the other sizes. Furthermore, the single law successfully captured local fracture behaviors by showing partial cohesive zone development at the fracture loads and matching the microscopic measurement of the separation values.

physics.app-ph

Experimental Investigation of the Structural Performance of Composite Structures Produced using Additive Manufacturing

This project is focused on investigating the structural performance of parts and structures produced using the latest additive manufacturing techniques. For additive manufacturing of test coupons, fused deposition modeling of fiber-reinforced polymer (RFP) and high-resolution low-force stereolithography (LFS) thermoset resin printing systems were employed. UV thermoset resin was used for LFS printing, while RFP printing adopted two different types of filaments: amorphous polycarbonate carbon fiber filaments and semi-crystalline Nylon 12 glass fiber filaments. For the experimental work, specimens were printed for tension, compression, and shear tests. Additionally, mode-II interlminar fracture in these specimens was explored. The elastic modulus and strength values of these specimens were compared with the data of oven-cured T700G/2510 composites. The experimental work herein will be extended to develop damage models for 3D-printed structural parts and structures for aerospace and space applications.

physics.app-ph

Experimental Characterization of Non-Associative Plasticity Flow Rule Coefficients for the LS-DYNA MAT213 Model

This project is focused on developing an experimental framework for characterizing non-associative plasticity flow rule coefficients through coupon-scale tests for the LS-DYNA MAT213 model. The main objective is to characterize these coefficients based on the multi-scale (i.e., both microscopic and macroscopic) full-field measurement of the evolution of strain and stress fields. This paper focuses on presenting the experimental work on characterizing the full-scale stress-strain curves of T700/LM-PAEK composites under tension, compression, and shear loads. The experimental data set was intended to build a deformation sub-model in the MAT213 model for the material. The strain data were collected using both microscopic and macroscopic digital image correlation techniques. The microscopic technique was particularly useful for fracture cases under small strains. A preliminary simulation result obtained from the MAT213 model is also presented in the paper. The experimental framework herein will be extended to characterize post-peak stress degradation in the composite material and to develop a damage sub-model for the material. This project will contribute to developing a simulation tool based on the MAT213 model for simulating the rate-dependent impact damages in composites under multi-axial loading.

physics.app-ph

Experimental multi-scale characterization of mode-II interlaminar fracture in geometrically scaled stitched and unstitched resin-infused composites

This work is focused on investigating the impact of out-of-plane stitches on enhancing mode-II interlaminar fracture toughness (or energy) and characterizing damage progression and crack arrestment in stitched resin-infused composites. For the experimental work, End-Notched Flexure (ENF) quasi-isotropic specimens were manufactured using +/-45 non-crimp carbon-fiber fabrics through a resin-infusion process. Both stitched and unstitched specimen sets were designed for comparison. For a size effect study, the ENF specimens were geometrically scaled with three scaling levels. Based on the load-displacement data (i.e., global analysis), the fracture energy of the specimen material was analyzed using the compliance calibration method and a size effect theory. The fracture energy values were compared between the stitched and unstitched cases to characterize the enhanced fracture toughness of stitched composites. For local analysis, two types of digital image correlation (DIC) systems were employed: microscopic and macroscopic (i.e., coupon-scale) DIC systems. By analyzing in-plane displacement through the thickness, separation development was characterized along predicted fracture process zones. The impact of out-of-plane stitches on separation propagation along fracture process zones was discussed based on the DIC analysis. This work will contribute to developing a high-fidelity damage model for stitched resin-infused composites in the form of a traction-separation for high-speed aircraft applications.

physics.app-ph

Group Generalized Mean Pooling for Vision Transformer

Vision Transformer (ViT) extracts the final representation from either class token or an average of all patch tokens, following the architecture of Transformer in Natural Language Processing (NLP) or Convolutional Neural Networks (CNNs) in computer vision. However, studies for the best way of aggregating the patch tokens are still limited to average pooling, while widely-used pooling strategies, such as max and GeM pooling, can be considered. Despite their effectiveness, the existing pooling strategies do not consider the architecture of ViT and the channel-wise difference in the activation maps, aggregating the crucial and trivial channels with the same importance. In this paper, we present Group Generalized Mean (GGeM) pooling as a simple yet powerful pooling strategy for ViT. GGeM divides the channels into groups and computes GeM pooling with a shared pooling parameter per group. As ViT groups the channels via a multi-head attention mechanism, grouping the channels by GGeM leads to lower head-wise dependence while amplifying important channels on the activation maps. Exploiting GGeM shows 0.1%p to 0.7%p performance boosts compared to the baselines and achieves state-of-the-art performance for ViT-Base and ViT-Large models in ImageNet-1K classification task. Moreover, GGeM outperforms the existing pooling strategies on image retrieval and multi-modal representation learning tasks, demonstrating the superiority of GGeM for a variety of tasks. GGeM is a simple algorithm in that only a few lines of code are necessary for implementation.

cs.CV

DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset

As sharing images in an instant message is a crucial factor, there has been active research on learning an image-text multi-modal dialogue models. However, training a well-generalized multi-modal dialogue model remains challenging due to the low quality and limited diversity of images per dialogue in existing multi-modal dialogue datasets. In this paper, we propose an automated pipeline to construct a multi-modal dialogue dataset, ensuring both dialogue quality and image diversity without requiring minimum human effort. In our pipeline, to guarantee the coherence between images and dialogue, we prompt GPT-4 to infer potential image-sharing moments - specifically, the utterance, speaker, rationale, and image description. Furthermore, we leverage CLIP similarity to maintain consistency between aligned multiple images to the utterance. Through this pipeline, we introduce DialogCC, a high-quality and diverse multi-modal dialogue dataset that surpasses existing datasets in terms of quality and diversity in human evaluation. Our comprehensive experiments highlight that when multi-modal dialogue models are trained using our dataset, their generalization performance on unseen dialogue datasets is significantly enhanced. We make our source code and dataset publicly available.

cs.CV

Learning with Memory-based Virtual Classes for Deep Metric Learning

The core of deep metric learning (DML) involves learning visual similarities in high-dimensional embedding space. One of the main challenges is to generalize from seen classes of training data to unseen classes of test data. Recent works have focused on exploiting past embeddings to increase the number of instances for the seen classes. Such methods achieve performance improvement via augmentation, while the strong focus on seen classes still remains. This can be undesirable for DML, where training and test data exhibit entirely different classes. In this work, we present a novel training strategy for DML called MemVir. Unlike previous works, MemVir memorizes both embedding features and class weights to utilize them as additional virtual classes. The exploitation of virtual classes not only utilizes augmented information for training but also alleviates a strong focus on seen classes for better generalization. Moreover, we embed the idea of curriculum learning by slowly adding virtual classes for a gradual increase in learning difficulty, which improves the learning stability as well as the final performance. MemVir can be easily applied to many existing loss functions without any modification. Extensive experimental results on famous benchmarks demonstrate the superiority of MemVir over state-of-the-art competitors. Code of MemVir is publicly available.

cs.CV

Back from the future: bidirectional CTC decoding using future information in speech recognition

In this paper, we propose a simple but effective method to decode the output of Connectionist Temporal Classifier (CTC) model using a bi-directional neural language model. The bidirectional language model uses the future as well as the past information in order to predict the next output in the sequence. The proposed method based on bi-directional beam search takes advantage of the CTC greedy decoding output to represent the noisy future information. Experiments on the Librispeechdataset demonstrate the superiority of our proposed method compared to baselines using unidirectional decoding. In particular, the boost inaccuracy is most apparent at the start of a sequence which is the most erroneous part for existing systems based on unidirectional decoding.

cs.CL

Proxy Synthesis: Learning with Synthetic Classes for Deep Metric Learning

One of the main purposes of deep metric learning is to construct an embedding space that has well-generalized embeddings on both seen (training) classes and unseen (test) classes. Most existing works have tried to achieve this using different types of metric objectives and hard sample mining strategies with given training data. However, learning with only the training data can be overfitted to the seen classes, leading to the lack of generalization capability on unseen classes. To address this problem, we propose a simple regularizer called Proxy Synthesis that exploits synthetic classes for stronger generalization in deep metric learning. The proposed method generates synthetic embeddings and proxies that work as synthetic classes, and they mimic unseen classes when computing proxy-based losses. Proxy Synthesis derives an embedding space considering class relations and smooth decision boundaries for robustness on unseen classes. Our method is applicable to any proxy-based losses, including softmax and its variants. Extensive experiments on four famous benchmarks in image retrieval tasks demonstrate that Proxy Synthesis significantly boosts the performance of proxy-based losses and achieves state-of-the-art performance.

cs.CV

Audio Source Separation Using a Deep Autoencoder

This paper proposes a novel framework for unsupervised audio source separation using a deep autoencoder. The characteristics of unknown source signals mixed in the mixed input is automatically by properly configured autoencoders implemented by a network with many layers, and separated by clustering the coefficient vectors in the code layer. By investigating the weight vectors to the final target, representation layer, the primitive components of the audio signals in the frequency domain are observed. By clustering the activation coefficients in the code layer, the previously unknown source signals are segregated. The original source sounds are then separated and reconstructed by using code vectors which belong to different clusters. The restored sounds are not perfect but yield promising results for the possibility in the success of many practical applications.

cs.SD