SearcharxivSearch

arXiv subjects

Yuji Nozawa

Publications and source records attributed to Yuji Nozawa.

9 recordsLinked to original sources

Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

Recent work has shown that "Vision-Free'' approaches (representing images as text) can be effective for standard image retrieval tasks. However, it remains unclear whether this paradigm can effectively handle a more complex, multimodal task, Composed Image Retrieval (CIR), due to the inherent information loss in textual descriptions. In this paper, we introduce a Vision-Free CIR framework that addresses this challenge through two key techniques: (1) Attribute-Augmented Hybrid Scoring, which compensates for lost visual details via explicit attribute matching, and (2) LLM-Based Reranking, which verifies semantic consistency of top candidates. Experiments on the open-domain CIRR dataset show that our approach outperforms existing Zero-shot CIR methods (44.04% R@1, +8.79%). On FashionIQ, our results highlight the trade-off between semantic reasoning and fine-grained visual matching. Ablation studies reveal that both attribute-augmented scoring and LLM-Based Reranking consistently improve performance.

cs.CV

CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-historyconsistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query ateach turn progressively approaches the target image. Data are generated via a CIReVLbased retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multiturn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, highquality benchmark to facilitate future research on multi-turn CIR. The dataset is publicly available at https://huggingface.co/datasets/tk1441/CIRCLED, and the code at https://github.com/mti-lab/circled.

cs.CV

Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval

The goal of this paper is to enhance pretrained Vision Transformer (ViT) models for focus-oriented image retrieval with visual prompting. In real-world image retrieval scenarios, both query and database images often exhibit complexity, with multiple objects and intricate backgrounds. Users often want to retrieve images with specific object, which we define as the Focus-Oriented Image Retrieval (FOIR) task. While a standard image encoder can be employed to extract image features for similarity matching, it may not perform optimally in the multi-object-based FOIR task. This is because each image is represented by a single global feature vector. To overcome this, a prompt-based image retrieval solution is required. We propose an approach called Prompt-guided attention Head Selection (PHS) to leverage the head-wise potential of the multi-head attention mechanism in ViT in a promptable manner. PHS selects specific attention heads by matching their attention maps with user's visual prompts, such as a point, box, or segmentation. This empowers the model to focus on specific object of interest while preserving the surrounding visual context. Notably, PHS does not necessitate model re-training and avoids any image alteration. Experimental results show that PHS substantially improves performance on multiple datasets, offering a practical and training-free solution to enhance model performance in the FOIR task.

cs.CV

Improving Image Clustering with Artifacts Attenuation via Inference-Time Attention Engineering

The goal of this paper is to improve the performance of pretrained Vision Transformer (ViT) models, particularly DINOv2, in image clustering task without requiring re-training or fine-tuning. As model size increases, high-norm artifacts anomaly appears in the patches of multi-head attention. We observe that this anomaly leads to reduced accuracy in zero-shot image clustering. These artifacts are characterized by disproportionately large values in the attention map compared to other patch tokens. To address these artifacts, we propose an approach called Inference-Time Attention Engineering (ITAE), which manipulates attention function during inference. Specifically, we identify the artifacts by investigating one of the Query-Key-Value (QKV) patches in the multi-head attention and attenuate their corresponding attention values inside the pretrained models. ITAE shows improved clustering accuracy on multiple datasets by exhibiting more expressive features in latent space. Our findings highlight the potential of ITAE as a practical solution for reducing artifacts in pretrained ViT models and improving model performance in clustering tasks without the need for re-training or fine-tuning.

cs.CV

Revisiting Relevance Feedback for CLIP-based Interactive Image Retrieval

Many image retrieval studies use metric learning to train an image encoder. However, metric learning cannot handle differences in users' preferences, and requires data to train an image encoder. To overcome these limitations, we revisit relevance feedback, a classic technique for interactive retrieval systems, and propose an interactive CLIP-based image retrieval system with relevance feedback. Our retrieval system first executes the retrieval, collects each user's unique preferences through binary feedback, and returns images the user prefers. Even when users have various preferences, our retrieval system learns each user's preference through the feedback and adapts to the preference. Moreover, our retrieval system leverages CLIP's zero-shot transferability and achieves high accuracy without training. We empirically show that our retrieval system competes well with state-of-the-art metric learning in category-based image retrieval, despite not training image encoders specifically for each dataset. Furthermore, we set up two additional experimental settings where users have various preferences: one-label-based image retrieval and conditioned image retrieval. In both cases, our retrieval system effectively adapts to each user's preferences, resulting in improved accuracy compared to image retrieval without feedback. Overall, our work highlights the potential benefits of integrating CLIP with classic relevance feedback techniques to enhance image retrieval.

cs.CV

Generalized hydrodynamics study of the one-dimensional Hubbard model: Stationary clogging and proportionality of spin, charge, and energy currents

In our previous work [Y. Nozawa and H. Tsunetsugu, Phys. Rev. B 101, 035121 (2020)], we studied quench dynamics in the one-dimensional Hubbard model based on the generalized hydrodynamics theory for a partitioning protocol and showed the presence of a clogging phenomenon. Clogging is a phenomenon that vanishing charge current coexists with nonzero energy current, and was found when the protocol uses the initial condition that the left half of the system is prepared to be half filling at high temperatures with the right half being empty. Clogging occurs at all the sites in the left half and lasts for a time proportional to its distance from the connection point. In this paper, we use various different initial conditions and discuss two issues. The first issue is the possibility of clogging in a stationary state. When the electron density in the right half is initially set nonzero, we found that the left half-filled part expands for various sets of parameters in the initial condition. This means that the clogging phenomenon occurs at all the sites in the long-time stationary state, and we also discuss its origin. In addition, stationary clogging is accompanied by a back current, namely, particle density current flows towards the high-density region. We also found spin clogging occurs for some initial conditions, i.e., the vanishing spin current coexists with nonzero energy current. The second issue is the proportionality of spin and charge currents. We have found two spatio-temporal regions where the current ratio is fixed to a nonzero constant. We numerically studied how the current ratio depends on various initial conditions. We also studied the ratio of charge and energy currents.

cond-mat.stat-mech

Explicit Construction of Local Conserved Quantities in the XYZ Spin-1/2 Chain

We present a rigorous explicit expression for an extensive number of local conserved quantities in the XYZ spin-1/2 chain with general coupling constants. All the coefficients of operators in each local conserved quantity are calculated. We also confirm that our result can be applied to the case of the XXZ chain with a magnetic field in the z-axis direction.

cond-mat.stat-mech

Noncommutative generalized Gibbs ensemble in isolated integrable quantum systems

The generalized Gibbs ensemble (GGE), which involves multiple conserved quantities other than the Hamiltonian, has served as the statistical-mechanical description of the long-time behavior for several isolated integrable quantum systems. The GGE may involve a noncommutative set of conserved quantities in view of the maximum entropy principle, and show that the GGE thus generalized (noncommutative GGE, NCGGE) gives a more qualitatively accurate description of the long-time behaviors than that of the conventional GGE. Providing a clear understanding of why the (NC)GGE well describes the long-time behaviors, we construct, for noninteracting models, the exact NCGGE that describes the long-time behaviors without an error even at finite system size. It is noteworthy that the NCGGE involves nonlocal conserved quantities, which can be necessary for describing long-time behaviors of local observables. We also give some extensions of the NCGGE and demonstrate how accurately they describe the long-time behaviors of few-body observables.

cond-mat.stat-mech

Generalized Hydrodynamic approach to charge and energy currents in the one-dimensional Hubbard model

We have studied nonequilibrium dynamics of the one-dimensional Hubbard model using the generalized hydrodynamic theory. We mainly investigated the spatio-temporal profile of charge density, energy density and their currents using the partitioning protocol; the initial state consists of two semi-infinite different thermal equilibrium states joined at the origin. In this protocol, there appears around the origin a transient region where currents flow. We examined how density and current profiles depend on initial conditions and have found a clogged region where charge current is zero but nonvanishing energy current flows. This region appears when one of the initial states has half-filled electron density. We have proved analytically the existence of the clogged region in the infinite temperature case of the half-filled initial state. The existence is confirmed also for finite temperatures by numerical calculations. A similar analytical proof is also given for a clogged region of spin current when magnetic field is applied to one and only one of the two initial states. A universal proportionality of charge and spin currents is also proved for a special region, for general initial conditions of electron density and magnetic field. Except for the clogged region, charge and energy densities are in good proportion to each other, and their ratio depends on the initial conditions. This proportionality in nonequilibrium dynamics is reminiscent of Wiedemann-Franz law in thermal equilibrium. The long-time stationary values of charge and energy currents were also studied with varying initial conditions. We have compared the results with the values of non-interacting systems and discussed the effects of electron correlations. We have found that the temperature dependence of the ratio of these stationary currents is strongly suppressed by electron correlations and even reversed.

cond-mat.stat-mech