SearcharxivSearch

arXiv subjects

Shota Sato

Publications and source records attributed to Shota Sato.

4 recordsLinked to original sources

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains. We analyze this mismatch from the perspective of the decision structure in zero-shot classification, i.e. selecting the most similar class-text prototype for an input image. Zero-shot accuracy depends not only on average image--text alignment, but also on class-wise decision margins. Using Linear correction as an analytically tractable case, we show that modality gap correction can alter the relative decision structure among classes and cause predictions to concentrate on a small subset of classes. We refer to this output-space failure mode as prediction-level hubness. Furthermore, experiments across multiple datasets show that accuracy degradation under gap correction is consistently associated with increased prediction concentration, both for Linear correction and for learning-based correction methods. This provides a systematic explanation of why modality gap reduction does not consistently improve CLIP zero-shot accuracy from the perspective of downstream decision structure. Our results suggest that gap correction should be evaluated not only by average alignment, but also by its impact on downstream prediction structure.

cs.CL

M-IFEval: Multilingual Instruction-Following Evaluation

Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Following Evaluation (IFEval) benchmark from the literature does this using objective criteria, offering a measure of LLM performance without subjective AI or human judgement. However, it only includes English instructions, limiting its ability to assess LLMs in other languages. We propose the Multilingual Instruction Following Evaluation (M-IFEval) benchmark, expanding the evaluation to French, Japanese, and Spanish, with both general and language-specific instructions. Applying this benchmark to 8 state-of-the-art LLMs, we find that benchmark performance across languages and instruction types can vary widely, underscoring the importance of a multilingual benchmark for evaluating LLMs in a diverse cultural context.

cs.CL

MASTER OT J030227.28+191754.5: an unprecedentedly energetic dwarf nova outburst

We present a detailed study of the MASTER OT J030227.28+191754.5 outburst in 2021-2022, reaching an amplitude of 10.2 mag and a duration of 60 d. The detections of (1) the double-peaked optical emission lines, and (2) the early and ordinary superhumps, established that MASTER OT J030227.28+191754.5 is an extremely energetic WZ Sge-type dwarf nova (DN). Based on the superhump observations, we obtained its orbital period and mass ratio as 0.05986(1) d and 0.063(1), respectively. These are within a typical range of low-mass-ratio DNe. According to the binary parameters derived based on the thermal-tidal instability model, our analyses showed that (1) the standard disk model requires an accretion rate $\simeq$ 10$^{20}$ g s$^{-1}$ to explain its peak optical luminosity and (2) large mass was stored in the disk at the outburst onset. These cannot be explained solely by the impact of its massive ($\gtrsim$ 1.15 M$_\odot$) primary white dwarf implied by Kimura et al. (2023). Instead, we propose that the probable origin of this enormously energetic DN outburst is the even lower quiescence viscosity than other WZ Sge-type DNe. This discussion is qualitatively valid for most possible binary parameter spaces unless the inclination is low ($\lesssim 40^\circ$) enough for the disk to be bright explaining the outburst amplitude. Such low inclinations, however, would not allow detectable amplitude of early superhumps in the current thermal-tidal instability model. The optical spectra at outburst maximum showed the strong emission lines of Balmer, He I, and He II series whose core is narrower than $\sim 800$ km s$^{-1}$. Considering its binary parameters, a Keplerian disk cannot explain this narrow component, but the presumable origin is disk winds.

astro-ph.SR

Two double-angle formulas of generalized trigonometric functions

With respect to generalized trigonometric functions, since the discovery of double-angle formula for a special case by Edmunds, Gurka and Lang in 2012, no double-angle formulas have been found. In this paper, we will establish new double-angle formulas of generalized trigonometric functions in two special cases.

math.CA