Searcharxiv⌕ Search

arXiv subjects

Tomoyuki Okuno

Publications and source records attributed to Tomoyuki Okuno.

At least 19 recordsLinked to original sources

Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When applying AFR to structured pruning, three major problems arise: distribution mismatch between heterogeneous pruning scores, loss of sign information indicating optimization direction consistency, and influence of outliers. To address these issues, we propose a unified approach combining power transformation for nonlinear distribution alignment, sign-preserving score aggregation, and percentile-based outlier removal. Experiments on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B demonstrate that our method maintains accuracy comparable to unstructured pruning while achieving practical inference speedup through structured pruning.

cs.CL↗

MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM

Dynamic runtime latency and memory constraints necessitate flexible large language model (LLM) deployment, where an LLM can be inferred with various quantization precisions based on available computational resources. Recent work on such any-precision quantization either relies on hardware-inefficient vector quantization or induces additional scaling factors when switching between bit-widths. Meanwhile, existing post-training quantization (PTQ) methods calibrated for a fixed low precision show poor generalizability under runtime precision change. In this work, we attribute the source of poor generalization across bit-widths to a precision-dependent \textit{outlier migration} phenomenon where the distribution of PTQ-sensitive tokens changes across precisions. Motivated by this observation, we propose \texttt{MoBiQuant}, a novel any-precision Mixture-of-Bits quantization framework that adjusts weight precision for flexible LLM inference based on token sensitivity. Specifically, we propose a many-in-one recursive residual quantization that can iteratively reconstruct higher-precision weights at runtime and mitigates \textit{outlier migration} with a token-aware router to dynamically select the optimal inference precision of each token.Extensive experiments show that \texttt{MoBiQuant} matches or surpasses frontier single-precision PTQ while exhibiting strong elasticity, achieving significant memory savings and throughput gains of up to $1.34\times$ over state-of-the-art any-precision methods.

cs.LG↗

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

Spatial intelligence in vision-language models (VLMs) attracts research interest with the practical demand to reason in the 3D world.Despite promising results, most existing methods follow the conventional 2D pipeline in VLMs and use pixel-aligned representations for the vision modality. However, correspondence-based models with implicit 3D scene understanding often fail to achieve spatial consistency, and representation-based models with 3D geometric priors lack efficiency in vision sequence serialization. To address this, we propose a Proxy3D method with compact yet comprehensive 3D proxy representations for the vision modality. Given only video frames as input, we employ semantic and geometric encoders to extract scene features and then perform their semantic-aware clustering to obtain a set of proxies in the 3D space. For representation alignment, we further curate the SpaceSpan dataset and apply multi-stage training to adopt the proposed 3D proxy representations with the VLM. When using shorter sequences for vision information, our method achieves competitive or state-of-the-art performance in 3D visual question answering, visual grounding and general spatial intelligence benchmarks.

cs.CV↗

ODE$_t$(ODE$_l$): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling

Continuous normalizing flows (CNFs) and diffusion models (DMs) generate high-quality data from a noise distribution. However, their sampling process demands multiple iterations to solve an ordinary differential equation (ODE) with high computational complexity. State-of-the-art methods focus on reducing the number of discrete time steps during sampling to improve efficiency. In this work, we explore a complementary direction in which the quality-complexity tradeoff can also be controlled in terms of the neural network length. We achieve this by rewiring the blocks in the transformer-based architecture to solve an inner discretized ODE w.r.t. its depth. Then, we apply a length consistency term during flow matching training, and as a result, the sampling can be performed with an arbitrary number of time steps and transformer blocks. Unlike others, our ODE$_t$(ODE$_l$) approach is solver-agnostic in time dimension and reduces both latency and, importantly, memory usage. CelebA-HQ and ImageNet generation experiments show a latency reduction of up to $2\times$ in the most efficient sampling mode, and FID improvement of up to $2.8$ points for high-quality sampling when applied to prior methods. We open-source our code and checkpoints at github.com/gudovskiy/odelt.

cs.LG↗

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens using certain training data. Differently, we propose a text-guided training-free token optimization mechanism dubbed SparseVLM that eliminates the need of extra parameters or fine-tuning costs. Given that visual tokens complement text tokens in VLM's linguistic reasoning, we select relevant text tokens to rate the significance of visual tokens using self-attention matrices and, then, prune visual tokens using the proposed strategy to maximize sparsity while retaining information. In particular, we introduce a rank-based strategy to adaptively determine the sparsification ratio for each layer, alongside a token recycling method that compresses pruned tokens into more compact representations. Experimental results show that SparseVLM increases the efficiency of various VLMs in a number of image and video understanding tasks. For example, LLaVA when equipped with SparseVLM achieves 54% reduction in FLOPs, 37% decrease in CUDA latency while maintaining 97% of its original accuracy. Our code is available at https://github.com/Gumpest/SparseVLMs.

cs.CV↗

DFM: Interpolant-free Dual Flow Matching

Continuous normalizing flows (CNFs) can model data distributions with expressive infinite-length architectures. But this modeling involves computationally expensive process of solving an ordinary differential equation (ODE) during maximum likelihood training. Recently proposed flow matching (FM) framework allows to substantially simplify the training phase using a regression objective with the interpolated forward vector field. In this paper, we propose an interpolant-free dual flow matching (DFM) approach without explicit assumptions about the modeled vector field. DFM optimizes the forward and, additionally, a reverse vector field model using a novel objective that facilitates bijectivity of the forward and reverse transformations. Our experiments with the SMAP unsupervised anomaly detection show advantages of DFM when compared to the CNF trained with either maximum likelihood or FM objectives with the state-of-the-art performance metrics.

cs.LG↗

Fisher-aware Quantization for DETR Detectors with Critical-category Objectives

The impact of quantization on the overall performance of deep learning models is a well-studied problem. However, understanding and mitigating its effects on a more fine-grained level is still lacking, especially for harder tasks such as object detection with both classification and regression objectives. This work defines the performance for a subset of task-critical categories, i.e. the critical-category performance, as a crucial yet largely overlooked fine-grained objective for detection tasks. We analyze the impact of quantization at the category-level granularity, and propose methods to improve performance for the critical categories. Specifically, we find that certain critical categories have a higher sensitivity to quantization, and are prone to overfitting after quantization-aware training (QAT). To explain this, we provide theoretical and empirical links between their performance gaps and the corresponding loss landscapes with the Fisher information framework. Using this evidence, we apply a Fisher-aware mixed-precision quantization scheme, and a Fisher-trace regularization for the QAT on the critical-category loss landscape. The proposed methods improve critical-category metrics of the quantized transformer-based DETR detectors. They are even more significant in case of larger models and higher number of classes where the overfitting becomes more severe. For example, our methods lead to 10.4% and 14.5% mAP gains for, correspondingly, 4-bit DETR-R50 and Deformable DETR on the most impacted critical classes in the COCO Panoptic dataset.

cs.CV↗

ContextFlow++: Generalist-Specialist Flow-based Generative Models with Mixed-Variable Context Encoding

Normalizing flow-based generative models have been widely used in applications where the exact density estimation is of major importance. Recent research proposes numerous methods to improve their expressivity. However, conditioning on a context is largely overlooked area in the bijective flow research. Conventional conditioning with the vector concatenation is limited to only a few flow types. More importantly, this approach cannot support a practical setup where a set of context-conditioned (specialist) models are trained with the fixed pretrained general-knowledge (generalist) model. We propose ContextFlow++ approach to overcome these limitations using an additive conditioning with explicit generalist-specialist knowledge decoupling. Furthermore, we support discrete contexts by the proposed mixed-variable architecture with context encoders. Particularly, our context encoder for discrete variables is a surjective flow from which the context-conditioned continuous variables are sampled. Our experiments on rotated MNIST-R, corrupted CIFAR-10C, real-world ATM predictive maintenance and SMAP unsupervised anomaly detection benchmarks show that the proposed ContextFlow++ offers faster stable training and achieves higher performance metrics. Our code is publicly available at https://github.com/gudovskiy/contextflow.

cs.LG↗

Split-Ensemble: Efficient OOD-aware Ensemble via Task and Model Splitting

Uncertainty estimation is crucial for machine learning models to detect out-of-distribution (OOD) inputs. However, the conventional discriminative deep learning classifiers produce uncalibrated closed-set predictions for OOD data. A more robust classifiers with the uncertainty estimation typically require a potentially unavailable OOD dataset for outlier exposure training, or a considerable amount of additional memory and compute to build ensemble models. In this work, we improve on uncertainty estimation without extra OOD data or additional inference costs using an alternative Split-Ensemble method. Specifically, we propose a novel subtask-splitting ensemble training objective, where a common multiclass classification task is split into several complementary subtasks. Then, each subtask's training data can be considered as OOD to the other subtasks. Diverse submodels can therefore be trained on each subtask with OOD-aware objectives. The subtask-splitting objective enables us to share low-level features across submodels to avoid parameter and computational overheads. In particular, we build a tree-like Split-Ensemble architecture by performing iterative splitting and pruning from a shared backbone model, where each branch serves as a submodel corresponding to a subtask. This leads to improved accuracy and uncertainty estimation across submodels under a fixed ensemble computation budget. Empirical study with ResNet-18 backbone shows Split-Ensemble, without additional computation cost, improves accuracy over a single model by 0.8%, 1.8%, and 25.5% on CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively. OOD detection for the same backbone and in-distribution datasets surpasses a single model baseline by, correspondingly, 2.2%, 8.1%, and 29.6% mean AUROC.

cs.LG↗

VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness

Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data points results in diminished training efficiency. To address this drawback, we propose a novel approach, Vision-language Collaborative Active Finetuning (VeCAF). With the emerging availability of labels and natural language annotations of images through web-scale crawling or controlled generation, VeCAF makes use of these information to perform parametric data selection for PVM finetuning. VeCAF incorporates the finetuning objective to select significant data points that effectively guide the PVM towards faster convergence to meet the performance goal. This process is assisted by the inherent semantic richness of the text embedding space which we use to augment image features. Furthermore, the flexibility of text-domain augmentation allows VeCAF to handle out-of-distribution scenarios without external data. Extensive experiments show the leading performance and high computational efficiency of VeCAF that is superior to baselines in both in-distribution and out-of-distribution image classification tasks. On ImageNet, VeCAF uses up to 3.3x less training batches to reach the target performance compared to full finetuning, and achieves an accuracy improvement of 2.7% over the state-of-the-art active finetuning method with the same number of batches.

cs.CV↗

Efficient Deweather Mixture-of-Experts with Uncertainty-aware Feature-wise Linear Modulation

The Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conventional MoE architecture with parallel Feed Forward Network (FFN) experts leads to significant parameter and computational overheads that hinder its efficient deployment. In addition, the naive MoE linear router is suboptimal in assigning task-specific features to multiple experts which limits its further scalability. In this work, we propose an efficient MoE architecture with weight sharing across the experts. Inspired by the idea of linear feature modulation (FM), our architecture implicitly instantiates multiple experts via learnable activation modulations on a single shared expert block. The proposed Feature Modulated Expert (FME) serves as a building block for the novel Mixture-of-Feature-Modulation-Experts (MoFME) architecture, which can scale up the number of experts with low overhead. We further propose an Uncertainty-aware Router (UaR) to assign task-specific features to different FM modules with well-calibrated weights. This enables MoFME to effectively learn diverse expert functions for multiple tasks. The conducted experiments on the multi-deweather task show that our MoFME outperforms the baselines in the image restoration quality by 0.1-0.2 dB and achieves SOTA-compatible performance while saving more than 72% of parameters and 39% inference time over the conventional MoE counterpart. Experiments on the downstream segmentation and classification tasks further demonstrate the generalizability of MoFME to real open-world applications.

cs.CV↗

Concurrent Misclassification and Out-of-Distribution Detection for Semantic Segmentation via Energy-Based Normalizing Flow

Recent semantic segmentation models accurately classify test-time examples that are similar to a training dataset distribution. However, their discriminative closed-set approach is not robust in practical data setups with distributional shifts and out-of-distribution (OOD) classes. As a result, the predicted probabilities can be very imprecise when used as confidence scores at test time. To address this, we propose a generative model for concurrent in-distribution misclassification (IDM) and OOD detection that relies on a normalizing flow framework. The proposed flow-based detector with an energy-based inputs (FlowEneDet) can extend previously deployed segmentation models without their time-consuming retraining. Our FlowEneDet results in a low-complexity architecture with marginal increase in the memory footprint. FlowEneDet achieves promising results on Cityscapes, Cityscapes-C, FishyScapes and SegmentMeIfYouCan benchmarks in IDM/OOD detection when applied to pretrained DeepLabV3+ and SegFormer semantic segmentation models.

cs.CV↗

MTTrans: Cross-Domain Object Detection with Mean-Teacher Transformer

Recently, DEtection TRansformer (DETR), an end-to-end object detection pipeline, has achieved promising performance. However, it requires large-scale labeled data and suffers from domain shift, especially when no labeled data is available in the target domain. To solve this problem, we propose an end-to-end cross-domain detection Transformer based on the mean teacher framework, MTTrans, which can fully exploit unlabeled target domain data in object detection training and transfer knowledge between domains via pseudo labels. We further propose the comprehensive multi-level feature alignment to improve the pseudo labels generated by the mean teacher framework taking advantage of the cross-scale self-attention mechanism in Deformable DETR. Image and object features are aligned at the local, global, and instance levels with domain query-based feature alignment (DQFA), bi-level graph-based prototype alignment (BGPA), and token-wise image feature alignment (TIFA). On the other hand, the unlabeled target domain data pseudo-labeled and available for the object detection training by the mean teacher framework can lead to better feature extraction and alignment. Thus, the mean teacher framework and the comprehensive multi-level feature alignment can be optimized iteratively and mutually based on the architecture of Transformers. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance in three domain adaptation scenarios, especially the result of Sim10k to Cityscapes scenario is remarkably improved from 52.6 mAP to 57.9 mAP. Code will be released.

cs.CV↗

Rapid Deceleration of Blast Waves Witnessed in Tycho's Supernova Remnant

In spite of their importance as standard candles in cosmology and as major major sites of nucleosynthesis in the Universe, what kinds of progenitor systems lead to type Ia supernovae (SN) remains a subject of considerable debate in the literature. This is true even for the case of Tycho's SN exploded in 1572 although it has been deeply studied both observationally and theoretically. Analyzing X-ray data of Tycho's supernova remnant (SNR) obtained with Chandra in 2003, 2007, 2009, and 2015, we discover that the expansion before 2007 was substantially faster than radio measurements reported in the past decades and then rapidly decelerated during the last ~ 15 years. The result is well explained if the shock waves recently hit a wall of dense gas surrounding the SNR. Such a gas structure is in fact expected in the so-called single-degenerate scenario, in which the progenitor is a binary system consisting of a white dwarf and a stellar companion, whereas it is not generally predicted by a competing scenario, the double-degenerate scenario, which has a binary of two white dwarfs as the progenitor. Our result thus favors the former scenario. This work also demonstrates a novel technique to probe gas environments surrounding SNRs and thus disentangle the two progenitor scenarios for Type Ia SNe.

astro-ph.HE↗

Time Variability of Nonthermal X-ray Stripes in Tycho's Supernova Remnant with Chandra

Analyzing Chandra data of Tycho's supernova remnant (SNR) taken in 2000, 2003, 2007, 2009, and 2015, we search for time variable features of synchrotron X-rays in the southwestern part of the SNR, where stripe structures of hard X-ray emission were previous found. By comparing X-ray images obtained at each epoch, we discover a knot-like structure in the northernmost part of the stripe region became brighter particularly in 2015. We also find a bright filamentary structure gradually became fainter and narrower as it moved outward. Our spectral analysis reveal that not only the nonthermal X-ray flux but also the photon indices of the knot-like structure change from year to year. During the period from 2000 to 2015, the small knot shows brightening of $\sim 70\%$ and hardening of $ΔΓ\sim 0.45$. The time variability can be explained if the magnetic field is amplified to $\sim 100~\mathrm{μG}$ and/or if magnetic turbulence significantly changes with time.

astro-ph.HE↗

Measurement of Charge Cloud Size in X-ray SOI Pixel Sensors

We report on a measurement of the size of charge clouds produced by X-ray photons in X-ray SOI (Silicon-On-Insulator) pixel sensor named XRPIX. We carry out a beam scanning experiment of XRPIX using a monochromatic X-ray beam at 5.0 keV collimated to $\sim 10$ $μ$m with a 4-$μ$m$ϕ$ pinhole, and obtain the spatial distribution of single-pixel events at a sub-pixel scale. The standard deviation of charge clouds of 5.0 keV X-ray is estimated to be $σ_{\rm cloud} = 4.30 \pm 0.07$ $μ$m. Compared to the detector response simulation, the estimated charge cloud size is well explained by a combination of photoelectron range, thermal diffusion, and Coulomb repulsion. Moreover, by analyzing the fraction of multi-pixel events in various energies, we find that the energy dependence of the charge cloud size is also consistent with the simulation.

astro-ph.IM↗

Subpixel Response of SOI Pixel Sensor for X-ray Astronomy with Pinned Depleted Diode: First Result from Mesh Experiment

We have been developing a monolithic active pixel sensor, ``XRPIX``, for the Japan led future X-ray astronomy mission ``FORCE`` observing the X-ray sky in the energy band of 1-80 keV with angular resolution of better than 15``. XRPIX is an upper part of a stack of two sensors of an imager system onboard FORCE, and covers the X-ray energy band lower than 20 keV. The XRPIX device consists of a fully depleted high-resistivity silicon sensor layer for X-ray detection, a low resistivity silicon layer for CMOS readout circuit, and a buried oxide layer in between, which is fabricated with 0.2 $μ$ m CMOS silicon-on-insulator (SOI) technology. Each pixel has a trigger circuit with which we can achieve a 10 $μ$ s time resolution, a few orders of magnitude higher than that with X-ray astronomy CCDs. We recently introduced a new type of a device structure, a pinned depleted diode (PDD), in the XRPIX device, and succeeded in improving the spectral performance, especially in a readout mode using the trigger function. In this paper, we apply a mesh experiment to the XRPIX devices for the first time in order to study the spectral response of the PDD device at the subpixel resolution. We confirmed that the PDD structure solves the significant degradation of the charge collection efficiency at the pixel boundaries and in the region under the pixel circuits, which is found in the single SOI structure, the conventional type of the device structure. On the other hand, the spectral line profiles are skewed with low energy tails and the line peaks slightly shift near the pixel boundaries, which contribute to a degradation of the energy resolution.

astro-ph.IM↗

Performance of SOI Pixel Sensors Developed for X-ray Astronomy

We have been developing monolithic active pixel sensors for X-rays based on the silicon-on-insulator technology. Our device consists of a low-resistivity Si layer for readout CMOS electronics, a high-resistivity Si sensor layer, and a SiO$_2$ layer between them. This configuration allows us both high-speed readout circuits and a thick (on the order of $100~μ{\rm m}$) depletion layer in a monolithic device. Each pixel circuit contains a trigger output function, with which we can achieve a time resolution of $\lesssim 10~μ{\rm s}$. One of our key development items is improvement of the energy resolution. We recently fabricated a device named XRPIX6E, to which we introduced a pinned depleted diode (PDD) structure. The structure reduces the capacitance coupling between the sensing area in the sensor layer and the pixel circuit, which degrades the spectral performance. With XRPIX6E, we achieve an energy resolution of $\sim 150$~eV in full width at half maximum for 6.4-keV X-rays. In addition to the good energy resolution, a large imaging area is required for practical use. We developed and tested XRPIX5b, which has an imaging area size of $21.9~{\rm mm} \times 13.8~{\rm mm}$ and is the largest device that we ever fabricated. We successfully obtain X-ray data from almost all the $608 \times 384$ pixels with high uniformity.

astro-ph.IM↗