Searcharxiv⌕ Search

arXiv subjects

Takahiro Toizumi

Publications and source records attributed to Takahiro Toizumi.

18 recordsLinked to original sources

IA-CLAHE: Image-Adaptive Clip Limit Estimation for CLAHE

This paper proposes image-adaptive contrast limited adaptive histogram equalization (IA-CLAHE). Conventional CLAHE is widely used to boost the performance of various computer vision tasks and to improve visual quality for human perception in practical industrial applications. CLAHE applies contrast limited histogram equalization to each local region to enhance local contrast. However, CLAHE often leads to over-enhancement, because the contrast-limiting parameter clip limit is fixed regardless of the histogram distribution of each local region. Our IA-CLAHE addresses this limitation by adaptively estimating tile-wise clip limits from the input image. To achieve this, we train a lightweight clip limits estimator with a differentiable extension of CLAHE, enabling end-to-end optimization. Unlike prior learning-based CLAHE methods, IA-CLAHE does not require pre-searched ground-truth clip limits or task-specific datasets, because it learns to map input image histograms toward a domain-invariant uniform distribution, enabling zero-shot generalization across diverse conditions. Experimental results show that IA-CLAHE consistently improves recognition performance, while simultaneously enhancing visual quality for human perception, without requiring any task-specific training data.

cs.CV↗

CURVE: CLIP-Utilized Reinforcement Learning for Visual Image Enhancement via Simple Image Processing

Low-Light Image Enhancement (LLIE) is crucial for improving both human perception and computer vision tasks. This paper addresses two challenges in zero-reference LLIE: obtaining perceptually 'good' images using the Contrastive Language-Image Pre-Training (CLIP) model and maintaining computational efficiency for high-resolution images. We propose CLIP-Utilized Reinforcement learning-based Visual image Enhancement (CURVE). CURVE employs a simple image processing module which adjusts global image tone based on Bézier curve and estimates its processing parameters iteratively. The estimator is trained by reinforcement learning with rewards designed using CLIP text embeddings. Experiments on low-light and multi-exposure datasets demonstrate the performance of CURVE in terms of enhancement quality and processing speed compared to conventional methods.

cs.CV↗

Rethinking Image Histogram Matching for Image Classification

This paper rethinks image histogram matching (HM) and proposes a differentiable and parametric HM preprocessing for a downstream classifier. Convolutional neural networks have demonstrated remarkable achievements in classification tasks. However, they often exhibit degraded performance on low-contrast images captured under adverse weather conditions. To maintain classifier performance under low-contrast images, histogram equalization (HE) is commonly used. HE is a special case of HM using a uniform distribution as a target pixel value distribution. In this paper, we focus on the shape of the target pixel value distribution. Compared to a uniform distribution, a single, well-designed distribution could have potential to improve the performance of the downstream classifier across various adverse weather conditions. Based on this hypothesis, we propose a differentiable and parametric HM that optimizes the target distribution using the loss function of the downstream classifier. This method addresses pixel value imbalances by transforming input images with arbitrary distributions into a target distribution optimized for the classifier. Our HM is trained on only normal weather images using the classifier. Experimental results show that a classifier trained with our proposed HM outperforms conventional preprocessing methods under adverse weather conditions.

cs.CV↗

Target Driven Adaptive Loss For Infrared Small Target Detection

We propose a target driven adaptive (TDA) loss to enhance the performance of infrared small target detection (IRSTD). Prior works have used loss functions, such as binary cross-entropy loss and IoU loss, to train segmentation models for IRSTD. Minimizing these loss functions guides models to extract pixel-level features or global image context. However, they have two issues: improving detection performance for local regions around the targets and enhancing robustness to small scale and low local contrast. To address these issues, the proposed TDA loss introduces a patch-based mechanism, and an adaptive adjustment strategy to scale and local contrast. The proposed TDA loss leads the model to focus on local regions around the targets and pay particular attention to targets with smaller scales and lower local contrast. We evaluate the proposed method on three datasets for IRSTD. The results demonstrate that the proposed TDA loss achieves better detection performance than existing losses on these datasets.

cs.CV↗

Improving Low-Light Image Recognition Performance Based on Image-adaptive Learnable Module

In recent years, significant progress has been made in image recognition technology based on deep neural networks. However, improving recognition performance under low-light conditions remains a significant challenge. This study addresses the enhancement of recognition model performance in low-light conditions. We propose an image-adaptive learnable module which apply appropriate image processing on input images and a hyperparameter predictor to forecast optimal parameters used in the module. Our proposed approach allows for the enhancement of recognition performance under low-light conditions by easily integrating as a front-end filter without the need to retrain existing recognition models designed for low-light conditions. Through experiments, our proposed method demonstrates its contribution to enhancing image recognition performance under low-light conditions.

cs.CV↗

Recognition-Oriented Low-Light Image Enhancement based on Global and Pixelwise Optimization

In this paper, we propose a novel low-light image enhancement method aimed at improving the performance of recognition models. Despite recent advances in deep learning, the recognition of images under low-light conditions remains a challenge. Although existing low-light image enhancement methods have been developed to improve image visibility for human vision, they do not specifically focus on enhancing recognition model performance. Our proposed low-light image enhancement method consists of two key modules: the Global Enhance Module, which adjusts the overall brightness and color balance of the input image, and the Pixelwise Adjustment Module, which refines image features at the pixel level. These modules are trained to enhance input images to improve downstream recognition model performance effectively. Notably, the proposed method can be applied as a frontend filter to improve low-light recognition performance without requiring retraining of downstream recognition models. Experimental results demonstrate that our method improves the performance of pretrained recognition models under low-light conditions and its effectiveness.

cs.CV↗

ERUP-YOLO: Enhancing Object Detection Robustness for Adverse Weather Condition by Unified Image-Adaptive Processing

We propose an image-adaptive object detection method for adverse weather conditions such as fog and low-light. Our framework employs differentiable preprocessing filters to perform image enhancement suitable for later-stage object detections. Our framework introduces two differentiable filters: a Bézier curve-based pixel-wise (BPW) filter and a kernel-based local (KBL) filter. These filters unify the functions of classical image processing filters and improve performance of object detection. We also propose a domain-agnostic data augmentation strategy using the BPW filter. Our method does not require data-specific customization of the filter combinations, parameter ranges, and data augmentation. We evaluate our proposed approach, called Enhanced Robustness by Unified Image Processing (ERUP)-YOLO, by applying it to the YOLOv3 detector. Experiments on adverse weather datasets demonstrate that our proposed filters match or exceed the expressiveness of conventional methods and our ERUP-YOLO achieved superior performance in a wide range of adverse weather conditions, including fog and low-light conditions.

cs.CV↗

Adaptive Deep Iris Feature Extractor at Arbitrary Resolutions

This paper proposes a deep feature extractor for iris recognition at arbitrary resolutions. Resolution degradation reduces the recognition performance of deep learning models trained by high-resolution images. Using various-resolution images for training can improve the model's robustness while sacrificing recognition performance for high-resolution images. To achieve higher recognition performance at various resolutions, we propose a method of resolution-adaptive feature extraction with automatically switching networks. Our framework includes resolution expert modules specialized for different resolution degradations, including down-sampling and out-of-focus blurring. The framework automatically switches them depending on the degradation condition of an input image. Lower-resolution experts are trained by knowledge-distillation from the high-resolution expert in such a manner that both experts can extract common identity features. We applied our framework to three conventional neural network models. The experimental results show that our method enhances the recognition performance at low-resolution in the conventional methods and also maintains their performance at high-resolution.

cs.CV↗

Fast Eye Detector Using Siamese Network for NIR Partial Face Images

This paper proposes a fast eye detection method that is based on a Siamese network for near infrared (NIR) partial face images. NIR partial face images do not include the whole face of a subject since they are captured using iris recognition systems with the constraint of frame rate and resolution. The iris recognition systems such as the iris on the move (IOTM) system require fast and accurate eye detection as a pre-process. Our goal is to design eye detection with high speed, high discrimination performance between left and right eyes, and high positional accuracy of eye center. Our method adopts a Siamese network and coarse to fine position estimation with a fast lightweight CNN backbone. The network outputs features of images and the similarity map indicating coarse position of an eye. A regression on a portion of a feature with high similarity refines the coarse position of the eye to obtain the fine position with high accuracy. We demonstrate the effectiveness of the proposed method by comparing it with conventional methods, including SOTA, in terms of the positional accuracy, the discrimination performance, and the processing speed. Our method achieves superior performance in speed.

cs.CV↗

Segmentation-free Direct Iris Localization Networks

This paper proposes an efficient iris localization method without using iris segmentation and circle fitting. Conventional iris localization methods first extract iris regions by using semantic segmentation methods such as U-Net. Afterward, the inner and outer iris circles are localized using the traditional circle fitting algorithm. However, this approach requires high-resolution encoder-decoder networks for iris segmentation, so it causes computational costs to be high. In addition, traditional circle fitting tends to be sensitive to noise in input images and fitting parameters, causing the iris recognition performance to be poor. To solve these problems, we propose an iris localization network (ILN), that can directly localize pupil and iris circles with eyelid points from a low-resolution iris image. We also introduce a pupil refinement network (PRN) to improve the accuracy of pupil localization. Experimental results show that the combination of ILN and PRN works in 34.5 ms for one iris image on a CPU, and its localization performance outperforms conventional iris segmentation methods. In addition, generalized evaluation results show that the proposed method has higher robustness for datasets in different domain than other segmentation methods. Furthermore, we also confirm that the proposed ILN and PRN improve the iris recognition accuracy.

cs.CV↗

Rollable Latent Space for Azimuth Invariant SAR Target Recognition

This paper proposes rollable latent space (RLS) for an azimuth invariant synthetic aperture radar (SAR) target recognition. Scarce labeled data and limited viewing direction are critical issues in SAR target recognition.The RLS is a designed space in which rolling of latent features corresponds to 3D rotation of an object. Thus latent features of an arbitrary view can be inferred using those of different views. This characteristic further enables us to augment data from limited viewing in RLS. RLS-based classifiers with and without data augmentation and a conventional classifier trained with target front shots are evaluated over untrained target back shots. Results show that the RLS-based classifier with augmentation improves an accuracy by 30% compared to the conventional classifier.

cs.CV↗

MAXI observations of GRBs

Monitor of all-sky image (MAXI) Gas Slit Camera (GSC) detects gamma-ray bursts (GRBs) including the bursts with soft spectra, such as X-ray flashes (XRFs). MAXI/GSC is sensitive to the energy range from 2 to 30 keV. This energy range is lower than other currently operating instruments which is capable of detecting GRBs. Since the beginning of the MAXI operation on August 15, 2009, GSC observed 35 GRBs up to the middle of 2013. One third of them are also observed by other satellites. The rest of them show a trend to have soft spectra and low fluxes. Because of the contribution of those XRFs, the MAXI GRB rate is about three times higher than those expected from the BATSE log N - log P distribution. When we compare it to the observational results of the Wide-field X-ray Monitor on the High Energy Transient Explorer 2, which covers the the same energy range to that of MAXI/GSC, we find a possibility that many of MAXI bursts are XRFs with Epeak lower than 20 keV. We discuss the source of soft GRBs observed only by MAXI. The MAXI log N - log S distribution suggests that the MAXI XRFs distribute in closer distance than hard GRBs. Since the distributions of the hardness of galactic stellar flares and X-ray bursts overlap with those of MAXI GRBs, we discuss a possibility of a confusion of those galactic transients with the MAXI GRB samples.

astro-ph.HE↗

Spectral Evolution of a New X-ray Transient MAXI J0556-332 Observed by MAXI, Swift, and RXTE

We report on the spectral evolution of a new X-ray transient, MAXI J0556-332, observed by MAXI, Swift, and RXTE. The source was discovered on 2011 January 11 (MJD=55572) by MAXI Gas Slit Camera all-sky survey at (l,b)=(238.9deg, -25.2deg), relatively away from the Galactic plane. Swift/XRT follow-up observations identified it with a previously uncatalogued bright X-ray source and led to optical identification. For more than one year since its appearance, MAXI J0556-332 has been X-ray active, with a 2-10 keV intensity above 30 mCrab. The MAXI/GSC data revealed rapid X-ray brightening in the first five days, and a hard-to-soft transition in the meantime. For the following ~ 70 days, the 0.5-30 keV spectra, obtained by the Swift/XRT and the RXTE/PCA on an almost daily basis, show a gradual hardening, with large flux variability. These spectra are approximated by a cutoff power-law with a photon index of 0.4-1 and a high-energy exponential cutoff at 1.5-5 keV, throughout the initial 10 months where the spectral evolution is mainly represented by a change of the cutoff energy. To be more physical, the spectra are consistently explained by thermal emission from an accretion disk plus a Comptonized emission from a boundary layer around a neutron star. This supports the source identification as a neutron-star X-ray binary. The obtained spectral parameters agree with those of neutron-star X-ray binaries in the soft state, whose luminosity is higher than 1.8x10^37 erg s^-1. This suggests a source distance of >17 kpc.

astro-ph.HE↗

Outburst of LS V+44 17 Observed by MAXI and RXTE, and Discovery of a Dip Structure in the Pulse Profile

We report on the first observation of an X-ray outburst of a Be/X-ray binary pulsar LS V +44 17/RX J0440.9+4431, and the discovery of an absorption dip structure in the pulse profile. An outburst of this source was discovered by MAXI GSC in 2010 April. It was the first detection of the transient activity of LS V +44 17 since the source was identified as a Be/X-ray binary in 1997. From the data of the follow-up RXTE observation near the peak of the outburst, we found a narrow dip structure in its pulse profile which was clearer in the lower energy bands. The pulse-phase-averaged energy spectra in the 3$-$100 keV band can be fitted with a continuum model containing a power-law function with an exponential cutoff and a blackbody component, which are modified at low energy by an absorption component. A weak iron K$α$ emission line is also detected in the spectra. From the pulse-phase-resolved spectroscopy we found that the absorption column density at the dip phase was much higher than those in the other phases. The dip was not seen in the subsequent RXTE observations at lower flux levels. These results suggest that the dip in the pulse profile originates from the eclipse of the radiation from the neutron star by the accretion column.

astro-ph.HE↗

Long-term Monitoring of the Black Hole Binary GX 339-4 in the High/Soft State during the 2010 Outburst with MAXI/GSC

We present the results of monitoring the Galactic black hole candidate GX 339-4 with the Monitor of All-sky X-ray Image (MAXI) / Gas Slit Camera (GSC) in the high/soft state during the outburst in 2010. All the spectra throughout the 8-month period are well reproduced with a model consisting of multi-color disk (MCD) emission and its Comptonization component, whose fraction is <= 25% in the total flux. In spite of the flux variability over a factor of 3, the innermost disk radius is constant at R_in = 61 +/- 2 km for the inclination angle of i = 46 deg and the distance of d=8 kpc. This R_in value is consistent with those of the past measurements with Tenma in the high/soft state. Assuming that the disk extends to the innermost stable circular orbit of a non-spinning black hole, we estimate the black hole mass to be M = 6.8 +/- 0.2 M_sun for i = 46 deg and d = 8 kpc, which is consistent with that estimated from the Suzaku observation of the previous low/hard state. Further combined with the mass function, we obtain the mass constraint of 4.3 M_sun < M < 13.3 M_sun for the allowed range of d = 6-15 kpc and i < 60 deg. We also discuss the spin parameter of the black hole in GX 339-4 by applying relativistic accretion disk models to the Swift/XRT data.

astro-ph.HE↗

Revisit of Local X-ray Luminosity Function of Active Galactic Nuclei with the MAXI Extragalactic Survey

We construct a new X-ray (2--10 keV) luminosity function of Compton-thin active galactic nuclei (AGNs) in the local universe, using the first MAXI/GSC source catalog surveyed in the 4--10 keV band. The sample consists of 37 non-blazar AGNs at $z=0.002-0.2$, whose identification is highly ($>97%$) complete. We confirm the trend that the fraction of absorbed AGNs with $N_{\rm H} > 10^{22}$ cm$^{-2}$ rapidly decreases against luminosity ($L_{\rm X}$), from 0.73$\pm$0.25 at $L_{\rm X} = 10^{42-43.5}$ erg s$^{-1}$ to 0.12$\pm0.09$ at $L_{\rm X} = 10^{43.5-45.5}$ erg s$^{-1}$. The obtained luminosity function is well fitted with a smoothly connected double power-law model whose indices are $γ_1 = 0.84$ (fixed) and $γ_2 = 2.0\pm0.2$ below and above the break luminosity, $L_{*} = 10^{43.3\pm0.4}$ ergs s$^{-1}$, respectively. While the result of the MAXI/GSC agrees well with that of HEAO-1 at $L_{\rm X} \gtsim 10^{43.5}$ erg s$^{-1}$, it gives a larger number density at the lower luminosity range. Comparison between our luminosity function in the 2--10 keV band and that in the 14--195 keV band obtained from the Swift/BAT survey indicates that the averaged broad band spectra in the 2--200 keV band should depend on luminosity, approximated by $Γ\sim1.7$ for $L_{\rm X} \ltsim 10^{44}$ erg s$^{-1}$ while $Γ\sim 2.0$ for $L_{\rm X} \gtsim 10^{44}$ erg s$^{-1}$. This trend is confirmed by the correlation between the luminosities in the 2--10 keV and 14--195 keV bands in our sample. We argue that there is no contradiction in the luminosity functions between above and below 10 keV once this effect is taken into account.

astro-ph.HE↗

The First MAXI/GSC Catalog in the High Galactic-Latitude Sky

We present the first unbiased source catalog of the Monitor of All-sky X-ray Image (MAXI) mission at high Galactic latitudes ($|b| > 10^{\circ}$), produced from the first 7-month data (2009 September 1 to 2010 March 31) of the Gas Slit Camera in the 4--10 keV band. We develop an analysis procedure to detect faint sources from the MAXI data, utilizing a maximum likelihood image fitting method, where the image response, background, and detailed observational conditions are taken into account. The catalog consists of 143 X-ray sources above 7 sigma significance level with a limiting sensitivity of $\sim1.5\times10^{-11}$ ergs cm$^{-2}$ s$^{-1}$ (1.2 mCrab) in the 4--10 keV band. Among them, we identify 38 Galactic/LMC/SMC objects, 48 galaxy clusters, 39 Seyfert galaxies, 12 blazars, and 1 galaxy. Other 4 sources are confused with multiple objects, and one remains unidentified. The log $N$ - log $S$ relation of extragalactic objects is in a good agreement with the HEAO-1 A-2 result, although the list of the brightest AGNs in the entire sky has significantly changed since that in 30 years ago.

astro-ph.HE↗

Peculiarly Narrow SED of GRB 090926B with MAXI and Fermi/GBM

The monitor of all-sky X-ray image (MAXI) Gas Slit Camera (GSC) on the International Space Station (ISS) detected a gamma-ray burst (GRB) on 2009, September 26, GRB\,090926B. This GRB had extremely hard spectra in the X-ray energy range. Joint spectral fitting with the Gamma-ray Burst Monitor on the Fermi Gamma-ray Space Telescope shows that this burst has peculiarly narrow spectral energy distribution and is represented by Comptonized blackbody model. This spectrum can be interpreted as photospheric emission from the low baryon-load GRB fireball. Calculating the parameter of fireball, we found the size of the base of the flow $r_0 = (4.3 \pm 0.9) \times 10^{9} \, Y^{\prime \, -3/2}$ cm and Lorentz factor of the plasma $Γ= (110 \pm 10) \, Y^{\prime \, 1/4}$, where $Y^{\prime}$ is a ratio between the total fireball energy and the energy in the blackbody component of the gamma-ray emission. This $r_0$ is factor of a few larger, and the Lorentz factor of 110 is smaller by also factor of a few than other bursts that have blackbody components in the spectra.

astro-ph.HE↗