SearcharxivSearch

arXiv subjects

Donghun Ryu

Publications and source records attributed to Donghun Ryu.

9 recordsLinked to original sources

LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) enables real-time novel view synthesis but produces millions of primitives through adaptive densification, leading to significant storage overhead. Learned-mask pruning methods such as LP-3DGS address this by assigning each Gaussian a learnable mask to identify and prune redundant primitives. However, we identify a limitation of this paradigm: the steep slope of the Gumbel-Sigmoid activation drives mask values to the extremes within the short mask-training window, before the importance ranking has stabilized, producing a sharply bimodal distribution from which that ranking can no longer be reliably recovered. We propose LinearMask-GS, which replaces Gumbel-Sigmoid with a linear increment activation that keeps mask values in a mid-confidence regime throughout mask training, producing a stable, unimodal mask distribution whose ranking tracks importance. On Mip-NeRF 360, our method achieves 3.6x and 1.6x Gaussian reductions over 3DGS and LP-3DGS, respectively, while maintaining or improving rendering quality. For outdoor scenes, it yields a 1.6x reduction (from 2.18M to 1.36M) with notable gains in PSNR (+0.38 dB), SSIM (+0.025), and LPIPS (-0.029).

cs.CV

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA

cs.AI

CongNaMul: A Dataset for Advanced Image Processing of Soybean Sprouts

We present 'CongNaMul', a comprehensive dataset designed for various tasks in soybean sprouts image analysis. The CongNaMul dataset is curated to facilitate tasks such as image classification, semantic segmentation, decomposition, and measurement of length and weight. The classification task provides four classes to determine the quality of soybean sprouts: normal, broken, spotted, and broken and spotted, for the development of AI-aided automatic quality inspection technology. For semantic segmentation, images with varying complexity, from single sprout images to images with multiple sprouts, along with human-labelled mask images, are included. The label has 4 different classes: background, head, body, tail. The dataset also provides images and masks for the image decomposition task, including two separate sprout images and their combined form. Lastly, 5 physical features of sprouts (head length, body length, body thickness, tail length, weight) are provided for image-based measurement tasks. This dataset is expected to be a valuable resource for a wide range of research and applications in the advanced analysis of images of soybean sprouts. Also, we hope that this dataset can assist researchers studying classification, semantic segmentation, decomposition, and physical feature measurement in other industrial fields, in evaluating their models. The dataset is available at the authors' repository. (https://bhban.kr/data)

cs.CV

Calibration-free quantitative phase imaging using data-driven aberration modeling

We present a data-driven approach to compensate for optical aberration in calibration-free quantitative phase imaging (QPI). Unlike existing methods that require additional measurements or a background region to correct aberrations, we exploit deep learning techniques to model the physics of aberration in an imaging system. We demonstrate the generation of a single-shot aberration-corrected field image by using a U-net-based deep neural network that learns a translation between an optical field with aberrations and an aberration-corrected field. The high fidelity of our method is demonstrated on 2D and 3D QPI measurements of various confluent eukaryotic cells, benchmarking against the conventional method using background subtractions.

eess.IV

Machine learning approach to remove ion interference effect in agricultural nutrient solutions

High concentration agricultural facilities such as vertical farms or plant factories consider hydroponic techniques as optimal solutions. Although closed-system dramatically reduces water consumption and pollution issues, it has ion-ratio related problem. As the root absorbs individual ions with different rate, ion rate in a nutrient solution should be adjusted periodically. But traditional method only considers pH and electrical conductivity to adjust the nutrient solution, leading to ion imbalance and accumulation of excessive salts. To avoid those problems, some researchers have proposed ion-balancing methods which measure and control each ion concentration. However, those approaches do not overcome the innate limitations of ISEs, especially ion interference effect. An anion sensor is affected by other anions, and the error grows larger in higher concentration solution. A machine learning approach to modify ISE data distorted by ion interference effect is proposed in this paper. As measurement of TDS value is relatively robust than any other signals, we applied TDS as key parameter to build a readjustment function to remove the artifact. Once a readjustment model is established, application on ISE data can be done in real time. Readjusted data with proposed model showed about 91.6 ~ 98.3% accuracies. This method will enable the fields to apply recent methods in feasible status.

cs.LG

ODE network model for nonlinear and complex agricultural nutrient solution system

In closed hydroponic systems, periodic readjustment of nutrient solution is necessary to continuously provide stable environment to plant roots because the interaction between plant and nutrient solution changes the rate of ions in it. The traditional method is to repeat supplying small amount of premade concentrated nutrient solution, measuring total electric conductivity and pH of the tank only. As it cannot control the collapse of ion rates, recent researches try to measure the concentration of individual components to provide insufficient ions only. However, those approaches use titrationlike heuristic approaches, which repeat adding small amount of components and measuring ion density a lot of times for a single control input. Both traditional and recent methods are not only time-consuming, but also cannot predict chemical reactions related with control inputs because the nutrient solution is a nonlinear complex system, including many precipitation reactions and complicated interactions. We present a continuous network model of the nutrient solution system, whose reactions are described as differential equations. The model predicts molar concentration of each chemical components and total dissolved solids with low error. This model also can calculate the amount of chemical compounds needed to produce a desired nutrient solution, by reverse calculation from dissolved ion concentrations.

eess.SY

Deep learning-enabled image quality control in tomographic reconstruction: Robust optical diffraction tomography

In tomographic reconstruction, the image quality of the reconstructed images can be significantly degraded by defects in the measured two-dimensional (2D) raw image data. Despite the importance of screening defective 2D images for robust tomographic reconstruction, manual inspection and rule-based automation suffer from low-throughput and insufficient accuracy, respectively. Here, we present deep learning-enabled quality control for holographic data to produce robust and high-throughput optical diffraction tomography (ODT). The key idea is to distill the knowledge of an expert into a deep convolutional neural network. We built an extensive database of optical field images with clean/noisy annotations, and then trained a binary classification network based upon the data. The trained network outperformed visual inspection by non-expert users and a widely used rule-based algorithm, with > 90% test accuracy. Subsequently, we confirmed that the superior screening performance significantly improved the tomogram quality. To further confirm the trained model's performance and generalizability, we evaluated it on unseen biological cell data obtained with a setup that was not used to generate the training dataset. Lastly, we interpreted the trained model using various visualization techniques that provided the saliency map underlying each model inference.

eess.IV

Deep learning approach to coherent noise reduction in optical diffraction tomography

We present a deep neural network to reduce coherent noise in three-dimensional quantitative phase imaging. Inspired by the cycle generative adversarial network, the denoising network was trained to learn a transform between two image domains: clean and noisy refractive index tomograms. The unique feature of this network, distinct from previous machine learning approaches employed in the optical imaging problem, is that it uses unpaired images. The learned network quantitatively demonstrated its performance and generalization capability through denoising experiments of various samples. We concluded by applying our technique to reduce the temporally changing noise emerging from focal drift in time-lapse imaging of biological cells. This reduction cannot be performed using other optical methods for denoising.

physics.optics

Subsampled Phase Retrieval for On-chip Lensless Holographic Video

On-chip holographic video is a convenient way to monitor biological samples simultaneously at high spatial resolution and over a wide field-of-view. However, due to the limited readout rate of digital detector arrays, one often faces a tradeoff between the per-frame pixel count and frame rate of the captured video. In this report, we propose a subsampled phase retrieval (SPR) algorithm to overcome the spatial-temporal trade-off in holographic video. Compared to traditional phase retrieval approaches, our SPR algorithm uses over an order of magnitude less pixel measurements while maintaining suitable reconstruction quality. We use an on-chip holographic video setup with pixel sub-sampling to experimentally demonstrate a factor of 5.5 increase in sensor frame rate while monitoring the in vivo movement of Peranema microorganisms.

physics.optics