SearcharxivSearch

arXiv subjects

Yijun Zhao

Publications and source records attributed to Yijun Zhao.

10 recordsLinked to original sources

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations

Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformations alter the structure of visual information. This study examines changes in representation through spatial-frequency accessibility, measured by the linear recoverability of band-limited Fourier energy from model representations. To isolate effects beyond dimensionality reduction, we introduce Residual Spectral Loss (RSL), which evaluates changes relative to a dimension-matched random projection baseline. To reduce confounding effects from optimization, the analysis uses pretrained models with all parameters frozen. The experimental results show consistent frequency-dependent changes in accessibility across CLIP and DINOv2 on ImageNet and MS-COCO datasets. Spectral accessibility follows a non-monotonic trajectory across depth, peaking at intermediate layers before decreasing toward the output representation. The final transformation differs across architectures: CLIP's learned projection is spectrally neutral, with changes explained by compression, whereas DINOv2's [CLS] pooling induces a structured loss across the spectrum. These findings identify intermediate layers and pooling mechanisms as primary drivers of spectral transformation in modern vision encoders.

cs.CV

Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning

Language models (LMs) and classical machine learning methods offer complementary strengths for predictive modeling, yet their fundamentally different representations and training paradigms hinder effective integration: LMs rely on gradient-based optimization over textual data, whereas models such as Random Forests (RF) employ non-differentiable feature partitioning. This work introduces a reciprocal co-training framework that couples an LM with an RF classifier via reinforcement learning, creating an iterative feedback loop in which each model improves using signals from the other. Tabular data are reformulated into standardized textual representations for the LM, whose embeddings augment the RF feature space, while calibrated RF probability estimates provide feedback signals that guide reinforcement learning updates of the LM. Experiments across three medical datasets, evaluated with both a domain-adapted clinical encoder (ClinicalBERT) and a larger instruction-tuned language model (Qwen2-7B-Instruct), demonstrate consistent performance gains for both model components. Ablation analyses indicate that iterative refinement, hybrid reward design, and dimensionality control jointly contribute to these gains. SHAP analysis further confirms that LM-derived representations are among the most important inputs to the RF predictions. The proposed framework provides a general mechanism that allows incompatible model families to leverage each other's strengths through bidirectional adaptation.

cs.CL

MixTeX: Data-Efficient LaTeX OCR via Synthetic Pretraining and Limited Fine-Tuning

LaTeX OCR converts scientific document images into editable LaTeX code. Existing systems rely on large paired datasets, which are costly to collect and limited for low-resource languages. This paper presents MIXTEX, a data-efficient system using synthetic pretraining without real LaTeX sources. Unlike Nougat that depends on arXiv datasets, we generate training data by randomly pairing grammatical Wikipedia text with LaTeX formulas, requiring only syntactic correctness. This eliminates dependency on real document collections, enables scalable data generation (120M tokens), and supports low-resource languages. Following synthetic pretraining, adaptation requires only 400 real samples. Evaluation on a 977-sample benchmark with printed and handwritten English and Chinese shows that this two-stage strategy outperforms methods trained on large real datasets while requiring less human effort and computation. Data, code, and models are publicly available.

cs.CV

Close-range Human Following Control on a Cane-type Robot with Multi-camera Fusion

Cane-type robots have been utilized to assist and supervise the mobility-impaired population. One essential technique for cane-type robots is human following control, which allows the robot to follow the user. However, the limited perceptible information of humans by sensors at close range, combined with the occlusion caused by lower limb swing during normal walking, affect the localization of users. These limitations make it difficult to achieve human following at close range.To address these challenges, this study developed a new cane-type wheeled robot and proposed a novel human-following control with multi-camera fusion. This control system mainly consists of two parts: 1) a human following controller that locates a user by multi-camera fusion and generates control signals to follow the user. 2) a cane robot controller designed to steer the cane robot to a target position. The proposed strategy's effectiveness has been validated in outdoor experiments with six healthy subjects. The experimental scenarios included different terrains (i.e., straight, turning, and inclined paths), road conditions (i.e., flat and rough roads), and walking speeds. The obtained results showed that the average tracking error for position and orientation was less than 5 cm and 15° respectively across all scenarios. Moreover, the cane robot can effectively adapt to a wide range of individual gait patterns and achieve stable human following at daily walking speeds (0.75 m/s - 1.45 m/s).

eess.SY

Enhance Gender and Identity Preservation in Face Aging Simulation for Infants and Toddlers

Realistic age-progressed photos provide invaluable biometric information in a wide range of applications. In recent years, deep learning-based approaches have made remarkable progress in modeling the aging process of the human face. Nevertheless, it remains a challenging task to generate accurate age-progressed faces from infant or toddler photos. In particular, the lack of visually detectable gender characteristics and the drastic appearance changes in early life contribute to the difficulty of the task. We propose a new deep learning method inspired by the successful Conditional Adversarial Autoencoder (CAAE, 2017) model. In our approach, we extend the CAAE architecture to 1) incorporate gender information, and 2) augment the model's overall architecture with an identity-preserving component based on facial features. We trained our model using the publicly available UTKFace dataset and evaluated our model by simulating up to 100 years of aging on 1,156 male and 1,207 female infant and toddler face photos. Compared to the CAAE approach, our new model demonstrates noticeable visual improvements. Quantitatively, our model exhibits an overall gain of 77.0% (male) and 13.8% (female) in gender fidelity measured by a gender classifier for the simulated photos across the age spectrum. Our model also demonstrates a 22.4% gain in identity preservation measured by a facial recognition neural network.

cs.CV

Localized Motion Artifact Reduction on Brain MRI Using Deep Learning with Effective Data Augmentation Techniques

In-scanner motion degrades the quality of magnetic resonance imaging (MRI) thereby reducing its utility in the detection of clinically relevant abnormalities. We introduce a deep learning-based MRI artifact reduction model (DMAR) to localize and correct head motion artifacts in brain MRI scans. Our approach integrates the latest advances in object detection and noise reduction in Computer Vision. Specifically, DMAR employs a two-stage approach: in the first, degraded regions are detected using the Single Shot Multibox Detector (SSD), and in the second, the artifacts within the found regions are reduced using a convolutional autoencoder (CAE). We further introduce a set of novel data augmentation techniques to address the high dimensionality of MRI images and the scarcity of available data. As a result, our model was trained on a large synthetic dataset of 225,000 images generated from 375 whole brain T1-weighted MRI scans. DMAR visibly reduces image artifacts when applied to both synthetic test images and 55 real-world motion-affected slices from 18 subjects from the multi-center Autism Brain Imaging Data Exchange (ABIDE) study. Quantitatively, depending on the level of degradation, our model achieves a 27.8%-48.1% reduction in RMSE and a 2.88--5.79 dB gain in PSNR on a 5000-sample set of synthetic images. For real-world artifact-affected scans from ABIDE, our model reduced the variance of image voxel intensity within artifact-affected brain regions (p = 0.014).

eess.IV

Magnetoplasmon-surface phonon polaritons coupling effects in radiative heat transfer

In this letter, based on the quantum Hall regime of magneto-optical graphene, we have theoretically investigated the coupling of magnetoplasmon polaritons (MPP) to surface phonon polaritons (SPhPs) by investigating the radiative heat transfer between two graphene-coated SiO2 slabs. By applying an external magnetic field, the separated branches of intraband and interband MPP can both couple with SPhPs to form tunable modes, which remould the energy transport of the system. The heat transfer mechanism is completely changed from enhancement to attenuation due to the strong coupling, and the thermal stealthy is realized for the graphene. The letter has great significance for the graphene-based magneto-optical devices.

physics.optics

Active control of near-field radiative heat transfer by coating-twisting method

In this letter, active control of near-field radiative heat transfer (NFRHT) between two isotropic materials is realized by a coating-twisting method. The two slabs are coated with graphene gratings, and then the NFRHT can be not only enhanced, but also weakened, by tuning the twisted angle between the two gratings. The physical mechanism is attributed to the modes coupled by the graphene gratings and the isotropic material, which can vary with the twisted angle. The proposed method is also applicable for other kinds of anisotropic films, and may provide a way to realize high-precision nanoscale thermal management, nimble thermal communications and thermal switch.

physics.app-ph

Graphene-based thermal repeater

In this letter, we have demonstrated the possibility to efficiently relay the radiative heat flux between two nanoparticles by opening a smooth channel for heat transfer. By coating the nanoparticles with a silica shell and modifying the substrate with multilayered graphene sheets respectively, the localized phonon polaritons excited near the nanoparticles can couple with the multiple surface plasmon polaritons near the substrate to realize the heat relay in the long distance. The heat transfer can be enhanced by more than six orders of magnitude and the relay distance can be as high as 35 times in the far-field regime. The work may provide a way to realize the energy modulation or thermal communications especially in long distance.

cond-mat.mes-hall

Magnetic-tunable nanoscale thermal radiation between twisted graphene gratings

This paper presents a comprehensive theoretical study of the magnetic-tunable near-field radiative heat transfer (NFRHT) between two twisted graphene gratings. As a result of the quantum Hall regime of magneto-optical graphene and the grating effect, three types of graphene surface plasmon polaritons (SPPs) modes are observed in the system: near-zero modes, high-frequency hyperbolic modes, and elliptic modes. The elliptic SPPs modes, which are caused by the combined effect of magnetic field and grating, are observed in the graphene grating system for the first time. In addition, the near-zero modes can be greatly enhanced by the combined effect grating and magnetic field, rendering graphene devices promising for thermal communication at ultra-low frequency. In particular, the near-zero modes result in a unique enhancement region of heat transfer, no matter for any twisted angle between gratings. The combined effect of grating and magnetic field is investigated simultaneously. By changing the strength of magnetic field, the positions and intensities of the modes can be modulated, and hence the NFRHT can be tuned accordingly, no matter for parallel or twisted graphene gratings. The magnetic field endows the grating action (graphene filling factors and twisted angles) with a higher modulation ability to modulate the NFRHT compared with zero-field. Moreover, the modulation ability of twist can be tuned by the magnetic field at different twisted angles. In sum, the combined effect of magnetic field and grating provides a tunable way to realize the energy modulation or multi-frequency thermal communications related to graphene devices.

cond-mat.mes-hall