SearcharxivSearch

arXiv subjects

Zhijun Tu

Publications and source records attributed to Zhijun Tu.

At least 19 recordsLinked to original sources

Stable FP4 Training via Transposition-Invariant Block Quantization

Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization. We identify a fundamental source of this instability in existing microscaling approaches: scale inconsistency induced by tensor transposition. In conventional 1D block quantization, forward and backward passes assign di erent scaling factors to the same values after transposition, leading to biased and unstable gradient updates. To address this issue, we propose a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations. We further combine this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients. To handle the sensitivity of attention mechanisms, we adopt MXFP8 quantization for query and key projections, yielding a practical mixed-precision design. We evaluate our method on dense LLMs up to 7B parameters and a 30B Mixture-of-Experts model, trained on up to 100B tokens. Across all settings, our approach achieves stable end-to-end FP4 training and closely matches BF16 performance, with less than 1.3% degradation in perplexity and downstream accuracy. These results demonstrate that enforcing forwardbackward scaling consistency is su cient to enable practical FP4 training at scale, providing a simple and e ective pathway toward more e cient LLM training.

cs.LG

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts

Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However, VAR still suffers from inherent deficiencies in multi-scale representation learning. Specifically, lower scales primarily capture global semantics, while higher scales focus on fine-grained details. Employing a shared architecture across scales induces optimization conflicts. Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. In addition, we enhance semantic modeling at early scales by incorporating external self-supervised features. Unlike naive alignment, we analyse and design a residual feature aggregation scheme tailored to the VAR paradigm. Extensive experiments show that our method significantly improves both training efficiency and generation quality. On the ImageNet 256*256 benchmark, our model achieves a superior FID compared to the dense baseline while requiring only half of the default training epochs and a smaller parameter budget, with a merely marginal increase in training cost. Moreover, the performance gap further widens with larger training epochs.

cs.CV

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models

1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to fully leverage pre-trained models. This results in high training costs and notable accuracy degradation. We identify that the large gap between full precision and 1-bit representations makes naive adaptation difficult. In this paper, we introduce a consistent progressive training for both forward and backward, smoothly converting the full-precision weights into the binarized ones. Additionally, we incorporate binary-aware initialization and dual-scaling compensation to reduce the difficulty of progressive training and improve the performance. Experimental results on LLMs of various sizes demonstrate that our method outperforms existing approaches. Our results show that high-performance 1-bit LLMs can be achieved using pre-trained models, eliminating the need for expensive training from scratch.

cs.CL

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffusion-3 and FLUX.1. However, deploying these models on resource-constrained mobile devices entails prohibitive computational and memory overhead. While efficiency-driven approaches like Linear-DiT and static pruning alleviate bottlenecks, they often incur quality degradation. Unlike cloud environments, mobile constraints require a single-model paradigm that dynamically balances fidelity and latency. We introduce ElasticDiT, which achieves this dynamic trade-off by adjusting spatial compression ratios and DiT block depths. By integrating Shift Sparse Block Attention (SSBA) and a Tiny DWT-Distilled VAE (T-DVAE), ElasticDiT reduces inference latency and memory footprint while maintaining image quality. Experiments confirm that ElasticDiT effectively covers a wide range of fidelity-latency trade-offs within a single set of parameters. By jointly adjusting compression and depth, a single ElasticDiT model can be reconfigured on-the-fly to outperform task-specific baselines. Specifically, our flex lite variant achieves an HPS of 32.87, surpassing the Flux model, while maintaining competitive quality at 84.16 percent average sparsity through SSBA. Furthermore, the plug-and-play T-DVAE provides SD3-level reconstruction with only 1/8x the computational cost of standard VAEs, and Flow-GRPO boosts semantic alignment (GenEval: 66.93 to 73.62). These results demonstrate that ElasticDiT offers a versatile, hardware-adaptive solution that eliminates the need for multiple specialized models, providing a promising path for future high-resolution image generation on mobile devices.

cs.CV

A Search for Variability of Ultracool Dwarfs with the Zwicky Transient Facility

Rotationally modulated photometric variability of ultracool dwarfs encodes key information about cloud structure and temperature contrasts. Large homogeneous optical datasets are crucial for linking atmospheric heterogeneity to fundamental parameters such as rotation, mass, and age. We present a search for rotation periods in ultracool dwarfs using Zwicky Transient Facility (ZTF) optical light curves. By propagating the coordinates to the ZTF epoch and applying Lomb-Scargle analysis, we identified 226 periodic variables, including 32 robust detections and 194 tentative cases. Among the robust detections, 12 have no previously published periods, while 20 have literature counterparts, most of which are consistent with the published values. Most robust detections are M dwarfs, reflecting the optical sensitivity limits of ZTF. We find a trend of decreasing periods toward later spectral types in relatively old dwarfs (> 100 Myr), suggesting faster rotation for late-M types than for mid-M types. The age-period relation of our sample is broadly consistent with angular-momentum-conservation models at higher-mass regime of brown dwarfs, consistent with the M-dwarf bias of our catalog. Many additional candidates remain to be confirmed due to sparse sampling or low S/N. Future high-cadence, multi-wavelength monitoring and systematic mining of ZTF and upcoming surveys will be crucial for validating these periods, extend sensitivity to later (L/T) types, and better connect rotation with cloud physics across the stellar-substellar boundary.

astro-ph.SR

SPHEREx Ultracool Dwarf spectrophotometric Atlas (SUDA): Atmospheric and Fundamental Parameters of Ultracool Dwarfs

We present the SPHEREx Ultracool Dwarf spectrophotometric Atlas (SUDA), a homogeneous sample of 1675 field ultracool dwarfs with continuous 0.75--5~$\mu$m spectrophotometry from SPHEREx Quick Release 2 (QR2). Using the SAND and ATMO2020++ atmospheric grids, we derive atmospheric parameters, compute bolometric luminosities ($L_{\rm bol}$), and combine $T_{\mathrm{eff}}$ and radii with evolutionary tracks to estimate masses, ages, and evolutionary $\log g$. For sources at 1700--2500~K, spectroscopic $\log g$ is systematically lower than evolutionary $\log g$, with a median offset of $\sim$1.05~dex, likely driven primarily by cloud-related model uncertainties. We further construct an empirical atlas by binning the measurements in $T_{\mathrm{eff}}$ and evolutionary $\log g$, producing 57 spectrophotometric templates spanning $T_{\mathrm{eff}}\simeq700$--3000~K. Molecular indices trace a coherent atmospheric sequence: H$_2$O and CH$_4$ indices strengthen toward lower $T_{\mathrm{eff}}$, while CO and CO$_2$ indices increase below $\sim$1500~K and turn over near $\sim$1000~K. Model comparisons indicate that CO$_2$ index is more sensitive to metallicity than gravity at $T_{\mathrm{eff}}\sim800$--1200~K. SUDA provides a reference sample linking 0.75--5~$\mu$m spectrophotometric morphology to atmospheric and evolutionary trends in ultracool dwarfs.

astro-ph.SR

Physical properties of RhGe and CoGe single crystals synthesized under high pressure

Chiral topological semimetals hosting multifold fermions and exotic surface states represent a frontier in topological materials research. Among them, noncentrosymmetric cubic B20 compounds-notably transition-metal silicides and germanides-offer a unique platform for realizing symmetry-protected topological phases and unconventional optoelectronic responses. Here, we report the physical properties of RhGe and CoGe single crystals with B20 structure in detail. Transport measurements reveal metallic behavior with characteristic Fermi-liquid scaling at low temperatures, while magnetization results confirm paramagnetism in both compounds. In addition, both of materials exhibit low carrier concentrations with small electronic specific heat coefficient, indicating their semimetal feature with weak electronic correlations. Such high-quality CoGe and RhGe single crystals provide a material platform to explore the evolution of multifold fermions and the instability of helicoid-arc surface states with spin-orbit coupling and surface environment in B20 material systems.

cond-mat.str-el

A Morpho-kinematic Study of Galactic High-ADF PNe Based on the VLT/UVES Deep Spectroscopy

We report detailed analyses of deep, high-resolution spectra of three Galactic planetary nebulae (PNe) with high abundance discrepancy factors (ADFs), Hf2-2, M1-42 and NGC6153, obtained with the Ultraviolet and Visual Echelle Spectrograph (UVES) on the 8.2m Very Large Telescope (VLT). These spectra were carefully reduced, including rigorous absolute flux calibration, yielding detections of ~410-800 emission lines in each PN. Plasma diagnostics and abundance calculations were critically performed using nebular lines. In all three PNe, the electron temperatures derived using the collisionally excited lines (CELs) are higher than those yielded by the HI Balmer and Paschen jumps, while the temperatures yielded by the OII and NII optical recombination lines (ORLs) are very low, <2000 K, indicating that the heavy-element ORLs probe cold nebular regions. The ORL abundances of N, O and Ne are systematically higher than the corresponding CEL values, confirming high ADFs in the three objects. Position-velocity (PV) diagrams were created, and spatio-kinematical studies show that CELs come from the outer nebular regions, while the ORL-emitting regions are close to nebular center. Additionally, the velocity indicated by CEL line-splitting decreases with ionization potential, which was not obvious in ORLs. These spatial and kinematic differences support two distinct components of ionized gas: a cold, metal-rich component and a warmer component with normal metallicity. Heavy elements are strongly enriched in the cold gas, while its H^+ fraction is low but still produces significant HI emission, affecting CEL abundance estimates.

astro-ph.SR

Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution

The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-tuning pre-trained diffusion models through Low-Rank Adaptation (LoRA) module to reconstruct high-resolution (HR) images. However, these dense Real-ISR models are limited in their ability to adaptively capture the heterogeneous characteristics of complex real-world degraded samples or enable knowledge sharing between inputs under equivalent computational budgets. To address this, we investigate the integration of sparse MoE into Real-ISR and propose a Mixture-of-Ranks (MoR) architecture for single-step image super-resolution. We introduce a fine-grained expert partitioning strategy that treats each rank in LoRA as an independent expert. This design enables flexible knowledge recombination while isolating fixed-position ranks as shared experts to preserve common-sense features and minimize routing redundancy. Furthermore, we develop a degradation estimation module leveraging CLIP embeddings and predefined positive-negative text pairs to compute relative degradation scores, dynamically guiding expert activation. To better accommodate varying sample complexities, we incorporate zero-expert slots and propose a degradation-aware load-balancing loss, which dynamically adjusts the number of active experts based on degradation severity, ensuring optimal computational resource allocation. Comprehensive experiments validate our framework's effectiveness and state-of-the-art performance.

cs.CV

Coexistence of near-EF van Hove singularity and in-gap topological Dirac surface states in superconducting electrides

Superconducting electrides have attracted growing attention for their potential to achieve high superconducting transition temperatures (TC) under pressure. However, many known electrides are chemically reactive and unstable, making high-quality single-crystal growth, characterization, and measurements difficult, and most do not exhibit superconductivity at ambient pressure. In contrast, La3In stands out for its ambient-pressure superconductivity (TC ~ 9.4 K) and the availability of high-quality single crystals. Here, we investigate its low-energy electronic structure using angle-resolved photoemission spectroscopy and first-principles calculations. The bands near the Fermi energy are mainly derived from La 5d and In 5p orbitals. A saddle point is directly observed at the Brillouin zone (BZ) boundary, while a three-dimensional van Hove singularity crosses EF at the BZ corner. First-principles calculations further reveal topological Dirac surface states within the bulk energy gap above EF. The coexistence of a high density of states and in-gap topological surface states near EF suggests that La3In offers a promising platform for tuning superconductivity and exploring possible topological superconducting phases through doping or external pressure.

cond-mat.supr-con

Evidence for Anion-Free-Electron Duality and Enhanced Superconducting Role of Interstitial Anionic Electrons in Electrides

The discovery of superconducting electrides, characterized by interstitial anionic electrons (IAEs) residing in lattice cavities, has established a distinctive platform for investigating superconductors. Yet the superconducting origin and the fundamental role of IAEs in Cooper pairing formation remain poorly understood due to the challenges in directly observing IAEs. Here, combining angle-resolved photoemission spectroscopy (ARPES), transport measurements, and first-principles calculations, we certify that the IAEs in electride La3In (Tc = 9.4 K) exhibit a dual nature as both anions and free electrons. With the finite-depth potential well model, we trace that IAEs originate from electronic states near the Fermi level located above potential barriers, forming a Fermi sea susceptible to scattering by La-derived phonons, triggering superconductivity. ARPES combined with high-resolution XRD measurements on oxygen-treated samples directly reveals IAEs' spatial distribution and energy dispersion from interstitial sites with the consistent energy value predicted by our theory model. The concomitant diminution of free electrons upon oxygen treatment, leading to a marked reduction in superconductivity, further provides compelling experimental evidence that IAEs actively participate in electron-phonon coupling. Our findings resolve the long-standing ambiguity regarding the electronic nature of IAEs, elucidate their enhancing superconductivity in the phonon-mediated mechanism, and provide a foundation for exploring advanced electride-based superconductors.

cond-mat.supr-con

PyEMILI: A New Generation Computer-aided Spectral Line Identifier -- II. Emission-line Identification and Plasma Diagnostics of a Sample of Gaseous Nebulae

In order to test the robustness and reliability of the new generation spectral-line identifier PyEMILI, as initially introduced in Paper I, in line identification and establish a reference/benchmark dataset for future spectroscopic studies, we run the code on the line lists of a selected sample of emission-line nebulae, including planetary nebulae (PNe), HII regions, and Herbig-Haro (HH) objects with deep high-dispersion spectroscopic observations published over the past two decades. The automated line identifications by PyEMILI demonstrate significant improvements in both completeness and accuracy compared to the previous manual identifications in the literature. Since our last report of PyEMILI, the atomic transition database used by the code has been further expanded by cross-matching the Kurucz Line Lists. Moreover, to aid the PyEMILI identification of numerous faint optical recombination lines (ORLs) of CII, NII, OII and NeII, we compiled a new dataset of effective recombination coefficients for these nebular lines, and created a new subroutine in the code to generate theoretical spectra of heavy-element ORLs at various electron temperature and density cases; these theoretical spectra can be used to fit the observed recombination spectrum of a PN to obtain the electron temperature, density and ionic abundances using the Markov-Chain Monte Carlo (MCMC) method. We present MCMC-derived parameters for a sample of PNe. This work establishes PyEMILI as a robust and versatile tool for both line identification and plasma diagnostics in deep spectroscopy of gaseous nebulae.

astro-ph.SR

Superconductivity in kagome metal YRu3Si2 with strong electron correlations

We report the detailed physical properties of YRu3Si2 with the Ru kagome lattice at normal and superconducting states. The results of resistivity and magnetization show that YRu3Si2 is a type-II bulk superconductor with Tc ~ 3.0 K. The specific heat measurement further suggests that this superconductivity could originate from the weak or moderate electron-phonon coupling. On the other hand, both large Kadawaki-Woods ratio and Wilson ratio indicate that there is a strong electron correlation effect in this system, which may have a connection with the featured flat band of kagome lattice.

cond-mat.supr-con

Pronounced orbital-selective electron-electron correlation and electron-phonon coupling in V2Se2O

Orbital-selective many-body effects, in which electrons occupying different orbitals experience distinct interaction strengths, play a crucial role in correlated multiorbital materials. However, these effects usually manifest in a complex manner, obscuring their microscopic origins. Here, by combining angle-resolved photoemission spectroscopy measurements with theoretical calculations, we reveal pronounced orbital selectivity in both electron-electron correlation and electron-phonon coupling in the van der Waals material V2Se2O. Electron correlation induces distinct bandwidth renormalization exclusively in the V d_xy-derived band, while the bands mainly composed of the other d orbitals remain essentially unrenormalized. Orbital-resolved analyses identify that the filling number and the bandwidth are decisive factors governing orbital-dependent correlation. Simultaneously, the d_(xz/yz)-derived band exhibits a sharp kink anomaly, arising from enhanced coupling to high-energy phonon modes dominated by oxygen vibrations. Such pronounced orbital selectivity positions V2Se2O as a rare and prototypical platform for unravelling the microscopic mechanisms of orbital-selective electron-electron and electron-phonon interactions, and offers guiding principles for the design of correlated multiorbital materials.

cond-mat.str-el

A Large Sample of JWST/NIRSpec Brown Dwarfs: New Distant Discoveries

Brown dwarfs are essential probes of stellar and planetary formation, yet their low luminosities pose challenges for detection at large Galactic distances. The James Webb Space Telescope (JWST), with its unprecedented near-infrared sensitivity, enables the discovery and characterization of distant substellar objects, including those in the Milky Way's thick disk and halo. We conducted a systematic search using over 40,000 publicly available JWST/NIRSpec PRISM/CLEAR spectra and identified 68 brown dwarfs through spectral template matching and visual inspection. Among them, 12 are newly identified candidates, including 8 T dwarfs and 4 M/L dwarfs, most at distances exceeding 1 kpc. Remarkably, two sources -- JWST J001418.22-302223.2 and JWST J033240.07-274907.8 -- are found at distances greater than 5 kpc, making them the most distant brown dwarfs within the Milky Way. Spectral fits were performed using a nested sampling Monte Carlo algorithm with three model grids: Sonora Elf Owl, LOWZ, and SAND. The analysis reveals that cloud-free models are unable to reproduce L/T transition spectra, whereas the SAND model provides a more accurate representation of cloud effects in metal-poor environments. With the newly identified distant brown dwarfs, we also investigated the vertical metallicity gradient of brown dwarfs. Overall, the metallicities do not show an evident trend with Galactic height $|Z|$, due to the limited sample size and the uncertainties in metallicity measurements.

astro-ph.SR

Superconductivity in cubic La3Al with interstitial anionic electrons

We report the observation of superconductivity in cubic La3Al single crystal. It shows a metallic behavior at a normal state without observable structural transition and enters the superconducting state below Tc ~ 6.32 K. Detailed characterizations and analysis indicate that cubic La3Al is a bulk type-II BCS superconductor. Moreover, theoretical calculations show that it can host interstitial anionic electrons, which are located at the body center of cubic unit cell, and confirm the electron-phonon coupling as the superconducting mechamism. Thus, cubic La3Al can be regarded as an novel electride superconductor.

cond-mat.supr-con

One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation

Diffusion-based models have been widely used in various visual generation tasks, showing promising results in image super-resolution (SR), while typically being limited by dozens or even hundreds of sampling steps. Although existing methods aim to accelerate the inference speed of multi-step diffusion-based SR methods through knowledge distillation, their generated images exhibit insufficient semantic alignment with real images, resulting in suboptimal perceptual quality reconstruction, specifically reflected in the CLIPIQA score. These methods still have many challenges in perceptual quality and semantic fidelity. Based on the challenges, we propose VPD-SR, a novel visual perception diffusion distillation framework specifically designed for SR, aiming to construct an effective and efficient one-step SR model. Specifically, VPD-SR consists of two components: Explicit Semantic-aware Supervision (ESS) and High-Frequency Perception (HFP) loss. Firstly, the ESS leverages the powerful visual perceptual understanding capabilities of the CLIP model to extract explicit semantic supervision, thereby enhancing semantic consistency. Then, Considering that high-frequency information contributes to the visual perception quality of images, in addition to the vanilla distillation loss, the HFP loss guides the student model to restore the missing high-frequency details in degraded images that are critical for enhancing perceptual quality. Lastly, we expand VPD-SR in adversarial training manner to further enhance the authenticity of the generated content. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed VPD-SR achieves superior performance compared to both previous state-of-the-art methods and the teacher model with just one-step sampling.

cs.CV

Autoregressive Image Generation with Vision Full-view Prompt

In autoregressive (AR) image generation, models based on the 'next-token prediction' paradigm of LLMs have shown comparable performance to diffusion models by reducing inductive biases. However, directly applying LLMs to complex image generation can struggle with reconstructing the image's structure and details, impacting the generation's accuracy and stability. Additionally, the 'next-token prediction' paradigm in the AR model does not align with the contextual scanning and logical reasoning processes involved in human visual perception, limiting effective image generation. Prompt engineering, as a key technique for guiding LLMs, leverages specifically designed prompts to improve model performance on complex natural language processing (NLP) tasks, enhancing accuracy and stability of generation while maintaining contextual coherence and logical consistency, similar to human reasoning. Inspired by prompt engineering from the field of NLP, we propose Vision Full-view prompt (VF prompt) to enhance autoregressive image generation. Specifically, we design specialized image-related VF prompts for AR image generation to simulate the process of human image creation. This enhances contextual logic ability by allowing the model to first perceive overall distribution information before generating the image, and improve generation stability by increasing the inference steps. Compared to the AR method without VF prompts, our method shows outstanding performance and achieves an approximate improvement of 20%.

cs.CV