SearcharxivSearch

arXiv subjects

Hanyu Zhang

Publications and source records attributed to Hanyu Zhang.

At least 19 recordsLinked to original sources

SenseNova-U1.5: Towards Native Unified Visual Intelligence

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

cs.CV

A measurement of $H_0$ from DESI DR1 using energy densities

We present a new measurement of the Hubble constant, independent of standard rulers and robust to pre-recombination modifications such as Early Dark Energy (EDE), obtained by calibrating the total energy density of the Universe. We start using the present-day photon density as an anchor, and use the baryon-to-photon ratio from Big Bang Nucleosynthesis based measurements and the baryon-to-matter ratio from the baryons' imprint on galaxy clustering to translate to a physical matter density at present day. We then compare this to measurements of the ratio of the matter density to the critical density ($Ω_{\mathrm{m}}$), calculated using the relative positions of the baryon acoustic oscillations, to measure the critical density of the universe and hence $H_0$. The important measurements of the evolution of the energy density all happen at low redshift, so we consider this a low-redshift measurement. We validate our method both on a suite of $N$-body mocks and on noiseless theory vectors generated across a wide range of Hubble parameters in both $Λ$CDM and EDE cosmologies. Using DESI DR1 data combined with the angular CMB acoustic scale and the latest BBN constraints, we find $H_0 = 69.0 \pm 2.5$ km s$^{-1}$ Mpc$^{-1}$, consistent with existing early and late-time determinations of the Hubble constant. We consider the impact of non-standard dark energy evolution on our measurement. Future data, including that from further iterations of DESI and from Euclid, will add to these results providing a powerful test of the Hubble tension.

astro-ph.CO

Discovery of Three Glitches in the previously quiet pulsar PSR J1637$-$4642

We present the discovery and analysis of three rotational glitches in the young pulsar PSR J1637$-$4642. The timing observations span from 19 February 2009 to 6 October 2024 (MJD 54881$-$60589) from the Murriyang radio telescope of the Parkes Observatory. The first and strongest glitch occurred around MJD 58352 with a fractional frequency change of $Δν/ν\sim 2.7 \times 10^{-6}$, while two additional smaller glitches were detected at MJD 59443 and MJD 60445 with fractional changes of $2.2 \times 10^{-9}$ and $2.8 \times 10^{-8}$, respectively. Prior to this, the pulsar had shown no glitch activity since its discovery in the Parkes Multibeam survey. Only the first glitch exhibits detectable exponential recovery, with a decay timescale of $\sim$100 days and a small recovery fraction $\approx 0.015$, accompanied by a permanent increase in the magnitude of the spin-down rate. Modeling the post-glitch evolution of $\dotν$ within the vortex-creep framework using Bayesian inference gives a superfluid moment-of-inertia fraction $\approx 0.0187$, consistent with the inner-crust superfluid. These results reinforce the standard superfluid glitch paradigm and demonstrate that even ``quiet'' pulsars can still host substantial glitch activity.

astro-ph.HE

Frequentist Cosmological Constraints from Full-Shape Clustering Measurements in DESI DR1

We present a frequentist analysis of clustering measurements from Data Release 1 of the Dark Energy Spectroscopic Instrument (DESI) using the standard profile likelihood method. While Bayesian inferences for effective field theory models of galaxy clustering can be highly sensitive to prior choices for extended cosmological models, frequentist inferences are not susceptible to such effects. We compare frequentist and Bayesian constraints for the parameter set $\{σ_8, H_0, Ω_{\rm{m}}, w_0, w_a\}$ using the full-shape power spectrum multipoles, post-reconstruction baryon acoustic oscillation (BAO) measurements, and external datasets from the CMB and type Ia supernovae measurements. The frequentist confidence intervals are significantly shifted relative to the Bayesian credible intervals for the $w_0w_a$CDM model, unless supernovae data are included. When DESI full-shape and BAO data are fit jointly, we obtain the following $1σ$ frequentist confidence intervals for $Λ$CDM ($w_0w_a$CDM): $σ_8 = 0.863^{+0.048}_{-0.040} , \ H_0 = 68.96^{+0.81}_{-0.80} \ \rm{km \ s^{-1}Mpc^{-1}} , \ Ω_{\rm{m}} = 0.3034\pm0.0110$ ($σ_8 = 0.782^{+0.060}_{-0.036} , \ H_0 = 63.7^{+4.2}_{-2.0} \ \rm{km \ s^{-1}Mpc^{-1}} , \ Ω_{\rm{m}} = 0.378^{+0.024}_{-0.047} , \ w_0 = -0.16^{+0.10}_{-0.50} , \ w_a = -3.0^{+1.7}_{}$), corresponding to 0.8$σ$, 0.3$σ$, 0.7$σ$ (2.1$σ$, 4.1$σ$, 6.5$σ$, 6.3$σ$, 6.6$σ$) shifts between the maximum likelihood estimate and the Bayesian posterior mean for $Λ$CDM ($w_0w_a$CDM) respectively.

astro-ph.CO

Alleviating prior dependencies for DESI DR1 clustering fits through reparameterization

Bayesian analyses of the full-shape clustering of Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) exhibit prior-volume projection effects, whereby weakly constrained nuisance parameters of the Effective Field Theory of Large Scale Structure (EFTofLSS) shift marginalized cosmological posteriors away from the posterior maximum. We reanalyze DESI DR1 power spectrum multipoles using two complementary mitigation strategies: (i) nonlinear orthogonalization to decorrelate nuisance and cosmological parameter priors, and (ii) a fully reparameterization-invariant Jeffreys prior over all EFTofLSS coefficients, evaluated on-the-fly via closed-form Jacobians. Including data from DESI, Big-Bang Nuclesynthesis and a constraint on $n_{\mathrm{s}}$, baseline priors lead to multi-$σ$ projection in the Hubble parameter $H_{0}$ and dark energy equation of state parameters $w_{0}$ and $w_{a}$; the Jeffreys prior successfully recenters these posteriors to enclose the maximum a posteriori estimate within the 68\% credible regions, demonstrating clear mitigation of projection effects for these late-time expansion parameters. A hybrid Jeffreys+baseline-Gaussian configuration controls residual over-broad tails in the physical cold dark matter density $ω_{\mathrm{c}}$ while preserving the volume correction, and is our favoured approach. We compare the credible intervals derived using our methodology to those obtained using Halo Occupation Distribution (HOD)-informed priors and to confidence intervals derived using frequentist profile likelihood analyses, finding agreement in both central values and degeneracy directions in the $w_{0}$--$w_{a}$ plane. This demonstrates that, once projection effects are properly controlled, we can make robust inferences about the late-time cosmological expansion independent of the statistical framework adopted.

astro-ph.CO

Finetuning Lightweight LLMs for Control Flow Graph Generation

Control Flow Graph (CFG) is an important program representations for software analysis, code understanding, and software maintenance. Traditional CFG generation techniques mainly rely on bytecode or abstract syntax trees. However, these approaches usually require complete, compilable, and syntax error-free code, which limits their applicability to incomplete or erroneous code. Furthermore, they often depend on language specific tools, making it difficult to support multiple programming languages in a unified manner. To address these limitations, this paper investigates the use of fine-tuned lightweight large language models (LLMs) for CFG generation. We first design a unified CFG output format and a task-specific fine-tuning prompt for CFG generation. Then, we construct a dataset based on an existing LeetCode dataset through automatic CFG generation and error augmentation. We evaluate the proposed approach on six lightweight LLM models, including three code-specific LLMs: CodeLlama, QwenCoder, and DeepSeekCoder; and three general purpose LLMs: Llama3.2-3B, Qwen-4B, and Phi-4B. The experimental results show that, through fine-tuning, lightweight LLMs achieve promising results for CFG generation, particularly when the input code is incomplete or erroneous. It also demonstrates cross-language generalization capability on programming language not included in the fine-tuning data.

cs.SE

Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided inquiry and reveals solutions directly. Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering, leaving the internal representation drift that precedes trajectory-level collapse largely unaddressed. We propose Scaffold-Preserving Representation Alignment, a two-stage framework that first warms up a Socratic tutor with supervised fine-tuning, then combines trajectory-weighted direct preference optimization with a margin-preserving representation loss anchored to frozen reference states. Our method is designed to maintain separation between scaffold-preserving and collapse-inducing hidden states across dialogue turns. We evaluate our method across five STEM disciplines and five red-teaming attack strategies. On Qwen3-8B, our method lowers Collapse Rate to 32%, delays average collapse onset beyond nine turns, and keeps over-refusal low, suggesting that representation-level alignment can improve the robustness of long-horizon Socratic tutoring under our red-teaming protocol.

cs.AI

Nanoscopic Multiplexing Optical Data Storage via Chip Fabrication

The accelerating growth of global data generation demands data storage platforms that offer high capacity, long lifespan, and low energy consumption beyond the limits of electronic memory technologies. Optical storage provides an attractive alternative. However, its density is fundamentally constrained by the optical diffraction limit and the limited scalability from the point-by-point laser writing, as well as thermal accumulation during high-speed writing. Here, we introduce a large-scale optical data storage scheme that is compatible with the progress in chip fabrication by combining electron-beam lithography (EBL) and ion implantation to deterministically encode high-density data. The approach achieves precise control of ion number and spatial distribution, enabling multi-bit grayscale encoding and wavelength division multiplexing with chip-scale patterning over millimeter areas. Wavelength-selective readout is performed using downconversion and upconversion fluorescence detection, allowing crosstalk-free retrieval of multiplexed data channels. We further develop a neural network-based super-resolution algorithm that reconstructs data beyond the diffraction limit, further increasing the effective storage density. Using this integrated framework, we achieve an optical data density of 10 Gbit/cm$^2$ with high fidelity. Our results establish a micro/nano-fabrication-compatible route to large-scale, high-density optical memory and provide a foundation for next-generation cold data optical storage technologies.

physics.optics

Modeling Behavioral Intensity and Transitions for Generative Recommendation

Multi-behavior recommendation aims to predict user conversions by modeling various interaction types that carry distinct intent signals. Recently, generative sequence modeling methods have emerged as an important paradigm for multi-behavior recommendation by achieving flexible sequence generation. However, existing generative methods typically treat behaviors as auxiliary token features and feed them into unified attention mechanisms. These models implicitly assume uniform activation of dependencies among historical behaviors, thereby failing to discern differences in intensity or capture transition patterns. To address these limitations, we propose BITRec, a novel generative multi-behavior recommendation framework that introduces structured behavioral modeling through selective dependency activation. BITRec incorporates (i) Hierarchical Behavior Aggregation (HBA), which explicitly models behavioral intensity differences through separated exploration and commitment pathways, and (ii) Transition Relation Encoding (TRE), which encodes transition structures through explicit learnable relation matrices. Experiments on four large-scale datasets (RetailRocket, Taobao, Tmall, Insurance Dataset) with millions of interactions achieve consistent improvements of 15-23% across multiple metrics, with peak gains of 22.79% MRR on Tmall and 17.83% HR@10, 17.55% NDCG@10 on Taobao.

cs.IR

SACS: A Code Smell Dataset using Semi-automatic Generation Approach

Code smell is a great challenge in software refactoring, which indicates latent design or implementation flaws that may degrade the software maintainability and evolution. Over the past of decades, the research on code smell has received extensive attention. Especially the researches applied machine learning-technique have become a popular topic in recent studies. However, one of the biggest challenges to apply machine learning-technique is the lack of high-quality code smell datasets. Manually constructing such datasets is extremely labor-intensive, as identifying code smells requires substantial development expertise and considerable time investment. In contrast, automatically generated datasets, while scalable, frequently exhibit reduced label reliability and compromised data quality. To overcome this challenge, in this study, we explore a semi-automatic approach to generate a code smell dataset with high quality data samples. Specifically, we first applied a set of automatic generation rules to produce candidate smelly samples. We then employed multiple metrics to group the data samples into an automatically accepted group and a manually reviewed group, enabling reviewers to concentrate their efforts on ambiguous samples. Furthermore, we established structured review guidelines and developed a annotation tool to support the manual validation process. Based on the proposed semi-automatic generation approach, we created an open-source code smell dataset, SACS, covering three widely studied code smells: Long Method, Large Class, and Feature Envy. Each code smell category includes over 10,000 labeled samples. This dataset could provide a large-scale and publicly available benchmark to facilitate future studies on code smell detection and automated refactoring.

cs.SE

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive most fine-grained latent features for safety, a low-frequency concept domain, remains unexplored. Two key challenges exist: identifying SAEs with the greatest potential for generating safety domain-specific features, and the prohibitively high cost of detailed feature explanation. In this paper, we propose Safe-SAIL, a unified framework for interpreting SAE features in safety-critical domains to advance mechanistic understanding of large language models. Safe-SAIL introduces a pre-explanation evaluation metric to efficiently identify SAEs with strong safety domain-specific interpretability, and reduces interpretation cost by 55% through a segment-level simulation strategy. Building on Safe-SAIL, we train a comprehensive suite of SAEs with human-readable explanations and systematic evaluations for 1,758 safety-related features spanning four domains: pornography, politics, violence, and terror. Using this resource, we conduct empirical analyses and provide insights on the effectiveness of Safe-SAIL for risk feature identification and how safety-critical entities and concepts are encoded across model layers. All models, explanations, and tools are publicly released in our open-source toolkit and companion product.

cs.LG

Nonlinear Information from DESI Luminous Red Galaxies: An Emulator-Based Analysis of Pre- and Post-Reconstruction Power Spectra

We present joint measurements of the pre- and post-reconstruction power spectra, $P_{\rm pre}$ and $P_{\rm post}$, together with their cross-power spectrum, $P_{\rm cross}$, for the Luminous Red Galaxies (LRGs) in the DESI Data Release 1 (DR1). We jointly analyse these observables with an emulator-based full-shape modeling framework, thereby, for the first time, we extract complementary nonlinear information from the galaxy density field before and after reconstruction in real survey data. Specifically, including $P_{\rm post}$ and $P_{\rm cross}$ in addition to $P_{\rm pre}$ (hereafter $P_{\rm all}$) yields an improvement of approximately $18$-$27\%$ in the $σ_8$ constraint in both $Λ$CDM and $w$CDM, depending on the redshift bin, relative to the $P_{\rm pre}$-only analysis with the cosmic microwave background distance priors (hereafter CMB). In $w$CDM, the joint CMB+$P_{\rm all}$ analysis can tighten the constraints on $w$ by approximately $5$-$15\%$ across the two LRG redshift bins, compared to the CMB+$P_{\rm pre}$ combination. Further incorporating the Type Ia supernova dataset and comparing the cosmological constraints in $w$CDM from each individual power-spectrum component with those from the full combination, we find that $P_{\rm all}$ consistently provides the tightest constraints. From the joint CMB+$P_{\rm all}$+DES-Dovekie dataset, we obtain $Ω_m = 0.314 \pm 0.0048$ and $w = -0.988 \pm 0.023$ for the \texttt{LRG1} sample, and $Ω_m = 0.318 \pm 0.0046$ and $w = -0.988 \pm 0.025$ for \texttt{LRG2}. These results demonstrate that combining pre- and post-reconstruction power spectra with their cross-correlation enables DESI to harvest additional nonlinear information, leading to tighter constraints on cosmological parameters.

astro-ph.CO

GLM-5: from Vibe Coding to Agentic Engineering

We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous reinforcement learning infrastructure that drastically improves post-training efficiency by decoupling generation from training. Furthermore, we propose novel asynchronous agent RL algorithms that further improve RL quality, enabling the model to learn from complex, long-horizon interactions more effectively. Through these innovations, GLM-5 achieves state-of-the-art performance on major open benchmarks. Most critically, GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges. Code, models, and more information are available at https://github.com/zai-org/GLM-5.

cs.LG

Deception at Scale: Deceptive Designs in 1K LLM-Generated Ecommerce Components

Recent work has shown that front-end code generated by Large Language Models (LLMs) can embed deceptive designs. To assess the magnitude of this problem, identify the factors that influence deceptive design production, and test strategies for reducing deceptive designs, we carried out two studies which generated and analyzed 1,296 LLM-generated web components, along with a design rationale for each. The first study tested four LLMs for 15 common ecommerce components. Overall 55.8% of components contained at least one deceptive design, and 30.6% contained two or more. Occurence varied significantly across models, with DeepSeek-V3 producing the fewest. Interface interference emerged as the dominant strategy, using color psychology to influence actions and hiding essential information. The first study found that prompts emphasizing business interests (e.g., increasing sales) significantly increased deceptive designs, so a second study tested a variety of prompting strategies to reduce their frequency, finding a values-centered approach the most effective. Our findings highlight risks in using LLMs for coding and offer recommendations for LLM developers and providers.

cs.HC

Tabular Incremental Inference

Tabular data is a fundamental form of data structure. The evolution of table analysis tools reflects humanity's continuous progress in data acquisition, management, and processing. The dynamic changes in table columns arise from technological advancements, changing needs, data integration, etc. However, the standard process of training AI models on tables with fixed columns and then performing inference is not suitable for handling dynamically changed tables. Therefore, new methods are needed for efficiently handling such tables in an unsupervised manner. In this paper, we introduce a new task, Tabular Incremental Inference (TabII), which aims to enable trained models to incorporate new columns during the inference stage, enhancing the practicality of AI models in scenarios where tables are dynamically changed. Furthermore, we demonstrate that this new task can be framed as an optimization problem based on the information bottleneck theory, which emphasizes that the key to an ideal tabular incremental inference approach lies in minimizing mutual information between tabular data and representation while maximizing between representation and task labels. Under this guidance, we design a TabII method with Large Language Model placeholders and Pretrained TabAdapter to provide external knowledge and Incremental Sample Condensation blocks to condense the task-relevant information given by incremental column attributes. Experimental results across eight public datasets show that TabII effectively utilizes incremental attributes, achieving state-of-the-art performance.

cs.AI

The Forecast Critic: Leveraging Large Language Models for Poor Forecast Identification

Monitoring forecasting systems is critical for customer satisfaction, profitability, and operational efficiency in large-scale retail businesses. We propose The Forecast Critic, a system that leverages Large Language Models (LLMs) for automated forecast monitoring, taking advantage of their broad world knowledge and strong ``reasoning'' capabilities. As a prerequisite for this, we systematically evaluate the ability of LLMs to assess time series forecast quality, focusing on three key questions. (1) Can LLMs be deployed to perform forecast monitoring and identify obviously unreasonable forecasts? (2) Can LLMs effectively incorporate unstructured exogenous features to assess what a reasonable forecast looks like? (3) How does performance vary across model sizes and reasoning capabilities, measured across state-of-the-art LLMs? We present three experiments, including on both synthetic and real-world forecasting data. Our results show that LLMs can reliably detect and critique poor forecasts, such as those plagued by temporal misalignment, trend inconsistencies, and spike errors. The best-performing model we evaluated achieves an F1 score of 0.88, somewhat below human-level performance (F1 score: 0.97). We also demonstrate that multi-modal LLMs can effectively incorporate unstructured contextual signals to refine their assessment of the forecast. Models correctly identify missing or spurious promotional spikes when provided with historical context about past promotions (F1 score: 0.84). Lastly, we demonstrate that these techniques succeed in identifying inaccurate forecasts on the real-world M5 time series dataset, with unreasonable forecasts having an sCRPS at least 10% higher than that of reasonable forecasts. These findings suggest that LLMs, even without domain-specific fine-tuning, may provide a viable and scalable option for automated forecast monitoring and evaluation.

cs.AI

What Galaxy Clusters Have to Say About Dynamical Dark Energy and $H_0$

We show that, in flat $Λ$CDM, low-redshift structure probes -- cluster abundances, 3$\times$2-point analyses, and full-shape clustering -- are mutually consistent, jointly delivering precise constraints on $σ_8$ and $Ω_{\rm m}$ that agree with geometrical datasets (CMB+BAO+SN). In $w_0w_a$CDM, adding clusters to the geometry dataset reduces the evidence for evolving dark energy while relaxing the $H_0$ tension, suggesting a $Λ$CDM evolution of the late-time Universe and a sound horizon that differs from its standard value.

astro-ph.CO

Baryon fraction from the BAO amplitude: a consistent approach to parameterizing perturbation growth

Galaxy clustering constrains the baryon fraction Omega_b/Omega_m through the amplitude of baryon acoustic oscillations and the suppression of perturbations entering the horizon before recombination. This produces a different pre-recombination distribution of baryons and dark matter. After recombination, the gravitational potential responds to both components in proportion to their mass, allowing robust measurement of the baryon fraction. This is independent of new-physics scenarios altering the recombination background (e.g. Early Dark Energy). The accuracy of such measurements does, however, depend on how baryons and CDM are modeled in the power spectrum. Previous template-based splitting relied on approximate transfer functions that neglected part of information. We present a new method that embeds an extra parameter controlling the balance between baryons and dark matter in the growth terms of the perturbation equations in the CAMB Boltzmann solver. This approach captures the baryonic suppression of CDM prior to recombination, avoids inconsistencies, and yields a clean parametrization of the baryon fraction in the linear power spectrum, separating out the simple physics of growth due to the combined matter potential. We implement this framework in an analysis pipeline using Effective Field Theory of Large-Scale Structure with HOD-informed priors and validate it against noiseless LCDM and EDE cosmologies with DESI-like errors. The new scheme achieves comparable precision to previous splitting while reducing systematic biases, providing a more robust way to baryon-fraction measurements. In combination with BBN constraints on the baryon density and Alcock-Paczynski estimates of the matter density, these results strengthen the use of baryon fraction measurements to derive a Hubble constant from energy densities, with future DESI and Euclid data expected to deliver competitive constraints.

astro-ph.CO