SearcharxivSearch

arXiv subjects

Hao He

Publications and source records attributed to Hao He.

At least 37 records · Page 2Linked to original sources

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this via post-training, they typically rely on a bidirectional teacher model or discriminator. To achieve an end-to-end solution, we introduce Resampling Forcing, a teacher-free framework that enables training autoregressive video models from scratch and at scale. Central to our approach is a self-resampling scheme that simulates inference-time model errors on history frames during training. Conditioned on these degraded histories, a sparse causal mask enforces temporal causality while enabling parallel training with frame-level diffusion loss. To facilitate efficient long-horizon generation, we further introduce history routing, a parameter-free mechanism that dynamically retrieves the top-k most relevant history frames for each query. Experiments demonstrate that our approach achieves performance comparable to distillation-based baselines while exhibiting superior temporal consistency on longer videos owing to native-length training.

cs.CV

Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects

Large language models (LLMs) have demonstrated the promise to revolutionize the field of software engineering. Among other things, LLM agents are rapidly gaining momentum in software development, with practitioners reporting a multifold increase in productivity after adoption. Yet, empirical evidence is lacking around these claims. In this paper, we estimate the causal effect of adopting a widely popular LLM agent assistant, namely Cursor, on development velocity and software quality. The estimation is enabled by a state-of-the-art difference-in-differences design comparing Cursor-adopting GitHub projects with a matched control group of similar GitHub projects that do not use Cursor. We find that the adoption of Cursor leads to a statistically significant, large, but transient increase in project-level development velocity, along with a substantial and persistent increase in static analysis warnings and code complexity. Further panel generalized-method-of-moments estimation reveals that increases in static analysis warnings and code complexity are major factors driving long-term velocity slowdown. Our study identifies quality assurance as a major bottleneck for early Cursor adopters and calls for it to be a first-class citizen in the design of agentic AI coding tools and AI-driven workflows.

cs.SE

Recursive Inverse Design Enables Hyper-spectral Photonic Integrated Circuits

Spectrum manipulation is central to photonic systems, where advanced computing and sensing applications often demand highly complex spectral responses to achieve high throughput. Conventional methods for enhancing spectral complexity typically rely on cascading discrete photonic components, resulting in a complexity that scales only linearly with the number of components. Here, we introduce hyper-spectral photonic integrated circuits (HS-PICs), in which spectral complexity scales exponentially with the number of components. This is achieved through recursive inverse design - a system-level inverse design strategy that exploits intricate inter-component interactions as design freedoms, thereby substantially expanding the design space for spectral engineering. Using this approach, we demonstrate that even a single waveguide structure can resolve spectra with sub-picometer resolution, surpassing the performance of current state-of-the-art spectrometers. This performance bridges optical and microwave frequencies in spectral analysis, enabling simultaneous monitoring of optical and radio signals within a single device. Our work establishes a transformative framework for next-generation computing and sensing technologies.

physics.optics

Surveying the Whirlpool at Arcseconds with NOEMA (SWAN): III. $^{13}$CO/C$^{18}$O ratio variations across the M51 galaxy

CO isotopologues are common tracers of the bulk molecular gas in extragalactic studies, providing insights into the physical and chemical conditions of the cold molecular gas, a reservoir for star formation. Since star formation occurs within molecular clouds, mapping CO isotopologues at cloud-scale is important to understanding the processes driving star formation. However, achieving this mapping at such scales is challenging and time-intensive. The Surveying the Whirlpool Galaxy at Arcseconds with NOEMA (SWAN) survey addresses this by using the Institut de radioastronomie millim\'etrique (IRAM) NOrthern Extended Millimeter Array (NOEMA) to map the $^{13}$CO(1-0) and C$^{18}$O(1-0) isotopologues, alongside several dense gas tracers, in the nearby star-forming galaxy M51 at high sensitivity and spatial resolution ($\approx$ 125 pc).We examine the $^{13}$CO(1-0) to C$^{18}$O(1-0) line emission ratio as a function of galactocentric radius and star formation rate surface density to infer how different chemical and physical processes affect this ratio at cloud scales across different galactic environments: nuclear bar, molecular ring, northern and southern spiral arms. In line with previous studies conducted at kiloparsec scales for nearby star-forming galaxies, we find a moderate positive correlation with galactocentric radius and a moderate negative correlation with star formation rate surface density across the field-of-view (FoV), with slight variations depending on the galactic environment. We propose that selective nucleosynthesis and changes in the opacity of the gas are the primary drivers of the observed variations in the ratio.

astro-ph.GA

An inexact variable metric proximal linearization method for composite optimization on manifolds

This paper concerns the minimization of the composition of a nonsmooth convex function and a $\mathcal{C}^{1,1}$ mapping $F$ over a $\mathcal{C}^2$-smooth embedded closed submanifold $\mathcal{M}$. For this class of nonconvex and nonsmooth problems, we propose an inexact variable metric proximal linearization method by leveraging its composite structure and the retraction and first-order information of $\mathcal{M}$, which at each iteration seeks an inexact solution to a subspace constrained strongly convex problem by a practical inexactness criterion. Under the boundedness assumption on the iterate sequence, we establish the $O(\epsilon^{-3})$ oracle complexity with a dual fast gradient method as the inner solver, and prove that any cluster point of the iterate sequence is a stationary point. If in addition the constructed potential function has the Kurdyka-Lojasiewicz (KL) property on the set of cluster points, the iterate sequence converges to a stationary point, and if the potential function has the KL property of exponent $q\in[\frac{1}{2},1)$, the local convergence rate is characterized. We also provide a condition only involving the original data to identify the KL property of the potential function with an exponent $q\in[0,1)$. Numerical comparisons with the existing methods validate the efficiency of the proposed method.

math.OC

Multiple quantum spin Hall states and topological current divider in Twisted Bilayer WSe$_2$

It has been demonstrated that topological quantum spin Hall (QSH) state exist in twisted bilayers of transition metal dichalcogenides. However, a comprehensive theoretical characterization of the topological edge states remains a topic of interest and an unresolved issue. Here, the topological transport properties of the twisted WSe$_2$ bilayers are investigated. Beyond the conventional single QSH, we identify emergent double and quartuple quantum spin Hall states, hosting two and four pairs of counter-propagating helical edge channels respectively. Furthermore, the charge carriers in these edge states are not localized at edge but rather the high potential point of the moire superlattice boundary, undergoing interlayer transitions and propagating forward continuously. We term these edge states as moire edge states. These edge states can survive in non-magnetic disorder, with the robustness of double QSH states surpassing that of single QSH states. At a twisting angle of 2.45$^\circ$, the transition between the single and double QSH states can be achieved by adjusting the gate on the surface. Based on this, we propose a five-terminal device to as a topological current devider. Our findings provide support for the development of dissipationless spintronics.

cond-mat.mes-hall

The Hierarchical Dynamical State of Molecular Gas from 3 to 300 pc in NGC 253

Understanding how the dynamical state of the interstellar medium (ISM) changes across spatial scales can provide important insights into how the gas is organized and ultimately collapses to form stars. To this end, we present ALMA $^{12}\mathrm{CO}(2-1)$ observations at $7$ pc ($0''.4$) spatial resolution across a $1.4~\mathrm{kpc}\times5.6~\mathrm{kpc}$ ($1'.3\times1'.3$) region located in the disk of the nearby ($D = 3.5$ Mpc), massive, star-forming galaxy NGC 253. We decompose this emission with a hierarchical, multiscale dendrogram algorithm to identify 2463 structures with deconvolved sizes ranging from $\sim3$ to $300$ pc, complete to a limiting mass of $10^4~M_\odot$. By comparing the virial parameter of these structures against physical properties including size, mass, surface density, velocity dispersion, and hierarchical position, we carry out a comprehensive search for a preferred scale at which gravitationally bound structures emerge. Ultimately, we do not identify evidence of an emergent scale for bound objects in our data, nor do we find a significant correlation between the virial parameter and structure sizes. These findings suggest that simple observational estimates of gravitational binding cannot be used to define molecular clouds and emphasize the need for multiscale approaches to characterize the ISM.

astro-ph.GA

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large language model's (LLM) vocabulary. By integrating vision and text into a unified space with an expanded vocabulary, our multimodal LLM, Tar, enables cross-modal input and output through a shared interface, without the need for modality-specific designs. Additionally, we propose scale-adaptive encoding and decoding to balance efficiency and visual detail, along with a generative de-tokenizer to produce high-fidelity visual outputs. To address diverse decoding needs, we utilize two complementary de-tokenizers: a fast autoregressive model and a diffusion-based model. To enhance modality fusion, we investigate advanced pre-training tasks, demonstrating improvements in both visual understanding and generation. Experiments across benchmarks show that Tar matches or surpasses existing multimodal LLM methods, achieving faster convergence and greater training efficiency. Code, models, and data are available at https://tar.csuhan.com

cs.CV

Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation

Existing large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to transform a pre-trained latent video diffusion model into a real-time, interactive video generator. Our model autoregressively generates a latent frame at a time using a single neural function evaluation (1NFE). The model can stream the result to the user in real time and receive interactive responses as controls to generate the next latent frame. Unlike existing approaches, our method explores adversarial training as an effective paradigm for autoregressive generation. This not only allows us to design an architecture that is more efficient for one-step generation while fully utilizing the KV cache, but also enables training the model in a student-forcing manner that proves to be effective in reducing error accumulation during long video generation. Our experiments demonstrate that our 8B model achieves real-time, 24fps, streaming video generation at 736x416 resolution on a single H100, or 1280x720 on 8xH100 up to a minute long (1440 frames). Visit our research website at https://seaweed-apt.com/2

cs.CV

Constraining resolved extragalactic $R_{21}$ variation with well calibrated ALMA observations

CO(1-0) and CO(2-1) are commonly used as bulk molecular gas tracers. The CO line ratios (especially CO(2-1)/CO(1-0) - $R_{21}$) vary within and among galaxies, yet previous studies on $R_{21}$ and alike often rely on measurements constructed by combining data from facilities with substantial relative calibration uncertainties that have the same order as physical line ratio variations. Hence robustly determining systematic $R_{21}$ variations is challenging. Here, we compare CO(1-0) and CO(2-1) mapping data from ALMA for 14 nearby galaxies, at a common physical resolution of 1.7 kpc. Our dataset includes new ALMA (7m+TP) CO(1-0) maps of 12 galaxies. We investigate $R_{21}$ variation to understand its dependence on global galaxy properties, kpc-scale environmental factors, and its correlation with star formation rate (SFR) surface density and metallicity. We find that the galaxy-to-galaxy scatter is 0.05 dex. This is lower than previous studies which reported over 0.1 dex variation, likely reflecting significant flux calibration uncertainties in single-dish surveys. Within individual galaxies, $R_{21}$ has a typical mean value of ~0.64 and 0.1 dex variation, with an increase to ~0.75 towards galactic centers. We find strong correlations between $R_{21}$ and various galactic parameters, particularly SFR surface density, which shows a power-law slope of 0.10-0.11 depending on the adopted binning/fitting methods. Our findings suggest that, for studies covering main sequence galaxy samples, assuming a fixed $R_{21}$=0.64 does not significantly bias kpc-scale molecular gas mass estimates from CO(2-1). Instead, systematic uncertainties from flux calibration and the CO-to-H$_2$ conversion factor account for more systematic scatter of CO-derived molecular gas properties.

astro-ph.GA

Communication Efficient Multiparty Private Set Intersection from Multi-Point Sequential OPRF

Multiparty private set intersection (MPSI) allows multiple participants to compute the intersection of their locally owned data sets without revealing them. MPSI protocols can be categorized based on the network topology of nodes, with the star, mesh, and ring topologies being the primary types, respectively. Given that star and mesh topologies dominate current implementations, most existing MPSI protocols are based on these two topologies. However, star-topology MPSI protocols suffer from high leader node load, while mesh topology protocols suffer from high communication complexity and overhead. In this paper, we first propose a multi-point sequential oblivious pseudorandom function (MP-SOPRF) in a multi-party setting. Based on MP-SOPRF, we then develop an MPSI protocol with a ring topology, addressing the challenges of communication and computational overhead in existing protocols. We prove that our MPSI protocol is semi-honest secure under the Hamming correlation robustness assumption. Our experiments demonstrate that our MPSI protocol outperforms state-of-the-art protocols, achieving a reduction of 74.8% in communication and a 6% to 287% improvement in computational efficiency.

cs.CR

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents

In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectively. The reward model, UI-Genie-RM, features an image-text interleaved architecture that efficiently pro- cesses historical context and unifies action-level and task-level rewards. To sup- port the training of UI-Genie-RM, we develop deliberately-designed data genera- tion strategies including rule-based verification, controlled trajectory corruption, and hard negative mining. To address the second challenge, a self-improvement pipeline progressively expands solvable complex GUI tasks by enhancing both the agent and reward models through reward-guided exploration and outcome verification in dynamic environments. For training the model, we generate UI- Genie-RM-517k and UI-Genie-Agent-16k, establishing the first reward-specific dataset for GUI agents while demonstrating high-quality synthetic trajectory gen- eration without manual annotation. Experimental results show that UI-Genie achieves state-of-the-art performance across multiple GUI agent benchmarks with three generations of data-model self-improvement. We open-source our complete framework implementation and generated datasets to facilitate further research in https://github.com/Euphoria16/UI-Genie.

cs.CL

Relationships between PAHs, Small Dust Grains, H$_2$, and HI in Local Group Dwarf Galaxies NGC 6822 and WLM Using JWST, ALMA, and the VLA

We present 0.7-3.3 pc resolution mid-infrared (MIR) JWST images at 7.7 $\mu$m (F770W) and 21 $\mu$m (F2100W) covering the main star-forming regions of two of the closest star-forming low-metallicity dwarf galaxies, NGC6822 and Wolf-Lundmark-Melotte (WLM). The images of NGC6822 reveal filaments, edge-brightened bubbles, diffuse emission, and a plethora of point sources. By contrast, most of the MIR emission in WLM is point-like, with a small amount of extended emission. Compared to solar metallicity galaxies, the ratio of 7.7 $\mu$m intensity ($I_\nu^{F770W}$), tracing polycyclic aromatic hydrocarbons (PAHs), to 21 $\mu$m intensity ($I_\nu^{F2100W}$), tracing small, warm dust grain emission, is suppressed in these low-metallicity dwarfs. Using ALMA CO(2-1) observations, we find that detected CO intensity versus $I_\nu^{F770W}$ at ~2 pc resolution in dwarfs follows a similar relationship to that at solar metallicity and lower resolution, while the CO versus $I_\nu^{F2100W}$ relationship in dwarfs lies significantly below that derived from solar metallicity galaxies at lower resolution, suggesting more pronounced destruction of CO molecules at low metallicity. Finally, adding in Local Group L-Band Survey VLA 21 cm HI observations, we find that $I_\nu^{F2100W}$ and $I_\nu^{F770W}$ vs. total gas ratios are suppressed in NGC6822 and WLM compared to solar metallicity galaxies. In agreement with dust models, the level of suppression appears to be at least partly accounted for by the reduced galaxy-averaged dust-to-gas and PAH-to-dust mass ratios in the dwarfs. Remaining differences are likely due to spatial variations in dust model parameters, which should be an exciting direction for future work in local dwarf galaxies.

astro-ph.GA

Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities

Recent advancements in Large Reasoning Models (LRMs), such as OpenAI's o1/o3 and DeepSeek-R1, have demonstrated remarkable performance in specialized reasoning tasks through human-like deliberative thinking and long chain-of-thought reasoning. However, our systematic evaluation across various model families (DeepSeek, Qwen, and LLaMA) and scales (7B to 32B) reveals that acquiring these deliberative reasoning capabilities significantly reduces the foundational capabilities of LRMs, including notable declines in helpfulness and harmlessness, alongside substantially increased inference costs. Importantly, we demonstrate that adaptive reasoning -- employing modes like Zero-Thinking, Less-Thinking, and Summary-Thinking -- can effectively alleviate these drawbacks. Our empirical insights underline the critical need for developing more versatile LRMs capable of dynamically allocating inference-time compute according to specific task characteristics.

cs.AI

CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos with large camera movement. We take an approach that progressively expands the generation of dynamic scenes -- first enhancing dynamic content within individual video clip, then extending this capability to create seamless explorations across broad viewpoint ranges. Specifically, we construct a dataset featuring a large degree of dynamics with camera parameter annotations for training while designing a lightweight camera injection module and training scheme to preserve dynamics of the pretrained models. Building on these improved single-clip techniques, we enable extended scene exploration by allowing users to iteratively specify camera trajectories for generating coherent video sequences. Experiments across diverse scenarios demonstrate that CameraCtrl Ii enables camera-controlled dynamic scene synthesis with substantially wider spatial exploration than previous approaches.

cs.CV

X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second

Sparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity of large-scale training datasets and the absence of direct and consistent 3D representations. In this paper, we propose an X-ray Large Reconstruction Model (X-LRM) for extremely sparse-view ($<$10 views) CT reconstruction. X-LRM consists of two key components: X-former and X-triplane. X-former can handle an arbitrary number of input views using an MLP-based image tokenizer and a Transformer-based encoder. The output tokens are then upsampled into our X-triplane representation, which models the 3D radiodensity as an implicit neural field. To support the training of X-LRM, we introduce Torso-16K, a large-scale dataset comprising over 16K volume-projection pairs of various torso organs. Extensive experiments demonstrate that X-LRM outperforms the state-of-the-art method by 1.5 dB and achieves 27$\times$ faster speed with better flexibility. Furthermore, the evaluation of lung segmentation tasks also suggests the practical value of our approach. Our code and dataset will be released at https://github.com/Richard-Guofeng-Zhang/X-LRM

eess.IV

Performance Evaluation of Large Language Models in Statistical Programming

The programming capabilities of large language models (LLMs) have revolutionized automatic code generation and opened new avenues for automatic statistical analysis. However, the validity and quality of these generated codes need to be systematically evaluated before they can be widely adopted. Despite their growing prominence, a comprehensive evaluation of statistical code generated by LLMs remains scarce in the literature. In this paper, we assess the performance of LLMs, including two versions of ChatGPT and one version of Llama, in the domain of SAS programming for statistical analysis. Our study utilizes a set of statistical analysis tasks encompassing diverse statistical topics and datasets. Each task includes a problem description, dataset information, and human-verified SAS code. We conduct a comprehensive assessment of the quality of SAS code generated by LLMs through human expert evaluation based on correctness, effectiveness, readability, executability, and the accuracy of output results. The analysis of rating scores reveals that while LLMs demonstrate usefulness in generating syntactically correct code, they struggle with tasks requiring deep domain understanding and may produce redundant or incorrect results. This study offers valuable insights into the capabilities and limitations of LLMs in statistical programming, providing guidance for future advancements in AI-assisted coding systems for statistical analysis.

stat.AP

Pinning Is Futile: You Need More Than Local Dependency Versioning to Defend against Supply Chain Attacks

Recent high-profile incidents in open-source software have greatly raised practitioner attention on software supply chain attacks. To guard against potential malicious package updates, security practitioners advocate pinning dependency to specific versions rather than floating in version ranges. However, it remains controversial whether pinning carries a meaningful security benefit that outweighs the cost of maintaining outdated and possibly vulnerable dependencies. In this paper, we quantify, through counterfactual analysis and simulations, the security and maintenance impact of version constraints in the npm ecosystem. By simulating dependency resolutions over historical time points, we find that pinning direct dependencies not only (as expected) increases the cost of maintaining vulnerable and outdated dependencies, but also (surprisingly) even increases the risk of exposure to malicious package updates in larger dependency graphs due to the specifics of npm's dependency resolution mechanism. Finally, we explore collective pinning strategies to secure the ecosystem against supply chain attacks, suggesting specific changes to npm to enable such interventions. Our study provides guidance for practitioners and tool designers to manage their supply chains more securely.

cs.SE