SearcharxivSearch

arXiv subjects

Junyi Zhang

Publications and source records attributed to Junyi Zhang.

At least 19 recordsLinked to original sources

Statistical Study of Solar Prominence Plumes Based on NVST H$\alpha$ Observations

Plumes are one of the most representative dynamic features observed in prominences and play a key role in mass and magnetic transport within them. However, their physical nature and triggering processes remain actively debated. Based on limb H$\alpha$ observations from the New Vacuum Solar Telescope (NVST) during 2013--2025, we statistically investigated 34 plumes with clear and complete evolutions by developing an automated image-processing pipeline. It is revealed that plume lifetimes mainly range from 300 s to 700 s, with vertical displacements between 3--7 Mm. The mean widths and velocities are concentrated in the range of 0.5--1.5 Mm and 10--20 km s$^{-1}$, respectively. Besides wide distribution ranges, plume parameters exhibit irregular evolution fluctuations, indicating that the formation and evolution of various plumes may exhibit different physical patterns. Correlation analysis among the parameters further reveals that: (1) Positive correlations were found among lifetime, vertical displacement, and mean width, indicating an intrinsic coupling between the temporal and spatial scales of plumes. (2) Trajectory curvature is negatively correlated with lifetime, vertical displacement, and velocity. Accelerating and width-contracting plumes typically have lower curvature, suggesting that curvature may reflect environmental influences and the stability of plumes. (3) Plumes with higher initial velocities were more likely to be accompanied by precursor brightening, suggesting that these plumes may be triggered by magnetic reconnection. Furthermore, we infer that some plumes in non-bubble regions may be inherently driven by mini-filament eruptions. These results establish a statistical framework for prominence plumes and reveal diversity in their dynamical evolution and triggering mechanisms.

astro-ph.SR

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, under-specified, and involve multiple layers of abstraction. We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning. In this position paper, we provide a systematic review of the field, covering datasets, auto-formalization, and proof synthesis. More importantly, we identify core limitations of existing systems in serving as mathematical research agents, examining issues across datasets, relational structure, mathematical exploration, tool ecosystem, and human-AI collaboration, outlining a strategic road-map for the future of AI4Math.

cs.CL

Ferguson's Dirichlet Process Breakthrough: A Lasting Legacy

Ferguson's 1973 introduction of the Dirichlet process marked a breakthrough in Bayesian nonparametric statistics. For the first time, a prior on the space of probability measures fulfilled two key desiderata: large support and analytical tractability. In this paper, we review three complementary constructions of the Dirichlet process, whose roots can be traced back to Ferguson: through finite-dimensional distributions, via normalization of a gamma process, and through predictive distributions. Each perspective not only deepens the understanding of the Dirichlet process but also provides a template for generalizations, from normalized random measures with independent increments to Gibbs--type priors and beyond. Over the past fifty years, the Dirichlet process has become the cornerstone of Bayesian nonparametric methodology and applications, while simultaneously inspiring the expansion of the landscape of nonparametric priors. Since de Finetti laid out the Bayesian nonparametric framework in the 1930s, the key obstacle had been the absence of a tractable nonparametric prior. Ferguson's contribution overcame this challenge, providing a solution to a decades-long open problem. In recognition of this decisive advance, it seems appropriate to refer to the Dirichlet process as the Ferguson--Dirichlet process.

math.ST

Playful Agentic Robot Learning

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and MolmoSpaces show that play-learned skills improve held-out downstream tasks over no-play and random-play baselines, with 20.6 and 17.0 percentage-point gains over CaP-Agent0 on LIBERO-PRO and MolmoSpaces, respectively. Moreover, the learned skills can be plugged into other inference-time Code-as-Policy agents by simply retrieving them into the context, improving RoboSuite and real-world transfer by 8.9 and 8.8 points, respectively, without finetuning the underlying model.

cs.RO

T-Rex: Tactile-Reactive Dexterous Manipulation

The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) models for robotic manipulation generally either overlook the tactile modality or are limited to encoders with static cues, due in part to the scarcity of diverse training data and standardized evaluation, architectural constraints in current VLA models, and limitations of static tactile encoders. In this paper, we push the frontier of tactile-reactive manipulation by addressing all of these limitations. We propose a large-scale, 100-hour tactile-rich dataset collected via a novel, data-efficient recipe that prioritizes elementary motor primitives. To effectively exploit naturally high-frequency touch signals without sacrificing the existing capabilities of existing VLAs, we introduce a variable-rate Mixture-of-Transformers (MoT) architecture equipped with a novel temporal tactile VQ-VAE encoder. We demonstrate the effectiveness of tactile-reactive policies on 12 manipulation tasks requiring delicate force control and deformable object manipulation, achieving over 30% higher average success rate than the strongest baseline.

cs.RO

Cambrian-P: Pose-Grounded Video Understanding

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) for video understanding, which process frames as isolated 2D snapshots, instead of the persistent scene humans perceive. We revisit pose as a lightweight supervisory signal and introduce Cambrian-P, a video MLLM augmented with per-frame learnable camera tokens and a pose regression head. With a carefully designed sampling scheme, the model achieves substantial gains of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench, generalizes across eight additional spatial and general video QA benchmarks, and, as a byproduct, achieves state of the art streaming pose estimation on ScanNet. Surprisingly, training on pseudo-annotated poses from in-the-wild video further improves general video QA benchmarks, showing pose helps beyond spatial reasoning. Together, these results position camera pose as a fundamental signal for video models that reason about the physical world.

cs.CV

Can We Distinguish the Source Region Location of Filament/Prominence Eruptions from the Sun-as-a-star H$\alpha$ Spectrum?

Solar filament/prominence eruptions can significantly perturb geospace when originating from favorable source locations and directions. While stellar analogs have been recently reported, the disk locations and magnetic environments of their source regions remain spatially unresolved on other stars. To bridge this gap, we investigate the typical Sun-as-a-star H$\alpha$ temporal spectral characteristics of solar filament/prominence eruptions with different source region locations (on-disk vs. limb, active region vs. quiet-Sun region). It is revealed that limb eruptions are characterized by blueshifted/redshifted emission caused by the bright off-limb erupting structures, whereas on-disk eruptions may show blueshifted absorptions due to the dark erupting filaments. Among the limb eruptions, front-side limb eruptions usually display line center emission before the blueshifted/redshifted emission, while far-side limb eruptions show the opposite sequence. Moreover, the magnetic environment at source also shapes the spectral characteristics. On-disk filament eruptions from active region exhibit much more intense flare-ribbon-dominated line center emission features compared with those from quiet-Sun region. Limb active region eruptions often show single-wing emissions, whereas large-scale quiet-Sun region (quiescent) prominence eruptions frequently display expansion-induced emission in both wings followed by line center absorption due to the disappearance of bright prominence. These distinct Sun-as-a-star H$\alpha$ spectral characteristics, dependent on eruption location, provide a diagnostic basis for inferring source regions of stellar filament/prominence eruptions from spatially unresolved H$\alpha$ spectra.

astro-ph.SR

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Diffusion based paradigms have emerged in recent years, offering novel perspectives for dataset distillation. However, they typically necessitate additional fine-tuning stages, and effective guidance mechanisms remain underexplored. To address these limitations, we rethink diffusion based dataset distillation and propose a Dual Matching Guided Diffusion (DMGD) framework, centered on efficient training-free guidance. We first establish Semantic Matching via conditional likelihood optimization, eliminating the need for auxiliary classifiers. Furthermore, we propose a dynamic guidance mechanism that enhances the diversity of synthetic data while maintaining semantic alignment. Simultaneously, we introduce an optimal transport (OT) based Distribution Matching approach to further align with the target distribution structure. To ensure efficiency, we develop two enhanced strategies for diffusion based framework: Distribution Approximate Matching and Greedy Progressive Matching. These strategies enable effective distribution matching guidance with minimal computational overhead. Experimental results on ImageNet-Woof, ImageNet-Nette, and ImageNet-1K demonstrate that our training-free approach achieves significant improvements, outperforming state-of-the-art (SOTA) methods requiring additional fine-tuning by average accuracy gains of 2.1%, 5.4%, and 2.4%, respectively.

cs.CV

Detectability and Systematic Bias from First-Order Phase-Transition Dephasing in Kerr EMRIs

We study gravitational-wave dephasing induced by an effective first-order phase transition in a Kerr extreme mass-ratio inspiral (EMRI). The transition is modeled phenomenologically as a finite-width restructuring of the dissipative flux sector, and its observational consequences are quantified with standard LISA matched-filter diagnostics. For a representative system with $M=2\times10^{5}M_\odot$, $\mu=1.4M_\odot$, and $\hat a=0.90$, we obtain $\rho_{\rm B}=5.064$, $\rho_{\rm T}=4.073$, $\rho_{\rm R}=1.051$, and a mismatch $\mathcal M=2.986\times10^{-3}$ after maximization over extrinsic time and phase shifts. Although the normalized mismatch remains small, the accumulated phase difference grows to $\Delta\Phi_{22}^{\rm SF}\sim 5\times10^{3}\,\mathrm{rad}$, indicating that a narrow transition window can generate a large coherent deformation of the inspiral clock while leaving the waveform globally close to the baseline branch in detector-weighted norm. The resulting signal therefore lies in a bias-sensitive regime, characterized by small mismatch, order-unity residual norm, and large cumulative dephasing. Our results suggest that the dominant consequence of the transition sector is not loss of detectability, but loss of faithfulness for precision inference. This motivates future LISA EMRI waveform models that incorporate parameterized transition sectors directly into the waveform manifold.

gr-qc

TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?

Automated theorem proving (ATP) benchmarks largely consist of problems formalized in MathLib, so current ATP training and evaluation are heavily biased toward MathLib's definitional framework. However, frontier mathematics is often exploratory and prototype-heavy, relying on bespoke constructions that deviate from standard libraries. In this work, we evaluate the robustness of current ATP systems when applied to a novel definitional framework, specifically examining the performance gap between standard library problems and bespoke mathematical constructions. We introduce TaoBench, an undergraduate-level benchmark derived from Terence Tao's Analysis I, which formalizes analysis by constructing core mathematical concepts from scratch, without relying on standard Mathlib definitions, as well as by mixing from-scratch and MathLib constructions. For fair evaluation, we build an agentic pipeline that automatically extracts a compilable, self-contained local environment for each problem. To isolate the effect of definitional frameworks, we additionally translate every problem into a mathematically equivalent Mathlib formulation, yielding paired TaoBench-Mathlib statements for direct comparison. While state-of-the-art ATP models perform capably within the MathLib framework, performance drops by an average of roughly 26% on the definitionally equivalent Tao formulation. This indicates that the main bottleneck is limited generalization across definitional frameworks rather than task difficulty. TaoBench thus highlights a gap between benchmark performance and applicability, and provides a concrete foundation for developing and testing provers better aligned with research mathematics.

cs.LG

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR (Long-context Geometric Reconstruction), a novel architecture that scales dense 3D reconstruction to extremely long sequences without post-optimization. LoGeR processes video streams in chunks, leveraging strong bidirectional priors for high-fidelity intra-chunk reasoning. To manage the critical challenge of coherence across chunk boundaries, we propose a learning-based hybrid memory module. This dual-component system combines a parametric Test-Time Training (TTT) memory to anchor the global coordinate frame and prevent scale drift, alongside a non-parametric Sliding Window Attention (SWA) mechanism to preserve uncompressed context for high-precision adjacent alignment. Remarkably, this memory architecture enables LoGeR to be trained on sequences of 128 frames, and generalize up to thousands of frames during inference. Evaluated across standard benchmarks and a newly repurposed VBR dataset with sequences of up to 19k frames, LoGeR substantially outperforms prior state-of-the-art feedforward methods--reducing ATE on KITTI by over 74%--and achieves robust, globally consistent reconstruction over unprecedented horizons.

cs.CV

Observation of a structurally driven, reversible topological phase transition in a distorted square net material

Topological materials hold immense promise for exhibiting exotic quantum phenomena, yet achieving controllable topological phase transitions remains challenging. Here, we demonstrate a structurally driven, reversible topological phase transition in the distorted square net material GdPS, induced via in situ potassium dosing. Using angle-resolved photoemission spectroscopy and first principles calculations, we demonstrate a cascade of topological phases in the sub-surface P layer: from a large, topologically trivial band gap to a gapless Dirac cone state with a 2 eV dispersion, and finally to a two-dimensional topological insulator as inferred from theory. This evolution is driven by subtle structural distortions in the first P layer caused by potassium adsorption, which in turn contribute to the band gap closure and topological phase transition. Furthermore, the ability to manipulate the topology of a sub-surface layer in GdPS offers a unique route for exploring and controlling topological states in bulk materials.

cond-mat.mtrl-sci

VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long horizons. We introduce VisGym, a gymnasium of 17 environments for evaluating and training VLMs. The suite spans symbolic puzzles, real-image understanding, navigation, and manipulation, and provides flexible controls over difficulty, input representation, planning horizon, and feedback. We also provide multi-step solvers that generate structured demonstrations, enabling supervised finetuning. Our evaluations show that all frontier models struggle in interactive settings, achieving low success rates in both the easy (46.6%) and hard (26.0%) configurations. Our experiments reveal notable limitations: models struggle to effectively leverage long context, performing worse with an unbounded history than with truncated windows. Furthermore, we find that several text-based symbolic tasks become substantially harder once rendered visually. However, explicit goal observations, textual feedback, and exploratory demonstrations in partially observable or unknown-dynamics settings for supervised finetuning yield consistent gains, highlighting concrete failure modes and pathways for improving multi-step visual decision-making. Code, data, and models can be found at: https://visgym.github.io/.

cs.CV

Spherical Geometry Diffusion: Generating High-quality 3D Face Geometry via Sphere-anchored Representations

A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D space, making it challenging for existing models to establish clean connectivity and resulting in suboptimal geometry. To address this, our core insight is to simplify the underlying geometric structure by constraining the distribution onto a simple and regular manifold, a topological sphere. Building on this, we first propose the Spherical Geometry Representation, a novel face representation that anchors geometric signals to uniform spherical coordinates. This guarantees a regular point distribution, from which the mesh connectivity can be robustly reconstructed. Critically, this canonical sphere can be seamlessly unwrapped into a 2D map, creating a perfect synergy with powerful 2D generative models. We then introduce Spherical Geometry Diffusion, a conditional diffusion framework built upon this 2D map. It enables diverse and controllable generation by jointly modeling geometry and texture, where the geometry explicitly conditions the texture synthesis process. Our method's effectiveness is demonstrated through its success in a wide range of tasks: text-to-3D generation, face reconstruction, and text-based 3D editing. Extensive experiments show that our approach substantially outperforms existing methods in geometric quality, textual fidelity, and inference efficiency.

cs.CV

Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment

Novel View Synthesis (NVS) has traditionally relied on models with explicit 3D inductive biases combined with known camera parameters from Structure-from-Motion (SfM) beforehand. Recent vision foundation models like VGGT take an orthogonal approach -- 3D knowledge is gained implicitly through training data and loss objectives, enabling feed-forward prediction of both camera parameters and 3D representations directly from a set of uncalibrated images. While flexible, VGGT features lack explicit multi-view geometric consistency, and we find that improving such 3D feature consistency benefits both NVS and pose estimation tasks. We introduce Selfi, a self-improving 3D reconstruction pipeline via feature alignment, transforming a VGGT backbone into a high-fidelity 3D reconstruction engine by leveraging its own outputs as pseudo-ground-truth. Specifically, we train a lightweight feature adapter using a reprojection-based consistency loss, which distills VGGT outputs into a new geometrically-aligned feature space that captures spatial proximity in 3D. This enables state-of-the-art performance in both NVS and camera pose estimation, demonstrating that feature alignment is a highly beneficial step for downstream 3D reasoning.

cs.CV

BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning

As retrieval-augmented generation (RAG) tackles complex tasks, increasingly expanded contexts offer richer information, but at the cost of higher latency and increased cognitive load on the model. To mitigate this bottleneck, especially for intricate multi-hop questions, we introduce BRIEF-Pro. It is a universal, lightweight compressor that distills relevant evidence for a given query from retrieved documents into a concise summary for seamless integration into in-context RAG. Using seed data consisting of relatively short contexts (fewer than 1k words), BRIEF-Pro is trained to perform abstractive compression of extended contexts exceeding 10k words across a wide range of scenarios. Furthermore, BRIEF-Pro offers flexible user control over summary length by allowing users to specify the desired number of sentences. Experiments on four open-domain multi-hop question-answering datasets show that BRIEF-Pro generates more concise and relevant summaries, enhancing performance across small, large, and proprietary language models. With the 70B reader model, 32x compression by BRIEF-Pro improves QA performance by 4.67% on average over LongLLMLingua's 9x, while requiring only 23% of its computational overhead.

cs.CL

Phase-sensitive evidence for pair density waves in a kagome superconductor

Pair density wave (PDW) exhibits periodic amplitude and sign modulations of the superconducting order parameter. Such a pairing state has long been proposed to be highly sensitive to nonmagnetic scattering, but its experimental realization remains elusive. Here we discover a nonmagnetic PDW-breaking effect in a kagome superconductor, using designer atomic nonmagnetic impurities and high-precision scanning tunneling microscopy (STM) at a base temperature of 30mK. We detect 2x2 pair density modulations by Josephson STM with a superconducting tip and 2x2 pairing gap modulations by normal STM. We find that the pairing modulations in both cases are substantially suppressed upon doping the kagome lattice with dilute isovalent nonmagnetic impurities, whereas the charge order and uniform superconductivity remain robust. We further identify the correlation between atomic dopants and the local suppression of PDW. We attribute these findings to a nonmagnetic pair-breaking effect, arising from the phase modulation of PDW in the kagome d-orbital. Taken together with its signatures in other state-of-the-art spectroscopy and transport measurements linked by theory, our findings support the ground state of the kagome superconductor as a correlated topological phase with superconducting loop currents.

cond-mat.supr-con

Spontaneous $\pi$ flux trapping in granular rings of unconventional superconductors

We study Josephson couplings in unconventional superconductors and generalize the Sigrist-Rice formula by incorporating symmetry constraints and interface orientation disorder. Applying this framework to granular superconducting rings, we establish a no-go result that single-band chiral superconductors cannot spontaneously trap a magnetic flux. This rules out chiral $p$-wave pairing in $\beta$-Bi$_2$Pd, in light of the half-quantum flux observed in recent Little-Parks experiments. Incorporating the full crystalline and time-reversal symmetries, we show that a fully gapped helical equal-spin pairing state, naturally stabilized by spin-orbit coupling arising from local inversion-symmetry breaking, is instead favored. We further find that granular rings of such superconductors can trap a spontaneous $\pi$ flux in a manner robust against interface disorder.

cond-mat.supr-con