SearcharxivSearch

arXiv subjects

Junyu Zhang

Publications and source records attributed to Junyu Zhang.

At least 19 recordsLinked to original sources

On the $\Omega(n)$ SDP relaxation gap for Submodular Box-Constrained Quadratic Programming

Continuous submodular optimization is an important class of nonconvex problems with global guarantees, for which the submodular box-constrained quadratic programming (BCQP) forms a fundamental subclass. It arises naturally as quadratic subproblems of general twice-differentiable submodular problems and also has direct business and management applications such as pricing. This paper studies the worst-case relaxation gap of semidefinite programming (SDP) relaxations with a finite family of valid linear cuts for submodular BCQP. We show that in dimension $n\geq4$, for any SDP relaxation formulated with a finite family of linear cuts, there always exists a submodular BCQP instance with a strictly positive relaxation gap. By restricting to the normalized BCQP instances that remove the problem's scaling factor, we further quantify this limitation with explicit lower bounds. In general dimension $n\geq4$, when the SDP relaxation is formulated with only Boolean-quadric-polytope (BQP) valid cuts, including the commonly used McCormick and triangle inequalities, we establish an $\Omega(n)$ relaxation gap lower bound. For an arbitrary family of $m$ valid linear cuts, we establish an $\Omega(n/m^2)$ lower bound. These results capture the fundamental limitations of SDP relaxations, especially SDP with BQP valid cuts, and provide guidance on what additional constraints may be needed for more effective convex approximations of submodular BCQP.

math.OC

Low Ly$\alpha$ Visibility in Galaxy Overdensities: Reionization Topology and Neutral-Fraction Ceilings from DIVER over $4.8<z<11$

Ly-alpha emission is widely used to trace cosmic reionization, but its interpretation depends on how Ly-alpha visibility varies with galaxy environment. We use deep JWST/NIRSpec observations from Deep Insights into UV Spectroscopy at the Epoch of Reionization (DIVER) in GOODS-N to measure Ly-alpha visibility for 250 galaxies at 4.8 25 A. We combine these measurements with H-alpha and [O III] emitters from JWST/NIRCam wide-field slitless spectroscopy to map the density field around each DIVER galaxy. Galaxies with high Ly-alpha equivalent widths (W_Lyalpha>25 A) or high effective Ly-alpha escape fractions (f_esc,Lyalpha^eff>0.05) tend to lie farther from nearby H-alpha and [O III] emitters than galaxies with lower Ly-alpha visibility. The clearest signal occurs near the prominent GOODS-N overdensity at z~5.2, where fewer than 15% of galaxies show strong Ly-alpha emission. This trend is opposite to the simplest inside-out reionization expectation that overdensities produce larger ionized regions and enhance Ly-alpha visibility. Possible explanations include circumgalactic and local intergalactic opacity, dense absorbers, and gas kinematics. We also derive an empirical upper envelope for f_esc,Lyalpha^eff and calibrate it with reionization simulations. Interpreting this envelope as a limiting IGM-attenuation signal gives neutral-fraction ceilings of _max=0.36, 0.76, 0.74, 0.84, and 1.0 at z~5.2, 5.8, 6.7, 7.7, and 9.8, respectively. The z~8 ceiling disfavors an almost completely neutral IGM at this epoch. These results support patchy reionization already underway by z~8 and show that galaxy Ly-alpha visibility encodes both large-scale ionization topology and near-source gas structure.

astro-ph.GA

Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework

Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation. This condition gives a direct way to rule out convergence to strict saddles, but it also oversimplifies the actual noise structure, and does not match many stochastic optimization regimes. In overparameterized or interpolation models, the noise may vanish near stationarity. In finite-sum problems, the stochastic gradient noise may lie in a low-dimensional, data-dependent subspace. In these (common) scenarios, UE is naturally not satisfied. In this paper, we prove an abstract almost sure avoidance theorem for stochastic recursions without UE. The theorem replaces UE-type requirements by verifiable pathwise conditions. In applications, these conditions follow, e.g., from local smoothness and finite-moment assumptions under standard i.i.d. sampling, or from the finite-sum structure under without-replacement sampling. Since the stochastically sampled maps generally do not share a fixed point, the celebrated center-stable manifold argument used in deterministic analyses is not directly applicable. Instead, we use a path-dependent change of variables together with a pathwise Lyapunov--Perron-based proof strategy. As applications, we obtain strict saddle avoidance for stochastic mirror descent (including SGD) and for random reshuffling. For nonsmooth composite objectives, we prove avoidance results for a proximal-type stochastic gradient method. Combining these insights with suitable iterate convergence guarantees, this allows establishing convergence to local minimizers of the original objective function.

math.OC

A quasar hatching from a buried red phase at z = 3.7

We present JADES-GS 209777, previously cataloged as CANDELS J033238.02-274626.2, hereafter "the Hatchling," a red quasar at $z=3.711$. While the source has been reported in earlier deep-field catalogs, our multiwavelength analysis reveals a visible active nucleus still embedded in a dense gas- and dust-rich environment. Red quasar continua are often attributed to dust attenuation, including non-standard extinction curves, but the highly comprehensive multiwavelength data for this source provide direct constraints on the material being cleared. Using JWST/NIRSpec, NIRCam, MIRI, HST, MUSE, Chandra, ALMA, and VLA data, we detect broad emission lines and strong X-ray emission, showing that the active nucleus is at least partially exposed. We also detect H$\alpha$ and He I absorption, indicating dense gas close to the nucleus. Kinematically disturbed O I, Mg II, Na D, and [O III] features, together with extended Ly$\alpha$ emission over $\gtrsim 20$ kpc, further show that multiphase gas is being accelerated from the nuclear region into the host-galaxy environment. The ALMA detection reveals strong dust emission, with the inferred infrared luminosity placing the system in the ULIRG regime. The continuum is red and sharply declining toward the rest-frame UV, resembling compact red AGNs, and may reflect extreme dust attenuation, gas reprocessing, possible BAL-like absorption, or a combination of these effects. Regardless of which mechanism dominates the continuum shape, the line diagnostics show that the visible nucleus remains partially obscured by nearby material. The Hatchling therefore represents a unique opportunity to explore a poorly known transition phase in which feedback is likely clearing an enshrouded quasar and allowing it to emerge toward a more unobscured active nucleus.

astro-ph.GA

PhotoIFU: NIRCam as a Photometric Integral Field Unit for Mapping Feedback in Galaxies

We present PhotoIFU, a workflow that uses deep multi-band imaging as a low-resolution photometric integral field unit. Applied to PSF-matched JWST/NIRCam imaging, PhotoIFU treats each spatial pixel as a coarse SED element and fits the pixel SEDs with Prospector to map resolved stellar-population and ISM-related properties. We apply this approach to three galaxies at $z=1.3$--3.7 in JADES: two systems with extended ionized line emission and one post-starburst galaxy with an exceptionally strong neutral outflow. Pixel-by-pixel SED fitting gives maps of stellar-mass surface density, specific star formation rate, dust attenuation, gas-phase metallicity, and recent star-formation history. We find that regions selected from the extended-emission or outflow geometry occupy distinct parts of the resolved SED-property distribution compared with the full host. In the systems with extended ionized emission, these regions are generally less dusty, consistent with ionized emission being observed along dust-poor, low-column-density pathways through the host. In the neutral-outflow system, the selected regions show enhanced recent star formation, suggesting that compact rejuvenation may mark the aftermath of an earlier energetic phase. These results show that galactic outflows and extended emission-line structures can be spatially associated with measurable differences in resolved host-galaxy stellar populations and ISM-related properties. PhotoIFU provides an imaging-based method for resolved SED mapping of feedback-related structures in larger galaxy samples where full spectroscopic integral-field mapping is unavailable.

astro-ph.GA

An (in)complete NIRSpec census of Balmer absorption in Type 1 AGN -- radiation-driven outflows in little red dots, quasars and variable stars

A notable achievement of the first generation of JWST surveys was the discovery of an abundance of faint high-redshift ($z > 4$) broad-line active galactic nuclei (AGN) in the $L_{\rm bol} < 10^{45}$~erg~s$^{-1}$ regime, previously only accessible at $z < 1$. The high prevalence of absorption features in their broad hydrogen lines is one of the peculiarities of a significant fraction of this new population. In this paper, we conduct a broad census of Balmer absorption in a sample of 47 Type-1 AGN spanning $2 < z < 7$. Accounting for incompleteness of JWST spectroscopic data, we estimate that $44_{-6}^{+21}$~\% of Little Red Dots (LRDs) have absorption in \Has while the incidence rate is $< 25$~\% (at 2$\sigma$) in Little Blue Dots (LBDs). Additionally, R1000 JWST/NIRSpec data is strongly resolution limited, implying an inability to disentangle Doppler broadening in the absorption from optical depth effects. This is alleviated with R2700 observations, although some degeneracies remain. We find a significant correlation between the velocity of the Balmer absorption and the [OIII]5007 narrow line luminosity, suggesting a common driving mechanism. We do not find any other significant correlations (and do not confirm previously claimed correlations) between Balmer absorption and other spectral properties of LRDs (except for the expected correlation between \Has and \Hbs absorption velocities). Comparing LRDs, quasars and stellar \Has absorbers, we establish that absorption in both LRDs and broad absorption line quasars is consistent with radiatively driven outflows, echoing the physics of variable stars yet occurring at vastly different scales.

astro-ph.GA

Schattor: Schatten-family methods for deep learning optimization

Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis. Motivated by the limitations of SGD and the success of adaptive optimizers, we propose {\it Schattor}, a family of adaptive first-order methods based on Schatten norms. Schattor unifies SGD and the recently proposed matrix-variate adaptive optimizer Muon within a single Schatten-norm-based framework. We establish dimension-free stationarity guarantees for methods in the Schattor family for stochastic matrix optimization problems via a novel matrix martingale moment bound. We also develop multi-block extensions that adaptively balance block-wise optimization progress and prove dimension-free stationarity guarantees in this more general setting.

math.OC

A Single-Loop Regularized Newton Method for Nonconvex-Strongly-Concave Minimax Optimization

For smooth nonconvex-strongly-concave minimax problems, existing second-order methods share a common double-loop structure where the inner maximization is solved to sufficiently high accuracy before each second-order step. First-order methods are often adopted in the inner loops for scalability, but they also undermine the condition-insensitivity of second-order methods, limiting these methods to instances with mild conditioning. To resolve this issue, we propose a novel single-loop framework based on an equivalent regularized minimization reformulation of the original problem. By deriving a new adaptive cubic-quadratic majorization to dynamically absorb the non-Lipschitz components of the reformulated Hessian, we establish a regularized Newton method with robust theoretical guarantees across multiple settings. For deterministic problems, our single-loop method matches the $\mathcal{O}(\varepsilon^{-1.5})$ global iteration complexity of the existing double-loop second-order methods, while automatically achieving a local superlinear rate that is unavailable in existing works due to the inner-loop bottleneck. For the stochastic setting, we achieve $\mathcal{O}(\varepsilon^{-3})$ gradient and $\mathcal{O}(\varepsilon^{-2})$ Hessian complexities by integrating a recursive variance reduction, strictly improving those of the double-loop methods by $\mathcal{O}(\varepsilon^{-0.5})$ factors. In both deterministic and stochastic experiments, our methods significantly outperform the benchmarks, offering substantial speedups over the double-loop methods even under mildly conditioned instances. As a byproduct of our analysis, we close a gap in stochastic second-order methods for nonconvex minimization, where the best known result contains a nontrivial technical issue.

math.OC

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies computed from restricted-game payoffs, which can lead to inefficient expansions that provide limited global improvement. We propose to guide population expansion by directly evaluating the post-expansion population quality. Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration--selection framework that explicitly minimizes PE during expansion. We instantiate this framework as Global PSRO, a practical DRL-based algorithm that efficiently generates candidate responses and estimates PE via parameter-sharing conditional neural networks. Experiments across multiple two-player zero-sum games show that Global PSRO achieves lower exploitability and approximates Nash equilibria with significantly fewer policy iterations than prior PSRO methods.

cs.AI

Clipped Stochastic Gradient Tracking For Locally Smooth Functions

Most stochastic gradient tracking (GT) methods adopt pre-scheduled stepsize rules, while a few recent works studied adaptive stepsizes that attempt to respond to the problem's local landscape. These methods are typically built upon the problem's global smoothness constant in both analysis and implementation, even for the adaptive ones. On the one hand, for many problems the local smoothness constant may vary drastically across the domain, and sometimes even unbounded, using the global upper bound of the local constants is too conservative. On the other hand, drastic stepsize changes can cause difficulties in the analysis of convergence and consensus of distributed algorithms, making the direct use of local smoothness constants risky and theoretically challenging. In this paper, we propose a \emph{Relative Uniform Continuity} (RUC) regularity condition for the local smoothness constant as a function of sets. The RUC condition covers most common growth functions for local smoothness constant, ranging from constant and logarithmic to polynomial and even exponential. For RUC-regular distributed optimization problems with finite-sum structure, we derive a clipped gradient tracking method with staggered variance reduction, which only relies on the local smoothness of objective functions, and an $\mathcal{O}(\sum_in_i^{1.5}+n_i^{0.5}\epsilon^{-1})$ complexity has been established for our algorithm.

math.OC

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models

Robust optimization (RO) provides a principled framework for decision-making under uncertainty, but its practical use is often limited by the need to manually reformulate uncertain optimization models into tractable deterministic counterparts. Recent large language models (LLMs) have been shown promising for automating optimization formulation, yet RO reformulation remains challenging because it requires precise multi-step reasoning and mathematically consistent transformations. To facilitate systematic evaluation of LLM-based reformulation, for which no dedicated benchmark currently exists, we develop AutoRO-Bench, a benchmark featuring an automated data generation pipeline for the core RO reformulation task and a curated dataset for the RO application task. To address the reformulation challenge, we propose Automated Reformulation with Experience Memory (AutoREM), a tuning-free memory-augmented framework that autonomously builds a structured textual experience memory by reflecting on past failed trajectories through a tailored offline adaptation procedure. AutoREM requires neither domain-specific expert knowledge nor parameter updates, and the resulting memory readily transfers across different base LLMs. Experimental results show that AutoREM consistently improves the accuracy and efficiency of RO reformulation across in-distribution datasets, out-of-distribution datasets, and diverse base LLMs.

cs.AI

Early metal-enriched baryon cycling before the midpoint of cosmic reionization

Models predict that chemical enrichment and gas redistribution should begin rapidly once star formation starts, but direct constraints at the earliest epochs have been scarce. Here we show that metal-enriched gas in multiple ionic phases was already present around galaxies before the midpoint of cosmic reionization. Using JWST/NIRSpec rest-frame ultraviolet spectroscopy from SPURS, we detect blueshifted metal absorption in three galaxies at $7.2<z<9.3$. The detected transitions span neutral, low-ionization, and high-ionization species, including O I, Si II, C II, Si IV, and C IV, with velocity offsets of $|\Delta v|\sim 50$--$250\,\mathrm{km\,s^{-1}}$ relative to nebular systemic redshifts. The ionic coexistence, overlapping velocity structure, and equivalent-width ratios are consistent with outflowing or otherwise kinematically disturbed galaxy-associated gas, implying rapid metal enrichment. These results show that key conditions for baryon cycling were established in at least a subset of luminous galaxies within the first several hundred million years of cosmic time, well before the completion of reionization.

astro-ph.GA

Adaptive Newton-CG methods with global and local analysis for unconstrained optimization with H\"older continuous Hessian

In this paper, we study Newton-conjugate gradient (Newton-CG) methods for minimizing a nonconvex function $f$ whose Hessian is $(H_f,\nu)$-H\"older continuous with modulus $H_f>0$ and exponent $\nu\in(0,1]$. Recently proposed Newton-CG methods for this problem adopt (i) non-adaptive regularization and (ii) a nested line-search procedure, where (i) often leads to inefficient early progress and the loss of local superlinear convergence, and (ii) may incur high computational cost due to multiple solves of the Newton system per iteration. To address these limitations, we propose two novel Newton-CG algorithms, depending on the availability of $\nu$, that adaptively regularize the Newton system by leveraging the auto-conditioning technique to eliminate the nested line search. The proposed algorithms achieve the best-known iteration complexity ${\mathcal O}\big(H_f^{1/(1+\nu)}\epsilon^{-(2+\nu)/(1+\nu)}\big)$ for finding an $\epsilon$-stationary point and, simultaneously, enjoy local superlinear convergence near nondegenerate local minimizers. Numerical experiments further demonstrate the practical advantages of our algorithms over existing approaches.

math.OC

Shuffling the Stochastic Mirror Descent via Dual Lipschitz Continuity and Kernel Conditioning

The global Lipschitz smoothness condition underlies most convergence and complexity analyses via two key consequences: the descent lemma and the gradient Lipschitz continuity. How to study the performance of optimization algorithms in the absence of Lipschitz smoothness remains an active area. The relative smoothness framework from Bauschke-Bolte-Teboulle (2017) and Lu-Freund-Nesterov (2018) provides an extended descent lemma, ensuring convergence of Bregman-based proximal gradient methods and their vanilla stochastic counterparts. However, many widely used techniques (e.g., momentum schemes, random reshuffling, and variance reduction) additionally require the Lipschitz-type bound for gradient deviations, leaving their analysis under relative smoothness an open area. To resolve this issue, we introduce the dual kernel conditioning (DKC) regularity condition to regulate the local relative curvature of the kernel functions. Combined with the relative smoothness, DKC provides a dual Lipschitz continuity for gradients: even though the gradient mapping is not Lipschitz in the primal space, it preserves Lipschitz continuity in the dual space induced by a mirror map. We verify that DKC is widely satisfied by popular kernels and is closed under affine composition and conic combination. With these novel tools, we establish the first complexity bounds as well as the iterate convergence of random reshuffling mirror descent for constrained nonconvex relative smooth problems.

math.OC

A New Kernel Regularity Condition for Distributed Mirror Descent: Broader Coverage and Simpler Analysis

Existing convergence analyses of distributed optimization methods in non-Euclidean geometries typically rely on kernel assumptions: (i) global Lipschitz smoothness and (ii) bi-convexity of the associated Bregman divergence function. Unfortunately, these conditions are violated by nearly all kernels used in practice, leaving a huge theory-practice gap. This work closes this gap by developing a unified analytical tool that guarantees convergence under mild conditions. Specifically, we introduce Hessian relative uniform continuity (HRUC), a regularity satisfied by nearly all standard kernels. Importantly, HRUC is closed under concatenation, positive scaling, composition, and various kernel combinations. Leveraging the geometric structure induced by HRUC, we derive convergence guarantees for mirror-descent-based gradient tracking without imposing any restrictive assumptions. More broadly, our analysis techniques extend seamlessly to other decentralized optimization methods in genuinely non-Euclidean and non-Lipschitz settings.

math.OC

Undermassive Hosts of $z = 4-6 $ AGN from JWST/NIRCam Image Decomposition with CONGRESS, FRESCO, and JADES

In the local Universe, supermassive black hole (SMBH) masses strongly correlate with their host-galaxies' stellar masses ($M_{*}$), but galaxies hosting faint AGN recently found by JWST may deviate from this relation. To constrain the M$_{\text{BH}}$-M$_{*}$ relation at high redshift, we performed AGN-host image decomposition for 17 low-luminosity AGN galaxies at $z$ $\sim$ 4-6 using NIRCam images in the JADES GOODS-N field. These sources are identified as AGNs from broad H$\alpha$ emission lines detected by the CONGRESS and FRESCO surveys. We used galfit+MCMC to fit spatial profiles in 7 wide-band images and detected extended emission in 9 sources out of 17. The close spatial alignment between the extended components and the AGN centers indicates that this emission likely originates from the host galaxies. These sources are extended at 0.9-2.0~$\mu$m, suggesting significant host-galaxy light in the rest-frame UV. For the sources with the host detection, the stellar mass inferred based on image decomposition result can be 1-2 dex lower than the results without image decomposition. The BH-to-stellar mass ratio spans $M_{\text{BH}}/M_\ast$ $\sim$ 0.01-1.48, placing them well above the local $M_{\text{BH}}$-$M_\ast$ relation. In contrast, the host-galaxy size-mass relation broadly agrees with previous measurements. Our results suggest that the host galaxies of these faint AGN are either genuinely under-massive compared to their black hole masses, or too compact to be spatially resolved.

astro-ph.GA

Clumps in High-Redshift Galaxies: Mass Scaling and Radial Trends from JADES

Massive star-forming clumps are a prominent feature of high-redshift galaxies and are thought to trace gravitational fragmentation, feedback, and bulge growth in gas-rich disks. We present a statistical analysis of clumps in $\sim$3600 galaxies spanning $2 \lesssim z \lesssim 8$ from deep JWST/NIRCam imaging in the JADES GOODS--South field. Clumps are identified as residual features after subtracting smooth S\'ersic profiles, enabling a uniform, rest-frame optical census of sub-galactic structure. We characterize their physical properties, size--mass relations, and spatial distributions to constrain models of sub-galactic structure formation and evolution. We find that clumps in our sample are typically low-mass ($10^{\sim7-8}M_\odot$), actively star-forming, and show diverse gas-phase metallicity, dust attenuation, and stellar population properties. Their sizes and average pairwise separations increase with cosmic time (toward lower redshift), consistent with inside-out disk growth. The clump mass function follows a power law with slope $\alpha = -1.50_{-0.17}^{+0.19}$, consistent with fragmentation in turbulent disks. We find a deficit of relatively young clumps near galaxy centers and a radial transition in the size--mass relation: outer clumps exhibit steeper, near-virial slopes ($R_{\rm e}\propto M_*^{\sim 0.3}$), while inner clumps follow flatter trends ($R_{\rm e}\propto M_*^{\sim 0.2}$), consistent with structural evolution via migration or disruption. These results provide new constraints on the formation, survival, and dynamical evolution of clumps, highlighting their role in shaping galaxy morphology during the peak of cosmic star formation.

astro-ph.GA

BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models

Do large language models (LLMs) genuinely understand abstract concepts, or merely manipulate them as statistical patterns? We introduce an abstraction-grounding framework that decomposes conceptual understanding into three capacities: interpretation of abstract concepts (Abstract-Abstract, A-A), grounding of abstractions in concrete events (Abstract-Concrete, A-C), and application of abstract principles to regulate concrete decisions (Concrete-Concrete, C-C). Using human values as a testbed - given their semantic richness and centrality to alignment - we employ probing (detecting value traces in internal activations) and steering (modifying representations to shift behavior). Across six open-source LLMs and ten value dimensions, probing shows that diagnostic probes trained solely on abstract value descriptions reliably detect the same values in concrete event narratives and decision reasoning, demonstrating cross-level transfer. Steering reveals an asymmetry: intervening on value representations causally shifts concrete judgments and decisions (A-C, C-C), yet leaves abstract interpretations unchanged (A-A), suggesting that encoded abstract values function as stable anchors rather than malleable activations. These findings indicate LLMs maintain structured value representations that bridge abstraction and action, providing a mechanistic and operational foundation for building value-driven autonomous AI systems with more transparent, generalizable alignment and control.

cs.CL