SearcharxivSearch

arXiv subjects

Shumin Wang

Publications and source records attributed to Shumin Wang.

8 recordsLinked to original sources

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present $\Phi$-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, $\Phi$-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.

cs.CL

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing intermediate 2D/3D regions. However, they remain vulnerable to two intertwined failure modes: the predicted or selected regions may miss the target interaction area or have unsuitable granularity, while language grounding may confuse visually similar alternatives under relational instructions. To this end, we introduce ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning. Specifically, the Affordance Proposal Generation module first uses learnable affordance prompts and multi-level visual features to predict interaction-conditioned heatmaps, extracting a variable number of fine-grained proposals without parsed object or part names as segmentation prompts. Visual-Prompted Affordance Reasoning then reasons over labeled proposal overlays using the full instruction, returning identifiers in a structured "think-then-answer" response. Moreover, Group Relative Policy Optimization uses proposal-level rewards from lifted 3D overlap to align VPAR selection with final 3D grounding. On the SceneFun3D validation split, ThinkAfford achieves 10.69% AP50 and 25.46% AP25 under the official evaluator, outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines. Module-level diagnostics further show that APG attains 77.5% recall at 25% intersection-over-union, while GRPO-trained VPAR achieves 72.1% selection accuracy on APG-covered queries, compared with 63.4% under supervised fine-tuning.

cs.CV

On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models

Entropy serves as a critical metric for measuring the diversity of outputs generated by large language models (LLMs), providing valuable insights into their exploration capabilities. While recent studies increasingly focus on monitoring and adjusting entropy to better balance exploration and exploitation in reinforcement fine-tuning (RFT), a principled understanding of entropy dynamics during this process is yet to be thoroughly investigated. In this paper, we establish a theoretical framework for analyzing the entropy dynamics during the RFT process, which begins with a discriminant expression that quantifies entropy change under a single logit update. This foundation enables the derivation of a first-order expression for entropy change, which can be further extended to the update formula of Group Relative Policy Optimization (GRPO). The corollaries and insights drawn from the theoretical analysis inspire the design of entropy control methods, and also offer a unified lens for interpreting various entropy-based methods in existing studies. We provide empirical evidence to support the main conclusions of our analysis and demonstrate the effectiveness of the derived entropy-discriminator clipping methods. This study yields novel insights into RFT training dynamics, providing theoretical support and practical strategies for optimizing the exploration-exploitation balance during LLM fine-tuning.

cs.LG

Multi-beam Beamforming in RIS-aided MIMO Subject to Reradiation Mask Constraints -- Optimization and Machine Learning Design

Reconfigurable intelligent surfaces (RISs) are an emerging technology for improving spectral efficiency and reducing power consumption in future wireless systems. This paper investigates the joint design of the transmit precoding matrices and the RIS phase shift vector in a multi-user RIS-aided multiple-input multiple-output (MIMO) communication system. We formulate a max-min optimization problem to maximize the minimum achievable rate while considering transmit power and reradiation mask constraints. The achievable rate is simplified using the Arimoto-Blahut algorithm, and the problem is broken into quadratic programs with quadratic constraints (QPQC) sub-problems using an alternating optimization approach. To improve efficiency, we develop a model-based neural network optimization that utilizes the one-hot encoding for the angles of incidence and reflection. We address practical RIS limitations by using a greedy search algorithm to solve the optimization problem for discrete phase shifts. Simulation results demonstrate that the proposed methods effectively shape the multi-beam radiation pattern towards desired directions while satisfying reradiation mask constraints. The neural network design reduces the execution time, and the discrete phase shift scheme performs well with a small reduction of the beamforming gain by using only four phase shift levels.

math.OC

The Double-Episode Jet Genesis of the eROSITA and Fermi Bubbles

The Fermi and eROSITA bubbles are giant gamma-ray and X-ray lobes in the Milky Way, extending up to $\sim$50{\deg} and ~$\sim$80{\deg} in galactic latitude, respectively, yet their origins remain debated. Using three-dimensional magnetohydrodynamic simulations, we investigate a scenario in which two temporally separated episodes of active galactic nucleus (AGN) jets launched from the Galactic center produce the bubbles, with each structure bounded by a forward shock. Our simulations reveal that the first jet pair, launched 15 Myr ago, forms the outer eROSITA bubbles (extending to $\sim$18 kpc), while the second, launched 5 Myr ago, creates the nested Fermi bubbles ($\sim$10 kpc height). This model broadly reproduces the observed elongated morphology, multi-band X-ray surface brightness distribution, O VIII/O VII line ratios, radio ridge structures, and gamma-ray emissions of the bubbles. Cosmic-ray electrons are accelerated \textit{in situ} at the shock fronts, explaining the sharp edges and nearly uniform gamma-ray surface brightness distribution of Fermi bubbles. The results suggest that the eROSITA and Fermi bubbles encode a time-resolved record of episodic AGN activity in the Galactic center, providing a physically motivated framework for interpreting their multi-wavelength properties.

astro-ph.HE

Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving

The significant achievements of pre-trained models leveraging large volumes of data in the field of NLP and 2D vision inspire us to explore the potential of extensive data pre-training for 3D perception in autonomous driving. Toward this goal, this paper proposes to utilize massive unlabeled data from heterogeneous datasets to pre-train 3D perception models. We introduce a self-supervised pre-training framework that learns effective 3D representations from scratch on unlabeled data, combined with a prompt adapter based domain adaptation strategy to reduce dataset bias. The approach significantly improves model performance on downstream tasks such as 3D object detection, BEV segmentation, 3D object tracking, and occupancy prediction, and shows steady performance increase as the training data volume scales up, demonstrating the potential of continually benefit 3D perception models for autonomous driving. We will release the source code to inspire further investigations in the community.

cs.CV

The Vapor-Solid-Solid Growth of Ge Nanowires on Ge (110) by Molecular Beam Epitaxy

We demonstrate Au-assisted vapor-solid-solid (VSS) growth of Ge nanowires (NWs) by molecular beam epitaxy (MBE) at 220 {\deg}C, which is compatible with the temperature window for Si-based integrated circuit. Low temperature grown Ge NWs hold a smaller size, similar uniformity and better fit with Au tips in diameter, in contrast to Ge NWs grown at around or above the eutectic temperature of Au-Ge alloy in the vapor-liquid-solid (VLS) growth. Three growth orientations were observed on Ge (110) by the VSS growth at 220 {\deg}C, differing from only one growth direction of Ge NWs by the VLS growth at a high temperature. The evolution of NWs dimension and morphology from the VLS growth to the VSS growth is qualitatively explained via analyzing the mechanism of the two growth modes.

physics.app-ph

Photoluminescence of InGaAs/GaAsBi/InGaAs type-II quantum well grown by gas source molecular beam epitaxy

InGaAs/GaAsBi/InGaAs quantum wells (QWs) were grown on GaAs substrates by gas source molecular beam epitaxy for realizing the type II band-edge line-up. Both type I and type II transitions were observed in the Bi containing W QWs and the photoluminescence intensity was enhanced in the sample with a high Bi content, which is mainly due to the improvement of carrier confinement. Blue-shift of type II transitions at high excitation power density was observed and ascribed to the band-bending effect. The calculated transition energies based on 8 band k.p model fit well with the experiment results. The experimental and theoretical results show that the type-II QW design is a new promising candidate for realizing long wavelength GaAs-based light emitting devices near 1.3 um.

cond-mat.mtrl-sci