SearcharxivSearch

arXiv subjects

Yi Ren

Publications and source records attributed to Yi Ren.

At least 37 records · Page 2Linked to original sources

A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators

Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based 3D-DRAM has been adopted in LLM accelerators. While this emerging technology provides strong performance gains over existing hardware, current 3D-DRAM accelerators (3D-Accelerators) rely on closed-source evaluation tools, limiting access to publicly available performance analysis methods. Moreover, existing designs are highly customized for specific scenarios, lacking a general and reusable full-stack modeling for 3D-Accelerators across diverse usecases. To bridge this fundamental gap, we present ATLAS, the first silicon-proven Architectural Three-dimesional-DRAM-based LLM Accelerator Simulation framework. Built on commercially deployed multi-layer 3D-DRAM technology, ATLAS introduces unified abstractions for both 3D-Accelerator system architecture and programming primitives to support arbitrary LLM inference scenarios. Validation against real silicon shows that ATLAS achieves $\le$8.57% simulation error and 97.26-99.96\% correlation with measured performance. Through design space exploration with ATLAS, we demonstrate its ability to guide architecture design and distill key takeaways for both 3D-DRAM memory system and 3D-Accelerator microarchitecture across scenarios. ATLAS will be open-sourced upon publication, enabling further research on 3D-Accelerators.

cs.AR

Towards Multi-Object Nonprehensile Transportation via Shared Teleoperation: A Framework Based on Virtual Object Model Predictive Control

Multi-object nonprehensile transportation in teleoperation demands simultaneous trajectory tracking and tray orientation control. Existing methods often struggle with model dependency, uncertain parameters, and multi-object adaptability. We propose a shared teleoperation framework where humans and robots share positioning control, while the robot autonomously manages orientation to satisfy dynamic constraints. Key contributions include: 1) A theoretical dynamic constraint analysis utilizing a novel virtual object (VO)-based method to simplify constraints for trajectory planning. 2) An MPC-based trajectory smoothing algorithm that enforces real-time constraints and coordinates user tracking with orientation control. 3) Validations demonstrating stable manipulation of nine objects at accelerations up to 2.4 m/s2. Compared to the baseline, our approach reduces sliding distance by 72.45% and eliminates tip-overs (0% vs. 13.9%), proving robust adaptability in complex scenarios.

cs.RO

Extinction Distributions in Nearby Star-resolved Galaxies. II. M33

Extinction maps are essential for tracing interstellar dust and enabling accurate stellar population studies in galaxies. Here, a high-resolution extinction distribution of nearby galaxy M33 is constructed by fitting multiband color indexes of the individually resolved red giant branch (RGB) stars from the Panchromatic Hubble Andromeda Treasury: Triangulum Extended Region (PHATTER) survey. Achieving an angular resolution of approximately 6$^{\prime\prime}$ ($\sim$ 24.4 pc), the extinction map reveals the intricate and heterogeneous distribution of dust throughout the entire disk of M33, with distinct delineation of spiral arms, inter-arm regions, and compact dust clouds. In addition, it exhibits strong spatial correspondence with the distributions of total hydrogen, H I, and CO, underscoring the reliability of the extinction map for tracing both diffuse and dense components of the interstellar medium. The derived $V$-band extinction reaches up to 2.5 mag per pixel, with a mean value of about 1.05 mag. Beyond providing new insights into the dust structure of M33, the extinction map offers a robust foundation for accurate extinction corrections and will support future studies, including upcoming observations with the Chinese Space Station Telescope.

astro-ph.GA

A Wireless World Model for AI-Native 6G Networks

Integrating AI into the physical layer is a cornerstone of 6G networks. However, current data-driven approaches struggle to generalize across dynamic environments because they lack an intrinsic understanding of electromagnetic wave propagation. We introduce the Wireless World Model (WWM), a multi-modal foundation framework predicting the spatiotemporal evolution of wireless channels by internalizing the causal relationship between 3D geometry and signal dynamics. Pre-trained on a massive ray-traced multi-modal dataset, WWM overcomes the data authenticity gap, further validated under real-world measurement data. Using a joint-embedding predictive architecture with a multi-modal mixture-of-experts Transformer, WWM fuses channel state information, 3D point clouds, and user trajectories into a unified representation. Across the five key downstream tasks supported by WWM, it achieves remarkable performance in seen environments, unseen generalization scenarios, and real-world measurements, consistently outperforming SOTA uni-modal foundation models and task-specific models. This paves the way for physics-aware 6G intelligence that adapts to the physical world.

cs.NI

End-to-End Dexterous Grasp Learning from Single-View Point Clouds via a Multi-Object Scene Dataset

Dexterous grasping in multi-object scene constitutes a fundamental challenge in robotic manipulation. Current mainstream grasping datasets predominantly focus on single-object scenarios and predefined grasp configurations, often neglecting environmental interference and the modeling of dexterous pre-grasp gesture, thereby limiting their generalizability in real-world applications. To address this, we propose DGS-Net, an end-to-end grasp prediction network capable of learning dense grasp configurations from single-view point clouds in multi-object scene. Furthermore, we propose a two-stage grasp data generation strategy that progresses from dense single-object grasp synthesis to dense scene-level grasp generation. Our dataset comprises 307 objects, 240 multi-object scenes, and over 350k validated grasps. By explicitly modeling grasp offsets and pre-grasp configurations, the dataset provides more robust and accurate supervision for dexterous grasp learning. Experimental results show that DGS-Net achieves grasp success rates of 88.63\% in simulation and 78.98\% on a real robotic platform, while exhibiting lower penetration with a mean penetration depth of 0.375 mm and penetration volume of 559.45 mm^3, outperforming existing methods and demonstrating strong effectiveness and generalization capability. Our dataset is available at https://github.com/4taotao8/DGS-Net.

cs.RO

CellE: Automated Standard Cell Library Extension via Equality Saturation

Automated standard cell library extension is crucial for maximizing Quality of Results (QoR) in modern VLSI design. We introduce CellE, a novel framework that leverages formal methods to achieve exhaustive discovery of functionally equivalent subcircuits. CellE applies equality saturation to the post-mapping netlist, generating an e-graph to cluster all functionally equivalent implementations. This canonical representation enables an efficient pattern mining algorithm to select the most area-optimal standard cells. Experimental results show a 15.41% average area reduction (up to 23.64% over prior work). Furthermore, characterization in a commercial flow demonstrates an 8.00% average delay reduction, confirming CellE's superior QoR optimization capabilities.

cs.AR

Identifying Red Supergiants in the Local Group Using JWST Photometry. I. NGC 6822, Sextans A, NGC 300, WLM, and IC 1613

Red supergiants (RSGs) are crucial for studying the properties and evolution of massive stars. It is representative to conduct a census of RSGs across the Local Group, which spans a broad metallicity range. However, identifying RSGs in distant and metal-poor galaxies remains challenging mainly due to contamination of foreground dwarfs and observational limitations. In this work, we perform PSF photometry on publicly released JWST/NIRCam images of five Local Group galaxies: NGC 6822, Sextans~A, NGC 300, WLM, and IC 1613 using the DOLPHOT NIRCam module. We find an optimal color-color diagram (CCD) for metal-poor environments, that is F115W $-$ F200W versus F356W $-$ F444W, which clearly separates RSGs from foreground dwarfs. By using the CCD, we identify 208, 135, and 22 RSG candidates in NGC 6822, Sextans A, and NGC 300, respectively, free from contamination by foreground dwarfs and oxygen-rich asymptotic giant branch stars (O-AGBs). In addition, 40 and 14 RSG candidates are directly selected on the CMD in WLM and IC 1613, respectively. Compared with previous works, the number of RSG candidates within the same luminosity range and sky region increases significantly, demonstrating the advantages of JWST in constructing a more complete RSG sample in the Local Group thanks to its high spatial resolution and photometric quality. In addition, catalogs of O-AGBs and carbon-rich AGBs (C-AGBs) are provided as by-products.

astro-ph.GA

Hierarchical Vision-Language Interaction for Facial Action Unit Detection

Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Interaction for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal attention architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross-modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language-based AU details, culminating in robust and semantically enriched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.

cs.CV

BiKC+: Bimanual Hierarchical Imitation with Keypose-Conditioned Coordination-Aware Consistency Policies

Robots are essential in industrial manufacturing due to their reliability and efficiency. They excel in performing simple and repetitive unimanual tasks but still face challenges with bimanual manipulation. This difficulty arises from the complexities of coordinating dual arms and handling multi-stage processes. Recent integration of generative models into imitation learning (IL) has made progress in tackling specific challenges. However, few approaches explicitly consider the multi-stage nature of bimanual tasks while also emphasizing the importance of inference speed. In multi-stage tasks, failures or delays at any stage can cascade over time, impacting the success and efficiency of subsequent sub-stages and ultimately hindering overall task performance. In this paper, we propose a novel keypose-conditioned coordination-aware consistency policy tailored for bimanual manipulation. Our framework instantiates hierarchical imitation learning with a high-level keypose predictor and a low-level trajectory generator. The predicted keyposes serve as sub-goals for trajectory generation, indicating targets for individual sub-stages. The trajectory generator is formulated as a consistency model, generating action sequences based on historical observations and predicted keyposes in a single inference step. In particular, we devise an innovative approach for identifying bimanual keyposes, considering both robot-centric action features and task-centric operation styles. Simulation and real-world experiments illustrate that our approach significantly outperforms baseline methods in terms of success rates and operational efficiency. Implementation codes can be found at https://github.com/JoanaHXU/BiKC-plus.

cs.RO

An Enhanced Sample of Galactic Red Supergiants Reveals Spiral Structures

Red supergiants (RSGs), representing a kind of massive young stellar population, have rarely been used to probe the structure of the Milky Way, mainly due to the long-standing scarcity of Galactic RSG samples. The Gaia BP/RP spectra (hereafter XP), which cover a broad wavelength range, provide a powerful tool for identifying RSGs. In this work, we develop a feedforward neural network classifier that assigns to each XP spectrum a probability of being an RSG, denoted as $\mathrm{P(RSG)}$. We perform ten independent runs with randomly divided training and validation sets, and apply each run to all XP spectra of stars with $G < 12$ mag. By selecting sources with $\mathrm{P(RSG)} \geq 0.9$, ten high-confidence candidate samples are obtained. A star is considered a ture Galactic RSG only if it appears in at least eight of these samples, yielding a final catalog of 2,436 objects. These RSGs show a clear spatial correlation with OB stars and trace the Galactic spiral arms well, confirming the reliability of our classification, and highlighting their potential to serve as powerful tracers of the Milky Way's structure.

astro-ph.GA

Discovery of an Extremely Luminous Type II Cepheid in the Andromeda Giant Stellar Stream: Evidence for a Hierarchical Triple with an Inner Binary Merger

We report the discovery of LAMOST J0041+3948, the most luminous post-AGB Type II Cepheid (TIIC) known, located in the Andromeda Giant Stellar Stream. Its spectral energy distribution (SED) exhibits a strong near-infrared excess, indicating the presence of a circumbinary dusty disk and hence binarity. SED fitting yields an effective temperature of $T_{\rm eff}=6738_{-262}^{+234}\,$K and a post-AGB luminosity of $\log(L/L_{\odot})=4.32_{-0.08}^{+0.07}$. Comparison with theoretical evolutionary tracks suggests a ~$2.0$-$4.0\,M_{\odot}$ progenitor when accounting for a possible scattered-light contribution. ZTF Light curves reveal a pulsation period of 89d that lies close to the period-luminosity relation for long-period RV Tauri stars. Follow-up spectroscopy reveals clear $s$-process enrichment and signatures consistent with an accretion disk around the companion. The inferred progenitor is significantly younger and more massive than a typical stream member, suggesting that an additional mechanism such as a stellar merger is required. We propose a formation channel in which the present post-AGB binary descends from a hierarchical triple system. In this scenario, the inner binary merged after the system was displaced to its current location by the galaxy merger event, and the resulting massive merger remnant subsequently evolved into the extremely luminous post-AGB star observed today.

astro-ph.SR

On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement

Tool-integrated (TI) reinforcement learning (RL) enables large language models (LLMs) to perform multi-step reasoning by interacting with external tools such as search engines and retrievers. Group Relative Policy Optimization (GRPO), exemplified by the recent Search-R1, offers fast convergence and a value-free formulation that makes it appealing for this setting, yet consistently suffers from training collapse. We identify Lazy Likelihood Displacement (LLD), a systematic reduction or stagnation in the likelihood of both correct and incorrect responses, as the core mechanism driving this failure. LLD emerges early and triggers a self-reinforcing LLD Death Spiral, where declining likelihood leads to low-confidence responses, inflating gradients, and ultimately causing collapse. We empirically characterize this process across models on a Search-R1-style, search-integrated question answering task, revealing a consistent three-phase trajectory: early stagnation, steady decay, and accelerated collapse. To address this, we propose a likelihood-preserving regularization LLDS that activates only when a response action's likelihood decreases, and regularizes only the tokens responsible. This fine-grained structure mitigates LLD with minimal interference. Our method stabilizes training, prevents gradient explosion, and yields substantial performance improvements across seven benchmarks, including relative improvements of +45.2% on Qwen2.5-3B and +37.1% on Qwen2.5-7B over vanilla GRPO training. Our results establish LLD as a previously overlooked bottleneck in GRPO-based TIRL and provide a practical path toward stable, scalable training of tool-integrated RL.

cs.CL

Hyperpolarized Molecular Nuclear Spins Achieve Magnetic Amplification

The use of nuclear spins as physical sensing systems is disadvantaged by their low signal responsivity, particularly when compared to sensing techniques based on electron spins. This primarily results from the small nuclear gyromagnetic ratio and the difficulties in achieving high spin polarization. Here we develop a new approach to investigating the response of hyperpolarized molecular nuclear spins to magnetic fields and demonstrate orders-of-magnitude enhanced magnetic responsivity over state-of-the-art proton and Overhauser magnetometers. Using hyperpolarized molecules with proton spins, we report the realization of magnetic amplification in linear and nonlinear types. We further extend this amplification to hyperpolarized scalar-coupled multi-spin molecules and observe substantial magnetic amplification exceeding 10%. Moreover, we observe an anomalous amplification with dispersive frequency dependence that originates from magnetic interference effects. Our work highlights the potential of hyperpolarized molecular nuclear spins for use in a new class of quantum sensors, with promising applications in both applied and fundamental physics, including highly accurate absolute magnetometry and the exploration of axion-nucleon exotic interactions.

quant-ph

Beyond the Sampled Token: Preserving Candidate Support in RLVR

We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction. We formally show that as probability concentrates on the top-$1$ candidate, the expected number of distinct responses collapses to one regardless of the sampling budget $K$. This theoretical implication is further verified by our empirical tracking of top-$N$ candidate probabilities during training, where the top-$1$ candidate progressively dominates while plausible alternatives are suppressed. These findings suggest a key desideratum for effective exploration: \emph{preserving non-negligible probability mass on the top-$N$ candidates}. To this end, we propose Candidate-aware Support Preservation (CaSP), with two complementary designs. Specifically, CaSP redistributes positive gradients among top-$N$ candidates for correct responses, and applies a stronger penalty to the top-$1$ candidate for incorrect responses. Unlike many exploration-oriented methods that improve pass@$K$ at the cost of pass@1, CaSP improves pass@$K$ across the full $K$ spectrum. These gains generalize to 6 math, 2 logical-reasoning, and 2 coding benchmarks, and scales to 32B-parameter models and sampling budgets up to $K=1024$, positioning it as a principled, candidate-level approach for RLVR exploration.

cs.AI

LightSAE: Parameter-Efficient and Heterogeneity-Aware Embedding for IoT Multivariate Time Series Forecasting

Modern Internet of Things (IoT) systems generate massive, heterogeneous multivariate time series data. Accurate Multivariate Time Series Forecasting (MTSF) of such data is critical for numerous applications. However, existing methods almost universally employ a shared embedding layer that processes all channels identically, creating a representational bottleneck that obscures valuable channel-specific information. To address this challenge, we introduce a Shared-Auxiliary Embedding (SAE) framework that decomposes the embedding into a shared base component capturing common patterns and channel-specific auxiliary components modeling unique deviations. Within this decomposition, we \rev{empirically observe} that the auxiliary components tend to exhibit low-rank and clustering characteristics, a structural pattern that is significantly less apparent when using purely independent embeddings. Consequently, we design LightSAE, a parameter-efficient embedding module that operationalizes these observed characteristics through low-rank factorization and a shared, gated component pool. Extensive experiments across 9 IoT-related datasets and 4 backbone architectures demonstrate LightSAE's effectiveness, achieving MSE improvements of up to 22.8\% with only 4.0\% parameter increase.

cs.LG

Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning

Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward exploration or exploitation remains an open problem. We introduce Token Hidden Reward (THR), a token-level metric that quantifies each token's influence on the likelihood of correct responses under Group Relative Policy Optimization (GRPO). We find that training dynamics are dominated by a small subset of tokens with high absolute THR values. Most interestingly, tokens with positive THR strengthen confidence in correct outputs, thus favoring exploitation, while tokens with negative THR preserve probability mass for alternative outputs, enabling exploration. This insight suggests a natural intervention: a THR-guided reweighting algorithm that modulates GRPO's learning signals to explicitly bias training toward exploitation or exploration. We validate the efficacy of this algorithm on diverse math reasoning benchmarks. By amplifying tokens with positive THR value and weakening negative ones, our algorithm improves greedy-decoding accuracy, favoring exploitation. The reverse strategy yields consistent gains in Pass@K accuracy, favoring exploration. We further demonstrate that our algorithm integrates seamlessly with other RL objectives such as GSPO and generalizes across architectures including Llama. These findings establish THR as a principled and fine-grained mechanism for dynamically controlling exploration and exploitation in RL-tuned LLMs, providing new tools for targeted fine-tuning in reasoning-intensive applications.

cs.LG

Learning Dynamics of Deep Learning -- Force Analysis of Deep Neural Networks

This thesis explores how deep learning models learn over time, using ideas inspired by force analysis. Specifically, we zoom in on the model's training procedure to see how one training example affects another during learning, like analyzing how forces move objects. We break this influence into two parts: how similar the two examples are, and how strong the updating force is. This framework helps us understand a wide range of the model's behaviors in different real systems. For example, it explains why certain examples have non-trivial learning paths, why (and why not) some LLM finetuning methods work, and why simpler, more structured patterns tend to be learned more easily. We apply this approach to various learning tasks and uncover new strategies for improving model training. While the method is still developing, it offers a new way to interpret models' behaviors systematically.

cs.LG

Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization

With the diminishing return from Moore's Law, system-technology co-optimization (STCO) has emerged as a promising approach to sustain the scaling trends in the VLSI industry. By bridging the gap between system requirements and technology innovations, STCO enables customized optimizations for application-driven system architectures. However, existing research lacks sufficient discussion on efficient STCO methodologies, particularly in addressing the information gap across design hierarchies and navigating the expansive cross-layer design space. To address these challenges, this paper presents Orthrus, a dual-loop automated framework that synergizes system-level and technology-level optimizations. At the system level, Orthrus employs a novel mechanism to prioritize the optimization of critical standard cells using system-level statistics. It also guides technology-level optimization via the normal directions of the Pareto frontier efficiently explored by Bayesian optimization. At the technology level, Orthrus leverages system-aware insights to optimize standard cell libraries. It employs a neural network-assisted enhanced differential evolution algorithm to efficiently optimize technology parameters. Experimental results on 7nm technology demonstrate that Orthrus achieves 12.5% delay reduction at iso-power and 61.4% power savings at iso-delay over the baseline approaches, establishing new Pareto frontiers in STCO.

cs.AR