SearcharxivSearch

arXiv subjects

Feng Zhou

Publications and source records attributed to Feng Zhou.

At least 19 recordsLinked to original sources

A review of simulation, measurement techniques, and development in chip thermal design

As integrated circuits advance toward higher power densities, three-dimensional integration, and heterogeneous packaging, chip thermal management has become a key bottleneck limiting device performance, reliability, and lifetime. This article systematically reviews numerical simulation methods and experimental measurement techniques for chip thermal design, with particular emphasis on the technical challenges associated with multiscale and multiphysics coupling, thermal boundary resistance measurement, and high-heat-flux cooling. We first introduce macro- and device-scale thermal simulation methods, including equivalent thermal-circuit models, the finite element method, and computational fluid dynamics, and discuss the application of phonon transport theory and molecular dynamics at microscopic scales. We then examine the advantages and limitations of infrared thermography, thermoreflectance, Raman thermometry, and embedded sensors. Current limitations include the enormous computational cost, inaccurate multiscale coupling, expensive experimental facilities, and the physical limits of conventional cooling technologies. Finally, we discuss emerging directions, including AI-accelerated thermal simulation, embedded microchannel liquid cooling, two-phase cooling, advanced high-thermal-conductivity materials, and multiphysics co-design, with the aim of advancing chip thermal management toward greater efficiency and intelligence.

cond-mat.mtrl-sci

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.

cs.CV

Efficient Temporal Point Processes via Monotone Alternating Splines

Temporal point processes (TPPs) have widespread applications across various domains. Compared to modeling the conditional intensity of a TPP, modeling its cumulative conditional intensity function (CCIF) improves computational efficiency and eliminates numerical approximation errors. However, current CCIF parameterizations uniformly rely on Monotone Neural Networks (MNNs), which we identify as suffering from three structural deadlocks--convexity restrictions, saturation limits, and violations of CCIF modeling requirements--that fundamentally restrict their representational capacity for complex temporal dynamics. To resolve these bottlenecks, this paper proposes a novel framework called Monotone Alternating Splines (MAS). By leveraging distinct interpolation and extrapolation components, MAS provides a flexible and efficient framework for modeling CCIFs. Theoretically, MAS's interpolation provides strong fitting accuracy, while its extrapolation supports robust generalization, reducing the irreducible approximation gaps of MNNs. Extensive experiments show that MAS achieves superior performance on both synthetic and real-world datasets.

cs.LG

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven manner. Flexformer builds on random Fourier feature-based linear attention and treats spectral frequencies as trainable parameters, enabling the model to learn a broad family of attention kernels. We develop both stationary and nonstationary variants, with the latter offering strictly greater expressiveness. Extensive experiments on language modeling and sequence classification demonstrate that Flexformer consistently outperforms baselines. Moreover, Flexformer can be effectively distilled from pretrained Transformers to recover softmax attention and exhibits strong kernel transferability across domains, achieving both high efficiency and competitive performance on long-sequence tasks.

cs.LG

Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding

Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios. In this paper, we propose the Cross-Head Attention Uplift Network (CHAUN) and Robust Adversarial Inverse Propensity Score (RA-IPS) method to address these limitations. CHAUN employs shared feature embeddings and cross-head attention mechanisms to dynamically integrate treatment-specific and control-specific representations, enhancing inter-group correlation modeling. Theoretically, we prove that access to the true propensity scores ensures ITE identifiability even with unobserved confounders. For practical scenarios lacking true propensity scores, RA-IPS adversarially optimizes propensity weights within constrained uncertainty sets to mitigate bias from unobserved variables. Experiments on public datasets (CRITEO-UPLIFT, LAZADA) and a production e-commerce dataset demonstrate CHAUN's superiority over state-of-the-art uplift models, achieving relative improvements of up to 25.6% in QINI scores. RA-IPS further enhances robustness, outperforming standard IPS by 5.4% under unobserved confounding. The results validate the effectiveness of our proposed methods in real-world causal inference tasks.

cs.LG

Sliding contact creates universal self-affine fractal surfaces

Surface roughness evolves during sliding, a process known as run-in, and the resulting topography controls friction, leakage, and failure from machines to geological faults. Yet the physical rule selecting this state remains unclear. We show that metals, rocks, and glasses develop universal self-similar roughness at short wavelengths, while retaining a material-dependent roll-off. A two-process model explains this behavior: junction formation and rupture drive universal roughening, whereas larger-scale deformation and/or fracture limit its growth.

cond-mat.soft

Structuring Human-AI Productive Interdependence by Strategic Level of Automation Selection for Qualitative Inquiry

While Large Language Models (LLMs) offer a solution to the scale-versus-depth dilemma in qualitative analysis, the paradigm of maximizing automation is fundamentally at odds with the interpretive nature of qualitative inquiry. We argue that effective Human-AI collaboration is not an automation problem, but an interdependence problem. This paper reframes the design of "co-data" systems through the lens of Interdependence Theory, proposing a formal framework to structure human-AI productive interdependence. The framework guides the selection of an appropriate Level of Automation (LoA) for different stages of the qualitative analysis process by assessing task risk and the cost of validation. We present a case study where this framework led to a deliberately interdependent workflow, fostering the calibrated trust necessary for rigorous analysis. We conclude by presenting three design principles that instantiate this framework, demonstrating how to leverage AI as a powerful partner while preserving the human researcher's irreplaceable role in the transformation process of meaning-making.

cs.HC

Physics-Aware 3D Gaussian Editing for Driving Scene Generation

3D Gaussian Splatting (3DGS) has shown great potential in autonomous driving simulation and data generation, enabling photorealistic reconstruction and flexible scene manipulation. However, existing 3DGS scene editing methods have limited support for road geometry editing (e.g., inserting speed humps or sunken roads), and generally do not couple such edits with plausible vehicle-road interaction dynamics. Such editing is essential for generating training data under extreme driving scenarios or evaluating system reliability under these road irregularities. Moreover, many optimization-based methods require minutes of per-edit refinement, while existing efficient alternatives mainly focus on appearance-level or object-level manipulation rather than physics-aware road irregularity editing. To address these limitations, we propose RoVES, a Road-and-Vehicle Editing System for physics-aware 3D Gaussian editing in driving scenes. RoVES enables single-image-driven road geometry insertion and couples the edited road profile with a 4-DOF half-car vehicle dynamics model to achieve physics-aware vehicle pose correction in vertical displacement and pitch. RoVES inserts road elements in a one-shot, optimization-free pipeline (1.84s), and the full pipeline (including color transfer and vehicle-dynamics-based pose correction) completes in 6.24s; it edits dynamic vehicles via pose editing and corrects poses frame-by-frame to approximate dynamics-consistent vertical displacement and pitch responses. Experiments on the Waymo dataset show that RoVES provides practical efficiency and competitive visual consistency for physics-aware driving scene generation.

cs.CV

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical external signals (e.g., a cat's purr) correspond to radically different internal states depending on physiological context. Existing Multimodal Large Language Models (MLLMs) are blind to high-frequency biological time-series data, restricting them to superficial behavioural pattern matching rather than genuine latent-state reasoning. To bridge this gap, we introduce Meow-Omni 1, the first open-source, quad-modal MLLM purpose-built for computational ethology. It natively fuses video, audio, and physiological time-series streams with textual reasoning. Through targeted architectural adaptation, we integrate specialized scientific encoders into a unified backbone and formalize intent inference via physiologically grounded cross-modal alignment. Evaluated on MeowBench, a novel, expert-verified quad-modal benchmark, Meow-Omni 1 achieves state-of-the-art intent-recognition accuracy (71.16%), substantially outperforming leading vision-language and omni-modal baselines. We release the complete open-source pipeline including model weights, training framework, and the Meow-10K dataset, to establish a scalable paradigm for inter-species intent understanding and to advance foundation models toward real-world veterinary diagnostics and wildlife conservation.

cs.CL

Autonomous operation of the DIAG0 diagnostic line for 6D phase-space monitoring at LCLS-II

Characterizing the full 6-dimensional phase-space distribution of beams from the LCLS-II photoinjector is essential for understanding and optimizing downstream accelerator performance. Long-term monitoring of this distribution is equally important for detecting drifts in machine state and implementing timely corrective actions. Continuous phase space characterization during routine operation demands reliable tomographic diagnostic measurements and fast, efficient reconstruction methods. In this work, we demonstrate the first fully autonomous 6-dimensional beam-tomography system deployed on the DIAG0 parasitic beamline at LCLS-II. Using machine-learning-based control algorithms, the system autonomously configures DIAG0 and executes tomographic manipulations within operational constraints, adaptively re-optimizing beamline parameters and scan ranges in response to changes in the incoming beam. Tomographic measurements are streamed to the S3DF computing cluster where generative analysis methods reconstruct the phase-space distribution. We demonstrate that this framework produces detailed 6-dimensional beam reconstructions at a cadence of one reconstruction every 5 to 10 minutes, enabling real-time, multi-hour monitoring of injector beam evolution with unprecedented fidelity. These results represent a significant step toward fully autonomous operation of accelerator beamlines with real-time beam diagnostics for current and next-generation accelerator facilities.

physics.acc-ph

Maximal hypersurfaces with prescribed light-like cones in Lorentz-Minkowski space

The purpose in this paper is to study the maximal hypersurfaces with multiple light-cones in Lorentz-Minkowski space by considering the weak solutions to the mean curvature equation with multiple Dirac masses. Such solutions are constructed via an approximation procedure, using regular solutions with smooth sources that converge weakly to the Dirac measures.

math.AP

Magnetoelastic instabilities in kagome antiferromagnet Mn3-xGa

We present a systematic study of the structural, magnetic, and transport properties of hexagonal Mn3-xGa alloys, revealing a series of composition-controlled emergent phenomena. By tuning the Mn concentration, we uncover distinct lattice responses, including a zero thermal expansion-like volume compensation behavior in Mn-poor compositions and a magnetoelastic-driven, field-assisted structural phase transition in Mn-rich samples. These lattice instabilities are accompanied by correlated magnetic and transport anomalies, including metamagnetic transitions, negative magnetoresistance, and anomalous Hall sign reversal. First-principles calculations demonstrate that the Hall sign reversal originates from crystal-symmetry breaking rather than magnetic reorientation alone. Our results establish composition as the key control parameter governing magnetoelastic coupling in Mn3-xGa, providing a unified framework to tailor structural, magnetic, and topological transport properties in kagome antiferromagnets and reconcile previously disparate experimental observations.

cond-mat.mtrl-sci

SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding

Autoregressive language models suffer from high inference latency due to their sequential decoding nature. Speculative decoding (SD) mitigates this by employing a lightweight draft model to propose candidate tokens, which are selectively verified by a larger target model. While existing methods either adopt multi-draft strategies to increase acceptance rates or block verification techniques to jointly verify multiple tokens, they remain limited by treating these improvements in isolation. In this work, we propose SpecTr-GBV, a novel SD method that unifies multi-draft and greedy block verification (GBV) into a single framework. By formulating the verification step as an optimal transport problem over draft and target token blocks, SpecTr-GBV improves both theoretical efficiency and empirical performance. We theoretically prove that SpecTr-GBV achieves the optimal expected acceptance length physically attainable within the framework of i.i.d. draft generation, and this bound improves as the number of drafts increases. Empirically, we evaluate SpecTr-GBV across five datasets and four baselines. Our method achieves superior speedup and significantly higher block efficiency while preserving output quality. In addition, we perform comprehensive ablation studies to evaluate the impact of various hyperparameters in the model.

cs.CL

Advanced Control of Electron Beams: Tailoring X-ray Production with Programmable Laser Shaping

Leveraging the full scientific capabilities of next-generation high-repetition-rate free-electron lasers requires programmable control over electron-beam properties at their source. The photoinjector drive laser defines the electron beam's initial six-dimensional phase-space distribution, yet has historically been limited to Gaussian or static flat-top profiles, with most manipulation occurring downstream. Here we demonstrate software-programmable ultraviolet pulse shaping at the LCLS-II photoinjector as a source-level actuator that complements traditional accelerator controls. Using a coupled architecture combining dispersion-controlled nonlinear frequency conversion with spatial-light-modulator spectral shaping, we generate user-defined temporal structures and observe their imprint on electron bunches through high-resolution time-domain diagnostics. Laser-imposed multi-peaked modulation persists through acceleration, magnetic compression, and undulator transport with shot-to-shot repeatability, producing clearly resolved current structure in the compressed beam. Variance-based reconstruction from transverse deflecting cavity measurements reveals structured X-ray emission profiles exhibiting temporal features consistent with the programmed laser waveform. By providing rapid, software-controlled reconfiguration of electron-beam initial conditions, this source-level control approach establishes a programmable upstream actuator for future adaptive optimization and autonomous facility operation at high-repetition-rate light sources.

physics.acc-ph

InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, whereas Mamba-style selective state-space models (SSMs) scale linearly but often struggle to capture high-rank and synchronous global interactions. We present a consistency boundary analysis that characterizes when diagonal short-memory SSMs can approximate causal attention and identifies structural gaps that remain. Motivated by this analysis, we propose InfoMamba, an attention-free hybrid architecture. InfoMamba replaces token-level self-attention with a concept bottleneck linear filtering layer that serves as a minimal-bandwidth global interface and integrates it with a selective recurrent stream through information-maximizing fusion (IMF). IMF dynamically injects global context into the SSM dynamics and encourages complementary information usage through a mutual-information-inspired objective. Extensive experiments on classification, dense prediction, and non-vision tasks show that InfoMamba consistently outperforms strong Transformer and SSM baselines, achieving competitive accuracy-efficiency trade-offs while maintaining near-linear scaling.

cs.LG

On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions

Transformer networks have achieved remarkable empirical success across a wide range of applications, yet their theoretical expressive power remains insufficiently understood. In this paper, we study the expressive capabilities of Transformer architectures. We first establish an explicit approximation of maxout networks by Transformer networks while preserving comparable model complexity. As a consequence, Transformers inherit the universal approximation capability of ReLU networks under similar complexity constraints. Building on this connection, we develop a framework to analyze the approximation of continuous piecewise linear functions by Transformers and quantitatively characterize their expressivity via the number of linear regions, which grows exponentially with depth. Our analysis establishes a theoretical bridge between approximation theory for standard feedforward neural networks and Transformer architectures. It also yields structural insights into Transformers: self-attention layers implement max-type operations, while feedforward layers realize token-wise affine transformations.

cs.LG

Long-range Modeling and Processing of Multimodal Event Sequences

Temporal point processes (TPPs) have emerged as powerful tools for modeling asynchronous event sequences. While recent advances have extended TPPs to handle textual information, existing approaches are limited in their ability to generate rich, multimodal content and reason about event dynamics. A key challenge is that incorporating multimodal data dramatically increases sequence length, hindering the ability of attention-based models to generate coherent, long-form textual descriptions that require long-range understanding. In this paper, we propose a novel framework that extends LLM-based TPPs to the visual modality, positioning text generation as a core capability alongside time and type prediction. Our approach addresses the long-context problem through an adaptive sequence compression mechanism based on temporal similarity, which reduces sequence length while preserving essential patterns. We employ a two-stage paradigm of pre-training on compressed sequences followed by supervised fine-tuning for downstream tasks. Extensive experiments, including on the challenging DanmakuTPP-QA benchmark, demonstrate that our method outperforms state-of-the-art baselines in both predictive accuracy and the quality of its generated textual analyses.

cs.CL

Diving into Kronecker Adapters: Component Design Matters

Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component structures. However, existing work largely treats the component structure as a fixed or heuristic design choice, leaving the dimensions and number of Kronecker components underexplored. In this paper, we identify component structure as a key factor governing the capacity of Kronecker adapters. We perform a fine-grained analysis of both the dimensions and number of Kronecker components. In particular, we show that the alignment between Kronecker adapters and full fine-tuning depends on component configurations. Guided by these insights, we propose Component Designed Kronecker Adapters (CDKA). We further provide parameter-budget-aware configuration guidelines and a tailored training stabilization strategy for practical deployment. Experiments across various architectures and modalities demonstrate the effectiveness of CDKA. Code is available at https://github.com/rainstonee/CDKA.

cs.LG