Searcharxiv⌕ Search

arXiv subjects

Zhizhong Chen

Publications and source records attributed to Zhizhong Chen.

13 recordsLinked to original sources

Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling

Capturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers. Traditional orthogonal Video IDs fail to capture content relationships and demand large embedding tables, while the quadratic complexity of self-attention restricts the maximum sequence length under strict industrial latency and resource constraints. In this work, we present a production-deployed framework for modeling ultra-long user behavior sequences at a billion-user scale. We first address the representation bottleneck by adopting content-native Semantic IDs. By utilizing depth-truncated, coarse-grained Semantic IDs, we shrink the embedding table size from corpus cardinality. This compact representation naturally generalizes to cold-start content through shared semantic prefixes. Second, to overcome the sequence scaling barrier, we introduce a Global-Aware Compression Transformer that leverages non-parametric temporal folding and unified global query integration to effectively condense the sequence, alleviating both the memory and computational bottlenecks of standard self-attention. Offline profiling on our computing infrastructure demonstrates an order-of-magnitude reduction in peak memory footprint and a drastic decrease in computational overhead. This efficiency gain enables supporting longer sequence lengths at an affordable cost in production, yielding substantial online gains in satisfied user engagement and satisfied content consumption in large-scale online A/B tests.

cs.IR↗

PartDexTOG: Generating Dexterous Task-Oriented Grasping via Language-driven Part Analysis

Task-oriented grasping is a crucial yet challenging task in robotic manipulation. Despite the recent progress, few existing methods address task-oriented grasping with dexterous hands. Dexterous hands provide better precision and versatility, enabling robots to perform task-oriented grasping more effectively. In this paper, we argue that part analysis can enhance dexterous grasping by providing detailed information about the object's functionality. We propose PartDexTOG, a method that generates dexterous task-oriented grasps via language-driven part analysis. Taking a 3D object and a manipulation task represented by language as input, the method first generates the category-level and part-level grasp descriptions w.r.t the manipulation task by LLMs. Then, a category-part conditional diffusion model is developed to generate a dexterous grasp for each part, respectively, based on the generated descriptions. To select the most plausible combination of grasp and corresponding part from the generated ones, we propose a measure of geometric consistency between grasp and part. We show that our method greatly benefits from the open-world knowledge reasoning on object parts by LLMs, which naturally facilitates the learning of grasp generation on objects with different geometry and for different manipulation tasks. Our method ranks top on the OakInk-shape dataset over all previous methods, improving the Penetration Volume, the Grasp Displace, and the P-FID over the state-of-the-art by $3.58\%$, $2.87\%$, and $41.43\%$, respectively. Notably, it demonstrates good generality in handling novel categories and tasks.

cs.RO↗

Frequency-Resolved Forward Capacitance in GaN-based LEDs

This study establishes a unified framework for interpreting dynamic capacitive responses in InGaN-based light-emitting diodes (LEDs) through forward-bias capacitance-voltage-frequency spectroscopy. A hybrid impedance model integrating series RL components and parallel C-G networks was developed to resolve distinct frequency-dependent capacitive regimes. The low-frequency regime (<1 kHz) is governed by interfacial capacitance with characteristic reciprocal frequency dependence, while the mid-frequency range(10 kHz-6.4 MHz) demonstrates carrier diffusion and recombination dynamics. At MHz frequencies, negative capacitance manifests due to delayed carrier emission mediated by deep-level traps. The model achieved sub-1% fitting errors (R^2 > 0.99)across a broad bandwidth(10 kHz-6.4 MHz) , conclusively attributing negative capacitance to intrinsic trap processes rather than extrinsic artifacts. Critical advances include quantum well cap thickness modulation reducing mid-frequency capacitance by 30% and the dominance of trap-mediated inductance over parasitic contributions by three orders of magnitude. This framework resolves persistent controversies in LED impedance interpretation. By bridging semiconductor physics with device engineering, this methodology provides essential tools for designing next-generation optoelectronic systems requiring ultralow-latency operation and precise charge-state control.

physics.app-ph↗

Red Emission from Strain-Relaxed Bulk InGaN Active Region

High-In-content InGaN quantum wells (QWs) in red light-emitting diodes (LEDs) are typically grown at low temperatures to ensure effective In incorporation. In this study, red LEDs based on bulk InGaN active region were demonstrated. The growth temperature of bulk InGaN was ~800C, which is over 100C higher than the typical growth temperature of red QWs. By introducing high-density trench structures in the underlying green multi-quantum wells (MQWs), the compressive strain in bulk InGaN was relaxed by ~96%. With strain relaxation, phase separation occurred in the bulk InGaN, forming low-In-content (blue) and high-In-content (red) phases. The red phase acted as carrier localization centers, enabling red light emission under electrical injection. The red LEDs based on bulk InGaN exhibited a peak wavelength of 645 nm at 20 mA, with on-wafer peak external quantum efficiency of 0.32%. This study presents a new epitaxial strategy for red InGaN LEDs.

cond-mat.mtrl-sci↗

Challenges and Opportunities in Searching for Rashba-Dresselhaus Materials for Efficient Spin-Charge Interconversion at Room Temperature

Spintronic logic devices require efficient spin-charge interconversion: converting charge current to spin current and spin current to charge current. In spin-orbit materials that are regarded as the most promising candidate for spintronic logic devices, one mechanism that is responsible for spin-charge interconversion is Edelstein and inverse Edelstein effects based on spin-momentum locking in materials with Rashba-type spin-orbit coupling. Over last decade, there has been rapid progresses for increasing interconversion efficiencies due to the Edelstein effect in a few Rashba-Dresselhaus materials and topological insulators, making Rashba spin-momentum locking a promising technological solution for spin-orbit logic devices. However, despite the rapid progress that leads to high spin-charge interconversion efficiency at cryogenic temperatures, the room-temperature efficiency needed for technological applications is still low. This paper presents our understanding on the challenges and opportunities in searching for Rashba-Dresselhaus materials for efficient spin-charge interconversion at room temperature by focusing on materials properties such as Rashba coefficients, momentum relaxation times, spin-momentum locking relations and electrical conductivities.

cond-mat.mtrl-sci↗

Effect of Grain Coalescence on Dislocation and Stress Evolution of GaN Films Grown on Nanoscale Patterned Sapphire Substrates

Two types of nucleation layers (NLs), including in-situ low-temperature grown GaN (LT-GaN) and ex-situ sputtered physical vapor deposition AlN (PVD-AlN), are applied on cone-shaped nanoscale patterned sapphire substrate (NPSS). The initial growth process of GaN on these two NLs is comparably investigated by a series of growth interruptions. The coalescence process of GaN grains is modulated by adjusting the three-dimensional (3D) temperatures. The results indicate that higher 3D temperatures reduce the edge dislocation density while increasing the residual compressive stress in GaN films. Compared to the LT-GaN NLs, the PVD-AlN NLs effectively resist Ostwald ripening and facilitate the uniform growth of GaN grains on NPSS. Furthermore, GaN films grown on NPSS with PVD-AlN NLs exhibit a reduction of over 50% in both screw and edge dislocation densities compared to those grown on LT-GaN NLs. Additionally, PVD-AlN NLs result in an increase of about 0.5 GPa in the residual compressive stress observed in GaN films.

cond-mat.mtrl-sci↗

Scaled indium oxide transistors fabricated using atomic layer deposition

In order to continue to improve integrated circuit performance and functionality, scaled transistors with short channel lengths and low thickness are needed. But the further scaling of silicon-based devices and the development of alternative semiconductor channel materials that are compatible with current fabrication processes is challenging. Here we report atomic-layer-deposited indium oxide transistors with channel lengths down to 8 nm, channel thicknesses down to 0.5 nm and equivalent dielectric oxide thickness down to 0.84 nm. Due to the scaled device dimensions and low contact resistance, the devices exhibit high on-state currents of 3.1 A/mm at a drain voltage of 0.5 V and a transconductance of 1.5 S/mm at a drain voltage 1 V. Our devices are a promising alternative channel material for scaled transistors with back-end-of-line processing compatibility.

cond-mat.mtrl-sci↗

PoolRank: Max/Min Pooling-based Ranking Loss for Listwise Learning & Ranking Balance

Numerous neural retrieval models have been proposed in recent years. These models learn to compute a ranking score between the given query and document. The majority of existing models are trained in pairwise fashion using human-judged labels directly without further calibration. The traditional pairwise schemes can be time-consuming and require pre-defined positive-negative document pairs for training, potentially leading to learning bias due to document distribution mismatch between training and test conditions. Some popular existing listwise schemes rely on the strong pre-defined probabilistic assumptions and stark difference between relevant and non-relevant documents for the given query, which may limit the model potential due to the low-quality or ambiguous relevance labels. To address these concerns, we turn to a physics-inspired ranking balance scheme and propose PoolRank, a pooling-based listwise learning framework. The proposed scheme has four major advantages: (1) PoolRank extracts training information from the best candidates at the local level based on model performance and relative ranking among abundant document candidates. (2) By combining four pooling-based loss components in a multi-task learning fashion, PoolRank calibrates the ranking balance for the partially relevant and the highly non-relevant documents automatically without costly human inspection. (3) PoolRank can be easily generalized to any neural retrieval model without requiring additional learnable parameters or model structure modifications. (4) Compared to pairwise learning and existing listwise learning schemes, PoolRank yields better ranking performance for all studied retrieval models while retaining efficient convergence rates.

cs.IR↗

The Cross-Lingual Arabic Information REtrieval (CLAIRE) System

Despite advances in neural machine translation, cross-lingual retrieval tasks in which queries and documents live in different natural language spaces remain challenging. Although neural translation models may provide an intuitive approach to tackle the cross-lingual problem, their resource-consuming training and advanced model structures may complicate the overall retrieval pipeline and reduce users engagement. In this paper, we build our end-to-end Cross-Lingual Arabic Information REtrieval (CLAIRE) system based on the cross-lingual word embedding where searchers are assumed to have a passable passive understanding of Arabic and various supporting information in English is provided to aid retrieval experience. The proposed system has three major advantages: (1) The usage of English-Arabic word embedding simplifies the overall pipeline and avoids the potential mistakes caused by machine translation. (2) Our CLAIRE system can incorporate arbitrary word embedding-based neural retrieval models without structural modification. (3) Early empirical results on an Arabic news collection show promising performance.

cs.IR↗

ExpertRank: A Multi-level Coarse-grained Expert-based Listwise Ranking Loss

The goal of information retrieval is to recommend a list of document candidates that are most relevant to a given query. Listwise learning trains neural retrieval models by comparing various candidates simultaneously on a large scale, offering much more competitive performance than pairwise and pointwise schemes. Existing listwise ranking losses treat the candidate document list as a whole unit without further inspection. Some candidates with moderate semantic prominence may be ignored by the noisy similarity signals or overshadowed by a few especially pronounced candidates. As a result, existing ranking losses fail to exploit the full potential of neural retrieval models. To address these concerns, we apply the classic pooling technique to conduct multi-level coarse graining and propose ExpertRank, a novel expert-based listwise ranking loss. The proposed scheme has three major advantages: (1) ExpertRank introduces the profound physics concept of coarse graining to information retrieval by selecting prominent candidates at various local levels based on model prediction and inter-document comparison. (2) ExpertRank applies the mixture of experts (MoE) technique to combine different experts effectively by extending the traditional ListNet. (3) Compared to other existing listwise learning approaches, ExpertRank produces much more reliable and competitive performance for various neural retrieval models with different complexities, from traditional models, such as KNRM, ConvKNRM, MatchPyramid, to sophisticated BERT/ALBERT-based retrieval models.

cs.IR↗

Are "Undocumented Workers" the Same as "Illegal Aliens"? Disentangling Denotation and Connotation in Vector Spaces

In politics, neologisms are frequently invented for partisan objectives. For example, "undocumented workers" and "illegal aliens" refer to the same group of people (i.e., they have the same denotation), but they carry clearly different connotations. Examples like these have traditionally posed a challenge to reference-based semantic theories and led to increasing acceptance of alternative theories (e.g., Two-Factor Semantics) among philosophers and cognitive scientists. In NLP, however, popular pretrained models encode both denotation and connotation as one entangled representation. In this study, we propose an adversarial neural network that decomposes a pretrained representation as independent denotation and connotation representations. For intrinsic interpretability, we show that words with the same denotation but different connotations (e.g., "immigrants" vs. "aliens", "estate tax" vs. "death tax") move closer to each other in denotation space while moving further apart in connotation space. For extrinsic application, we train an information retrieval system with our disentangled representations and show that the denotation vectors improve the viewpoint diversity of document rankings.

cs.CL↗

Data-driven prediction and origin identification of epidemics in population networks

Effective intervention strategies for epidemics rely on the identification of their origin and on the robustness of the predictions made by network disease models. We introduce a Bayesian uncertainty quantification framework to infer model parameters for a disease spreading on a network of communities from limited, noisy observations; the state-of-the-art computational framework compensates for the model complexity by exploiting massively parallel computing architectures. Using noisy, synthetic data, we show the potential of the approach to perform robust model fitting and additionally demonstrate that we can effectively identify the disease origin via Bayesian model selection. As disease-related data are increasingly available, the proposed framework has broad practical relevance for the prediction and management of epidemics.

q-bio.PE↗

Spinodal Decomposition-Enabled Halide Perovskite Double Heterostructure with Reduced Fröhlich Electron-Phonon Coupling

Epitaxial III-V semiconductor heterostructures are key components in modern microelectronics, electro-optics and optoelectronics. With superior semiconducting properties, halide perovskite materials are rising as promising candidates for coherent heterostructure devices. In this report, spinodal decomposition is proposed and experimentally implemented to produce epitaxial double heterostructures in halide perovskite system. Pristine epitaxial mixed halide perovskites rods and films were synthesized via Van der Waals epitaxy by chemical vapor deposition method. At room temperature, photon was applied as a knob to regulate the kinetics of spinodal decomposition and classic coarsening. By this approach, halide perovskite double heterostructures were created carrying epitaxial interfaces and outstanding optical properties. Reduced Fröhlich electron-phonon coupling was discovered in coherent halide double heterostructure, which is hypothetically attributed to the classic phonon confinement effect widely existing in III-V double heterostructures. The ability to develop coherent double heterostructures in halide perovskites paves an avenue to exploring halide perovskite-based quantum wells and superlattices for high-performance and low-cost optoelectronics, electro-optics and microelectronics.

cond-mat.mtrl-sci↗