SearcharxivSearch

arXiv subjects

Yizhi Zhang

Publications and source records attributed to Yizhi Zhang.

At least 19 recordsLinked to original sources

Inferring Urban Mobility Interactions from Aggregated Dynamics

Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. Here we show that this directional structure need not be observed to be known: aggregated counts which cities already collect retain enough information to reconstruct the temporal evolution of origin-destination (OD) matrix. Using an uncertainty-aware physics-informed framework, we infer future OD flows from area-level counts alone across twelve mobility datasets from cities in the United States and China, reaching accuracy comparable to models that take historical OD matrices as input. Probabilistic modeling corrects the systematic underestimation of sparse, high-value corridors and yields calibrated predictions consistent with observed flows. Architectures that respect the generation-before-assignment logic of transport planning recover interactions more faithfully, indicating that location-level spatial heterogeneity should be preserved before pairwise interactions are reconstructed. Because inference requires only aggregated observations after training, recovering interactions this way reduces reliance on continuous individual-level tracking, pointing toward a more deployable and less exposure-heavy basis for real-time urban intelligence.

cs.LG

A Complete Characterization of Realizable (Embedding Dimension, Multiplicity) Pairs for Complete Intersection Numerical Semigroups

We give a complete characterization of the positive integer pairs $(e,m)$ that occur as the (embedding dimension, multiplicity) of a complete intersection numerical semigroup, thereby settling an open problem in the theory of relations of numerical semigroups. We prove that there exists a complete intersection numerical semigroup $\Gamma$ with $e(\Gamma)=e$ and $m(\Gamma)=m$ if and only if $(e,m)=(1,1)$ or $e\ge 2$ and $m\ge 2^{e-1}$.

math.AC

Telogenesis: Goal Is All U Need

Goal-conditioned systems assume goals are provided externally. We ask whether attentional priorities can emerge endogenously from an agent's internal cognitive state. We propose a priority function that generates observation targets from three epistemic gaps: ignorance (posterior variance), surprise (prediction error), and staleness (temporal decay of confidence in unobserved variables). We validate this in two systems: a minimal attention-allocation environment (2,000 runs) and a modular, partially observable world (500 runs). Ablation shows each component is necessary. A key finding is metric-dependent reversal: under global prediction error, coverage-based rotation wins; under change detection latency, priority-guided allocation wins, with advantage growing monotonically with dimensionality (d = -0.95 at N=48, p < 10^-6). Detection latency follows a power law in attention budget, with a steeper exponent for priority-guided allocation (0.55 vs. 0.40). When the decay rate is made learnable per variable, the system spontaneously recovers environmental volatility structure without supervision (t = 22.5, p < 10^-6). We demonstrate that epistemic gaps alone, without external reward, suffice to generate adaptive priorities that outperform fixed strategies and recover latent environmental structure.

cs.AI

Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving

Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these modalities face a fundamental Consistency-Realism Dilemma. Low-level signals (e.g., edges, blurred images) ensure precise control but compromise realism by "baking in" synthetic artifacts, whereas high-level priors (e.g., depth, semantics, HDMaps) facilitate photorealism but lack the structural detail required for consistent guidance. In this work, we present Driving with DINO (DwD), a novel framework that leverages Vision Foundation Module (VFM) features as a unified bridge between the simulation and real-world domains. We first identify that these features encode a spectrum of information, from high-level semantics to fine-grained structure. To effectively utilize this, we employ Principal Subspace Projection to discard the high-frequency elements responsible for "texture baking," while concurrently introducing Random Channel Tail Drop to mitigate the structural loss inherent in rigid dimensionality reduction, thereby reconciling realism with control consistency. Furthermore, to fully leverage DINOv3's high-resolution capabilities for enhancing control precision, we introduce a learnable Spatial Alignment Module that adapts these high-resolution features to the diffusion backbone. Finally, we propose a Causal Temporal Aggregator employing causal convolutions to explicitly preserve historical motion context when integrating frame-wise DINO features, which effectively mitigates motion blur and guarantees temporal stability. Project page: https://albertchen98.github.io/DwD-project/

cs.CV

Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning

Universal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H&E enriches H&E-based features with diverse, biologically meaningful information. However, progress is limited by the scarcity of well-aligned multi-stain datasets. Inter-stain misalignment shifts corresponding tissue across slides, hindering consistent patch-level features and degrading slide-level embeddings. To address this, we curated a slide-level aligned, five-stain dataset (H&E, HER2, KI67, ER, PGR) to enable paired H&E-IHC learning and robust cross-stain representation. Leveraging this dataset, we propose Cross-Stain Contrastive Learning (CSCL), a two-stage pretraining framework with a lightweight adapter trained using patch-wise contrastive alignment to improve the compatibility of H&E features with corresponding IHC-derived contextual cues, and slide-level representation learning with Multiple Instance Learning (MIL), which uses a cross-stain attention fusion module to integrate stain-specific patch features and a cross-stain global alignment module to enforce consistency among slide-level embeddings across different stains. Experiments on cancer subtype classification, IHC biomarker status classification, and survival prediction show consistent gains, yielding high-quality, transferable H&E slide-level representations. The code and data are available at https://github.com/lily-zyz/CSCL.

cs.CV

Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

cs.CL

Pure Core Sets of $n \times n$ Matrices over Finite Fields

This paper studies the structure of core sets under different similarity classes. We investigate the influence of factors of the minimal polynomial with different degrees on the structure of core sets. When $F$ is a finite field of prime order, we study the upper bound on the size of a non-core set in a similarity class in $M_n(F)$. We prove that as $|F|$ increases, the proportion of pure core sets among subsets of $M_n(F)$ tends to $1$.

math.RA

Ultralow Voltage Operation of p- and n-FETs Enabled by Self-Formed Gate Dielectric and Metal Contacts on 2D Tellurium

The ongoing demand for more energy-efficient, high-performance electronics is driving the exploration of innovative materials and device architectures, where interfaces play a crucial role due to the continuous downscaling of device dimensions. Tellurium (Te), in its two-dimensional (2D) form, offers significant potential due to its high carrier mobility and ambipolar characteristics, with the carrier type easily tunable via surface modulation. In this study, we leverage atomically controlled material transformations in 2D Te to create intimate junctions, enabling near-ideal field-effect transistors (FETs) for both n-type and p-type operation. A NiTex-Te contact provides highly transparent interfaces, resulting in low contact resistance, while the TiOx-Te gate dielectric forms an ultraclean interface with a capacitance equivalent to 0.88 nm equivalent oxide thickness (EOT), where the quantum capacitance of Te is observed. Subthreshold slopes (SS) approach the Boltzmann limit, with a record-low SS of 3.5 mV/dec achieved at 10 K. Furthermore, we demonstrate 2D Te-based complementary metal-oxide-semiconductor (CMOS) inverters operating at an ultralow voltage of 0.08 V with a voltage gain of 7.1 V/V. This work presents a promising approach to forming intimate dielectric/semiconductor and metal/semiconductor junctions for next-generation low-power electronic devices.

physics.app-ph

NTIRE 2024 Challenge on Low Light Image Enhancement: Methods and Results

This paper reviews the NTIRE 2024 low light image enhancement challenge, highlighting the proposed solutions and results. The aim of this challenge is to discover an effective network design or solution capable of generating brighter, clearer, and visually appealing results when dealing with a variety of conditions, including ultra-high resolution (4K and beyond), non-uniform illumination, backlighting, extreme darkness, and night scenes. A notable total of 428 participants registered for the challenge, with 22 teams ultimately making valid submissions. This paper meticulously evaluates the state-of-the-art advancements in enhancing low-light images, reflecting the significant progress and creativity in this field.

cs.CV

Data Player: Automatic Generation of Data Videos with Narration-Animation Interplay

Data visualizations and narratives are often integrated to convey data stories effectively. Among various data storytelling formats, data videos have been garnering increasing attention. These videos provide an intuitive interpretation of data charts while vividly articulating the underlying data insights. However, the production of data videos demands a diverse set of professional skills and considerable manual labor, including understanding narratives, linking visual elements with narration segments, designing and crafting animations, recording audio narrations, and synchronizing audio with visual animations. To simplify this process, our paper introduces a novel method, referred to as Data Player, capable of automatically generating dynamic data videos with narration-animation interplay. This approach lowers the technical barriers associated with creating data videos rich in narration. To enable narration-animation interplay, Data Player constructs references between visualizations and text input. Specifically, it first extracts data into tables from the visualizations. Subsequently, it utilizes large language models to form semantic connections between text and visuals. Finally, Data Player encodes animation design knowledge as computational low-level constraints, allowing for the recommendation of suitable animation presets that align with the audio narration produced by text-to-speech technologies. We assessed Data Player's efficacy through an example gallery, a user study, and expert interviews. The evaluation results demonstrated that Data Player can generate high-quality data videos that are comparable to human-composed ones.

cs.HC

See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation

Humans use all of their senses to accomplish different tasks in everyday activities. In contrast, existing work on robotic manipulation mostly relies on one, or occasionally two modalities, such as vision and touch. In this work, we systematically study how visual, auditory, and tactile perception can jointly help robots to solve complex manipulation tasks. We build a robot system that can see with a camera, hear with a contact microphone, and feel with a vision-based tactile sensor, with all three sensory modalities fused with a self-attention model. Results on two challenging tasks, dense packing and pouring, demonstrate the necessity and power of multisensory perception for robotic manipulation: vision displays the global status of the robot but can often suffer from occlusion, audio provides immediate feedback of key moments that are even invisible, and touch offers precise local geometry for decision making. Leveraging all three modalities, our robotic system significantly outperforms prior methods.

cs.RO

A Nanometer-Thick Oxide Semiconductor Transistor with Ultra-High Drain Current

High drive current is a critical performance parameter in semiconductor devices for high-speed, low-power logic applications or high-efficiency, high-power, high-speed radio frequency (RF) analog applications. In this work, we demonstrate an In2O3 transistor grown by atomic layer deposition (ALD) at back-end-of-line (BEOL) compatible temperatures with a record high drain current exceeding 10 A/mm, the performance of which is 2-3 times better than all known transistors with semiconductor channels. A record high transconductance of 4 S/mm is also achieved among all transistors with a planar structure. It is found that a high carrier density and high electron velocity both contribute to this remarkably high on-state performance in ALD In2O3 transistors, which is made possible by the high-quality oxide/oxide interface, the metal-like charge-neutrality-level (CNL) alignment, and the high band velocities induced by the low density-of-state (DOS). Experimental Hall, I-V and split C-V measurements at room temperature confirm a high carrier density up to 6-7*10^13 /cm2 and a high velocity of about 10^7 cm/s. Ultra-thin oxide semiconductors, with a CNL located deep inside the conduction band, represent a promising new direction for the search of alternative channel materials for high-performance semiconductor devices.

cond-mat.mes-hall

EDAssistant: Supporting Exploratory Data Analysis in Computational Notebooks with In-Situ Code Search and Recommendation

Using computational notebooks (e.g., Jupyter Notebook), data scientists rationalize their exploratory data analysis (EDA) based on their prior experience and external knowledge such as online examples. For novices or data scientists who lack specific knowledge about the dataset or problem to investigate, effectively obtaining and understanding the external information is critical to carry out EDA. This paper presents EDAssistant, a JupyterLab extension that supports EDA with in-situ search of example notebooks and recommendation of useful APIs, powered by novel interactive visualization of search results. The code search and recommendation are enabled by state-of-the-art machine learning models, trained on a large corpus of EDA notebooks collected online. A user study is conducted to investigate both EDAssistant and data scientists' current practice (i.e., using external search engines). The results demonstrate the effectiveness and usefulness of EDAssistant, and participants appreciated its smooth and in-context support of EDA. We also report several design implications regarding code recommendation tools.

cs.HC

Intrinsic ferroelectricity in Y-doped HfO2 thin films

Ferroelectric HfO2-based materials hold great potential for widespread integration of ferroelectricity into modern electronics due to their robust ferroelectric properties at the nanoscale and compatibility with the existing Si technology. Earlier work indicated that the nanometer crystal grain size was crucial for stabilization of the ferroelectric phase of hafnia. This constraint caused high density of unavoidable structural defects of the HfO2-based ferroelectrics, obscuring the intrinsic ferroelectricity inherited from the crystal space group of bulk HfO2. Here, we demonstrate the intrinsic ferroelectricity in Y-doped HfO2 films of high crystallinity. Contrary to the common expectation, we show that in the 5% Y-doped HfO2 epitaxial thin films, high crystallinity enhances the spontaneous polarization up to a record-high 50 {\mu}C/cm2 value at room temperature. The high spontaneous polarization persists at reduced temperature, with polarization values consistent with our theoretical predictions, indicating the dominant contribution from the intrinsic ferroelectricity. The crystal structure of these films reveals the Pca21 orthorhombic phase with a small rhombohedral distortion, underlining the role of the anisotropic stress and strain. These results open a pathway to controlling the intrinsic ferroelectricity in the HfO2-based materials and optimizing their performance in applications.

cond-mat.mtrl-sci

Direct Observation of Room-Temperature Dislocation Plasticity in Diamond

It is well known that diamond does not deform plastically at room temperature and usually fails in catastrophic brittle fracture. Here we demonstrate room-temperature dislocation plasticity in sub-micrometer sized diamond pillars by in-situ mechanical testing in the transmission electron microscope. We document in unprecedented details of spatio-temporal features of the dislocations introduced by the confinement-free compression, including dislocation generation and propagation. Atom-resolved observations with tomographic reconstructions show unequivocally that mixed-type dislocations with Burgers vectors of 1/2<110> are activated in the non-close-packed {001} planes of diamond under uniaxial compression of <111> and <110> directions, respectively, while being activated in the {111} planes under the <100> directional loading, indicating orientation-dependent dislocation plasticity. These results provide new insights into the mechanical behavior of diamond and stimulate reconsideration of the basic deformation mechanism in diamond as well as in other brittle covalent crystals at low temperatures.

cond-mat.mtrl-sci

A Compact Design of Four-degree-of-freedom Transmission Electron Microscope Holder for Quasi-Four-Dimensional Characterization

Electron tomography (ET) has been demonstrated to be a powerful tool in addressing challenging problems, such as understanding 3D interactions among various microstructures. Advancing ET to broader applications requires novel instrumentation design to break the bottlenecks both in theory and in practice. In this work, we built a compact four-degree-of-freedom (three-directional positionings plus self-rotation) nano-manipulator dedicated to ET applications, which is called X-Nano transmission electron microscope (TEM) holder. All the movements of the four degrees of freedom are precisely driven by built-in piezoelectric actuators, minimizing the artefacts due to the vibration and drifting of the TEM stage. Full 360o rotation is realized with an accuracy of 0.05o in the whole range, which solves the missing wedge problem. Meanwhile, the specimen can move to the rotation axis with an integrated 3D nano-manipulator, greatly reducing the effort in tracking sample locations during tilting. Meanwhile, in-situ stimulation function can be seamlessly integrated into the X-Nano TEM holder so that dynamic information can be uncovered. We expect that more delicate researches, such as those about 3D microstructural evolution, can be carried out extensively by means of this holder in the near future.

physics.app-ph

GAIL---Guaranteed Automatic Integration Library in MATLAB: Documentation for Version 2.1

Automatic and adaptive approximation, optimization, or integration of functions in a cone with guarantee of accuracy is a relatively new paradigm. Our purpose is to create an open-source MATLAB package, Guaranteed Automatic Integration Library (GAIL), following the philosophy of reproducible research and sustainable practices of robust scientific software development. For our conviction that true scholarship in computational sciences are characterized by reliable reproducibility, we employ the best practices in mathematical research and software engineering known to us and available in MATLAB. This document describes the key features of functions in GAIL, which includes one-dimensional function approximation and minimization using linear splines, one-dimensional numerical integration using trapezoidal rule, and last but not least, mean estimation and multidimensional integration by Monte Carlo methods or Quasi Monte Carlo methods.

cs.MS