SearcharxivSearch

arXiv subjects

Jinpeng Liu

Publications and source records attributed to Jinpeng Liu.

15 recordsLinked to original sources

In-situ Indexing via Memristive Content-Addressable Memory

Processing-in-Memory (PIM) is a proven paradigm for overcoming the ``memory wall". However, while data indexing is severely bottlenecked by this same wall, it remains unclear how indexing can effectively benefit from PIM's unique capabilities. We present PATH, an in-situ indexing architecture that bridges this gap by leveraging the massive parallelism and inherent data-movement of PIMs. Specifically, we first reformulate the fundamental indexing operations, namely Insert, Search, Update, and Delete, into highly parallel in-situ content-addressable memory operations executed directly within memory arrays. Taking hash indexes as a typical case, we elaborate how PATH breaks the inherent trade-off among memory accesses, load factor, and process latency in conventional hashing schemes. By adopting ultra-large logical buckets and in-memory moving, PATH virtually eliminates the cost of hash collision resolution and significantly reduces resizing overhead. Compared with state-of-the-art schemes, PATH achieves $4.7-7.8\times$ higher throughput, $>14.5\times$ lower tail latency, and $>61.4\%$ fewer memory accesses under insertions, laying a scalable foundation for next-generation data-centric computing.

cs.AR

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, and cameras, making them unable to recover coherent geometry, stable motion, and physically aligned trajectories. These limitations motivate us to introduce a new task: unified human-scene-camera reconstruction from multi-view videos, which aims to jointly estimate dynamic humans, static scenes, and camera poses in one global coordinate frame. We propose TROPHIES--Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos-a unified framework tailored for this task. TROPHIES features a Human Branch that models humans through temporal and spatial reasoning, and a Scene Branch that reconstructs static geometry with human-aware attention. A global alignment and optimization module couples both branches by enforcing scale consistency, contact priors, and cross-view temporal coherence. Experiments on EgoHuman and EgoExo4D demonstrate that TROPHIES achieves globally aligned, physically plausible 4D reconstructions and consistently outperforms existing paradigms in both global fidelity and human-scene consistency.

cs.CV

Oxygen-vacancy quantum spin defects in silicon carbide

Optically addressable spin defects in wide-bandgap semiconductors are promising building blocks for quantum sensing and quantum networks. Establishing their atomic structure is essential for understanding functionality and enabling controlled engineering. In 4H-SiC, the PL5 and PL6 centers have long been recognized for their exceptional charge stability and room-temperature optically detected magnetic resonance (ODMR) performance, but their structural origin has remained elusive for over a decade. Here, we provide direct evidence for their oxygen-vacancy (${\rm O_C V_{Si}}$) origins through a combined chemical and isotopic control strategy. Under oxygen ion implantation, we observe over tenfold enhancement in the yield of PL5 and PL6 compared to nitrogen ion implantation. Furthermore, implantation with $^{17}{\rm O}$ ions produces PL5 and PL6 defects that exhibit a characteristic six-fold $^{17}{\rm O}$ hyperfine splitting in their ODMR spectra. These results affirm PL6 as the ${\rm O_C V_{Si}}$ defect in the $hh$ configuration. For PL5, the oxygen-related evidence, together with \textit{ab initio} calculations and additional measurements of the zero-field splitting and hyperfine structure, establishes it as the ${\rm O_C V_{Si}}$ defect in the $kh$ configuration. This unambiguous structural identification, achieved through materials-level chemical control, provides the microscopic foundation for deterministic engineering of these defects, paving the way for scalable photonic devices and high-sensitivity ensemble quantum sensors based on oxygen-vacancy centers.

cond-mat.mtrl-sci

Quantum Noise Spectroscopy of Nanoscale Charge Defects in Silicon Carbide at Room Temperature

The nanoscale charge environment critically influences semiconductor physics and device performance. While conventional bulk characterization techniques provide volume-averaged defect properties, they lack the spatial resolution to resolve nanoscale charge heterogeneity and identify microscopic noise sources. Here, we utilize single PL5 centers in 4H-SiC as room-temperature broadband quantum sensors to fill in the gap. We report the first real-time, nanoscale observation of singlecharge tunneling dynamics in a commercial semiconductor at room temperature, by monitoring the random telegraph noise using optically detected magnetic resonance (ODMR). This capability enables an electrical noise imaging technique, showing distinct noise variations across different wafer substrates. By employing dynamical decoupling, we extend noise spectroscopy from near-DC to MHz frequencies, uncovering significant noise spectral density correlations across frequency bands. Finally, we probe MHz-GHz noise and identify its origin via T1 relaxation spectroscopy, obtaining the first nanoscale electron paramagnetic resonance (EPR) spectroscopic fingerprint of charge defects in SiC. These techniques open avenues for characterizing noise environments in semiconductor devices, providing critical insights for optimizing SiC fabrication processes, defect control, and advancing quantum technologies.

quant-ph

Single-molecule Scale Nuclear Magnetic Resonance Spectroscopy using a Robust Near-Infrared Spin Sensor

Nuclear magnetic resonance (NMR) at the single-molecule level with atomic resolution holds transformative potential for structural biology and surface chemistry. Near-surface solid-state spin sensors with optical readout ability offer a promising pathway toward this goal. However, their extreme proximity to target molecules demands exceptional robustness against surface-induced perturbations. Furthermore, life science applications require these sensors to operate in biocompatible spectral ranges that minimize photodamage. In this work, we demonstrate that the PL6 quantum defect in 4H silicon carbide (4H-SiC) can serve as a robust near-infrared spin sensor. This sensor operates at tissue-transparent wavelengths and exhibits exceptional near-surface stability even at depth of 2 nm. Using shallow PL6 centers, we achieve nanoscale NMR detection of proton ($\mathrm{^{1}H}$) spins in immersion oil and fluorine ($\mathrm{^{19}F}$) spins in Fomblin, attaining a detection volume of $\mathrm{(3~nm)^3}$ and a sensitivity reaching the requirement for single-proton spin detection. This work establishes 4H-SiC quantum sensors as a compelling platform for nanoscale magnetic resonance, with promising applications in probing low-dimensional water phases, protein folding dynamics, and molecular interactions.

quant-ph

ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion

Joint reconstruction of human-object interaction marks a significant milestone in comprehending the intricate interrelations between humans and their surrounding environment. Nevertheless, previous optimization methods often struggle to achieve physically plausible reconstruction results due to the lack of prior knowledge about human-object interactions. In this paper, we introduce ScoreHOI, an effective diffusion-based optimizer that introduces diffusion priors for the precise recovery of human-object interactions. By harnessing the controllability within score-guided sampling, the diffusion model can reconstruct a conditional distribution of human and object pose given the image observation and object feature. During inference, the ScoreHOI effectively improves the reconstruction results by guiding the denoising process with specific physical constraints. Furthermore, we propose a contact-driven iterative refinement approach to enhance the contact plausibility and improve the reconstruction accuracy. Extensive evaluations on standard benchmarks demonstrate ScoreHOI's superior performance over state-of-the-art methods, highlighting its ability to achieve a precise and robust improvement in joint human-object interaction reconstruction.

cs.CV

Potentials of skip-free Markov chains

Potential theory has important applications in various fields such as physics, finance, and biology. In this paper, we investigate the potentials of two classic types of discrete-time skip-free Markov chains: upward skip-free and downward skip-free Markov chains. The key to deriving these potentials lies in the use of truncation approximation techniques. The results are then applied to GI/M/1 queues and M/G/1 queues, and further extended to continuous-time skip-free Markov chains.

math.PR

MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model

This work introduces MotionLCM, extending controllable motion generation to a real-time level. Existing methods for spatial-temporal control in text-conditioned motion generation suffer from significant runtime inefficiency. To address this issue, we first propose the motion latent consistency model (MotionLCM) for motion generation, building on the motion latent diffusion model. By adopting one-step (or few-step) inference, we further improve the runtime efficiency of the motion latent diffusion model for motion generation. To ensure effective controllability, we incorporate a motion ControlNet within the latent space of MotionLCM and enable explicit control signals (i.e., initial motions) in the vanilla motion space to further provide supervision for the training process. By employing these techniques, our approach can generate human motions with text and control signals in real-time. Experimental results demonstrate the remarkable generation and controlling capabilities of MotionLCM while maintaining real-time runtime efficiency.

cs.CV

NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model

We introduce NovelGS, a diffusion model for Gaussian Splatting (GS) given sparse-view images. Recent works leverage feed-forward networks to generate pixel-aligned Gaussians, which could be fast rendered. Unfortunately, the method was unable to produce satisfactory results for areas not covered by the input images due to the formulation of these methods. In contrast, we leverage the novel view denoising through a transformer-based network to generate 3D Gaussians. Specifically, by incorporating both conditional views and noisy target views, the network predicts pixel-aligned Gaussians for each view. During training, the rendered target and some additional views of the Gaussians are supervised. During inference, the target views are iteratively rendered and denoised from pure noise. Our approach demonstrates state-of-the-art performance in addressing the multi-view image reconstruction challenge. Due to generative modeling of unseen regions, NovelGS effectively reconstructs 3D objects with consistent and sharp textures. Experimental results on publicly available datasets indicate that NovelGS substantially surpasses existing image-to-3D frameworks, both qualitatively and quantitatively. We also demonstrate the potential of NovelGS in generative tasks, such as text-to-3D and image-to-3D, by integrating it with existing multiview diffusion models. We will make the code publicly accessible.

cs.CV

Hoeffding's inequality for continuous-time Markov chains

Hoeffding's inequality is a fundamental tool widely applied in probability theory, statistics, and machine learning. In this paper, we establish Hoeffding's inequalities specifically tailored for an irreducible and positive recurrent continuous-time Markov chain (CTMC) on a countable state space with the invariant probability distribution $π$ and an $\mathcal{L}^{2}(π)$-spectral gap $λ(Q)$. More precisely, for a function $g:E\to [a,b]$ with a mean $π(g)$, and given $t,\varepsilon>0$, we derive the inequality \[ \mathbb{P}_π\left(\frac{1}{t} \int_{0}^{t} g\left(X_{s}\right)\mathrm{d}s-π(g) \geq \varepsilon \right) \leq \exp\left\{-\frac{λ(Q)t\varepsilon^2}{(b-a)^2} \right\}, \] which can be viewed as a generalization of Hoeffding's inequality for discrete-time Markov chains (DTMCs) presented in [J. Fan et al., J. Mach. Learn. Res., 22(2022), pp. 6185-6219] to the realm of CTMCs. The key analysis enabling the attainment of this inequality lies in the utilization of the techniques of skeleton chains and augmented truncation approximations. Furthermore, we also discuss Hoeffding's inequality for a jump process on a general state space.

math.PR

Plan, Posture and Go: Towards Open-World Text-to-Motion Generation

Conventional text-to-motion generation methods are usually trained on limited text-motion pairs, making them hard to generalize to open-world scenarios. Some works use the CLIP model to align the motion space and the text space, aiming to enable motion generation from natural language motion descriptions. However, they are still constrained to generate limited and unrealistic in-place motions. To address these issues, we present a divide-and-conquer framework named PRO-Motion, which consists of three modules as motion planner, posture-diffuser and go-diffuser. The motion planner instructs Large Language Models (LLMs) to generate a sequence of scripts describing the key postures in the target motion. Differing from natural languages, the scripts can describe all possible postures following very simple text templates. This significantly reduces the complexity of posture-diffuser, which transforms a script to a posture, paving the way for open-world generation. Finally, go-diffuser, implemented as another diffusion model, estimates whole-body translations and rotations for all postures, resulting in realistic motions. Experimental results have shown the superiority of our method with other counterparts, and demonstrated its capability of generating diverse and realistic motions from complex open-world prompts such as "Experiencing a profound sense of joy". The project page is available at https://moonsliu.github.io/Pro-Motion.

cs.CV

FLAG3D: A 3D Fitness Activity Dataset with Language Instruction

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing hunger for data resources involved in high-quality data, fine-grained labels, and diverse environments. In this paper, we present FLAG3D, a large-scale 3D fitness activity dataset with language instruction containing 180K sequences of 60 categories. FLAG3D features the following three aspects: 1) accurate and dense 3D human pose captured from advanced MoCap system to handle the complex activity and large movement, 2) detailed and professional language instruction to describe how to perform a specific activity, 3) versatile video resources from a high-tech MoCap system, rendering software, and cost-effective smartphones in natural environments. Extensive experiments and in-depth analysis show that FLAG3D contributes great research value for various challenges, such as cross-domain human action recognition, dynamic human mesh recovery, and language-guided human action generation. Our dataset and source code are publicly available at https://andytang15.github.io/FLAG3D.

cs.CV

Toward Robust Diagnosis: A Contour Attention Preserving Adversarial Defense for COVID-19 Detection

As the COVID-19 pandemic puts pressure on healthcare systems worldwide, the computed tomography image based AI diagnostic system has become a sustainable solution for early diagnosis. However, the model-wise vulnerability under adversarial perturbation hinders its deployment in practical situation. The existing adversarial training strategies are difficult to generalized into medical imaging field challenged by complex medical texture features. To overcome this challenge, we propose a Contour Attention Preserving (CAP) method based on lung cavity edge extraction. The contour prior features are injected to attention layer via a parameter regularization and we optimize the robust empirical risk with hybrid distance metric. We then introduce a new cross-nation CT scan dataset to evaluate the generalization capability of the adversarial robustness under distribution shift. Experimental results indicate that the proposed method achieves state-of-the-art performance in multiple adversarial defense and generalization tasks. The code and dataset are available at https://github.com/Quinn777/CAP.

eess.IV

Matrix-analytic methods for solving Poisson's equation with applications to Markov chains of GI/G/1-type

In this paper, we are devoted to developing matrix-analytic methods for solving Poisson's equation for irreducible and positive recurrent discrete-time Markov chains (DTMCs). Two special solutions, including the deviation matrix D and the expected additive-type functional matrix K, will be considered. The results are applied to Markov chains of GI/G/1-type and MAP/G/1 queues with negative customers. Further extensions to continuous-time Markov chains (CTMCs) are also investigated.

math.PR

Augmented truncation approximations to the solution of Poisson's equation for Markov chains

Poisson's equation has a lot of applications in various areas. Usually it is hard to derive the explicit expression of the solution of Poisson's equation for a Markov chain on an infinitely many state space. We will present a computational framework for the solution for both discrete-time Markov chains (DTMCs) and continuous-time Markov chains (CTMCs), by developing the technique of augmented truncation approximations. The convergence to the solution is investigated in terms of the assumption about the monotonicity of the first return times, and is further established for two types of truncation approximation schemes: the censored chain and the linear augmented truncation. Moreover, truncation approximations to the variance constant in central limit theorems (CLTs) are also considered. The results obtained are applied to discrete-time single-birth processes and continuous-time single-death processes.

math.PR