SearcharxivSearch

arXiv subjects

Yujia Shi

Publications and source records attributed to Yujia Shi.

13 recordsLinked to original sources

Optimal and Deterministic Quantum Search on the Simplex of Complete Graphs

The simplex of complete graphs, also known as the first-order truncated simplex lattice, is a network of $M+1$ identical complete graphs, each with $M$ vertices, such that each clique contains an edge or bridge to every other clique. It contains $N = M(M+1)$ vertices, and previous asymptotic results using a continuous-time quantum walk to search this graph for a single marked vertex have either numerically demonstrated an optimal runtime of $O(\sqrt{N})$, or analytically proved a deterministic success probability of 1, but not both, even when the bridges are weighted. In this paper, we give the first analytical proof of optimal quantum search on this graph, proving that it occurs when the weight of the bridges equals $M$. In addition, we numerically show that the optimal runtime is achieved more broadly whenever the weight is at least $\sqrt{M}$. Furthermore, the algorithm is also deterministic when the weight scales between $\sqrt{M}$ and $M$, and this is the first example of quantum search on the simplex of complete graphs that is both asymptotically optimal and deterministic. In addition, for weights where the algorithm is nondeterministic, we give a way to find the marked vertex by inspecting neighboring vertices. Finally, while it is known that connectivity is not a reliable indicator of fast quantum search when comparing different graph families, we show that it is also unreliable within the graph family of weighted simplex of complete graphs.

quant-ph

EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with full-hand pressure supervision for diverse everyday objects, incorporating a bare-hand transfer subset to enable generalization to natural scenarios. Leveraging this benchmark, we first establish EgoPressureFormer as a discriminative baseline. Beyond this, to explicitly address the uncertainty in partial observations, we propose EgoPressureDiff, a conditional diffusion framework that adapts a large-scale pre-trained video diffusion backbone. By combining rich world knowledge priors with a Physically-Informed Feature Rectification layer to inject semantic constraints, our approach effectively infers plausible contact patterns and resolves visual-physical ambiguities. Extensive experiments demonstrate that our method achieves superior performance on the benchmark and robust transferability to in-the-wild scenarios. Our project page is available at https://egotactile.github.io/.

cs.CV

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing approaches often rely on pose estimators to extract intermediate representations, but such signals are prone to errors under occlusion or complex poses. Building on these observations, we present DirectAnimator, a framework that bypasses pose extraction and directly learns from raw driving videos. We introduce a Driving Cue Triplet consisting of pose, face, and location cues that captures motion, expression, and alignment in a semantically rich yet stable form, and we fuse them through a CueFusion DiT block for reliable control during denoising. To make learning dependable when the driving and reference identities differ, we devise a Same2X training strategy that aligns cross-ID features with those learned from same-ID data, regularizing optimization and accelerating convergence. Extensive experiments demonstrate that DirectAnimator attains state-of-the-art visual quality and identity preservation while remaining robust to occlusions and complex articulation, and it does so with fewer computational resources. Our project page is at https://directanimator.github.io/.

cs.CV

FreeAnimate: Training-Free Human Image Animation with Preview-Guided Denoising

Human Image Animation has seen significant advancements, primarily driven by diffusion models. However, existing methods typically demand substantial training data and resources to achieve high-quality results, limiting generalization and accessibility. In this work, we introduce \emph{FreeAnimate}, a training-free framework that leverages the inherent capabilities of image diffusion models to enable temporal consistency, identity preservation, and background stability. Our approach incorporates a novel preview generation strategy that provides temporal and structural priors from generated preview frames, effectively guiding pose alignment and background consistency without training. Additionally, FreeAnimate introduces Inversion-Boosted Attention and Reference-Anchored Self-Attention modules to guarantee temporal consistency and identity preservation. Experimental results demonstrate that FreeAnimate outperforms existing training-free competitors and training-based baseline methods, achieving generation quality comparable to state-of-the-art methods and offering robust generalization across diverse datasets. Our project page is at https://freeani.github.io/.

cs.CV

EgoPressDiff: Multimodal Video Diffusion for Egocentric UV-Domain Hand-Pressure Estimation

Estimating hand-surface contact pressure from an egocentric view is crucial for AR/VR devices, robotic imitation, and ergonomic analysis. Existing methods often discretize pressure signal and process frames independently, leading to quantization errors and temporal inconsistencies. We present \emph{EgoPressDiff}, a conditional video diffusion framework that generates UV-pressure maps from visual input. The core of our approach is a multi-modal conditioning strategy, introducing a PoseNet and a Vertex Encoder to efficiently extract features from hand pose and 3D mesh vertices. These signals, along with depth information, guide the generative process to ensure the pressure fields are physically grounded. To effectively fuse these heterogeneous features, we further propose a Distribution-Calibrated Spatial Layer, which aligns their statistical properties before combination. Evaluated on the EgoPressure ego-view setting, EgoPressDiff achieves state-of-the-art results, improving Volumetric IoU by over 34\% relative to prior baseline, while reducing MAE and maintaining high temporal accuracy. Our project page is at https://egopressdiff.github.io/.

cs.CV

Self-Trapping Bounds for Continuous-Time Nonlinear Quantum Walks on Path and Cycle Graphs

We explore a continuous-time quantum walk starting at a single vertex on the discrete path and cycle with a cubic nonlinearity. Such nonlinearities arise in Bose-Einstein condensates described by the Gross-Pitaevskii equation or by nonlinear optical waveguide arrays. When the nonlinearity is sufficiently strong, the walker remains localized at its initial vertex, a phenomenon known as self-trapping. This contrasts with linear quantum walks, which are known for spreading quickly in one dimension. While self-trapping has been known numerically, we introduce an analytical method that proves self-trapping and yields a quantitative relationship between the nonlinearity coefficient and the trapping probability. We propose that this trapping can be used for timing in quantum state transfer, where a qubit is held at a node until it is ready to be transferred, and it can also be held again at the receiving node. This scheme can also be interpreted as a form of quantum memory, with the trap and transfer corresponding to the storage and release of quantum information.

quant-ph

Fast state transfer via loop weights

We prove that almost-linear-time high-fidelity state transfer is achievable in a quantum spin chain using loop weights at the second and second-to-last nodes. We provide specific parameter values, and using a careful analysis of the eigenvectors we make precise quantitative estimates of the transfer time and strength.

quant-ph

Resilience by Design: A KPI for Heavy-Duty Megawatt Charging

We introduce a stressor-agnostic Resilience Key Performance Indicator (Resilience KPI) for megawatt charging stations (MSC) serving heavy-duty vehicles. Beyond routine performance statistics (e.g., availability, throughput), the KPI quantifies a site's ability to anticipate, operate under degradation, and recover from disruptions using observable signals already in the framework: ride-through capability, restoration speed, service under N-1, expected unserved charging energy, and queue impacts. The headline score is normalised to 0-100 for fair cross-site and cross-vendor benchmarking, with optional stressor-specific breakouts (grid, ICT, thermal, flooding, on-site incidents) for diagnostics and robustness checks. DATEX II provides a solid baseline for resilience KPIs centred on infrastructure inventory, status, and pricing, while additional KPIs, especially around grid capacity, on-site flexibility, heavy-vehicle geometry, environmental hardening, maintenance, and market exposure, are essential for a complete resilience picture and will require extensions or complementary data sources. The KPI is designed for monthly/quarterly reporting to support design and operational decisions and cost-benefit assessment of mitigations (e.g., backup power, spares, procedures). It offers a consistent, transparent methodology that consolidates heterogeneous logs and KPIs into a single, auditable indicator, making resilience comparable across sites, vendors, and jurisdictions.

cs.ET

Continuous-Time Quantum State Transfer with a Generalized Laplacian

Quantum walks generated by the adjacency matrix or the Laplacian are known to exhibit low transfer fidelity on general graphs. In this paper, we study continuous-time quantum walks governed by the generalized Laplacian operator L_k = A+kD, where A is the adjacency matrix, D is the degree matrix, and k is a real-valued parameter. Recent work of Duda, McLaughlin, and Wong showed that in the single-excitation Heisenberg (XYZ) spin model, one can realize walks generated by this family of operators on signed weighted graphs. Motivated by earlier studies on vertex-weighted graphs, we demonstrate that for certain graphs, tuning the parameter k can significantly enhance the fidelity of state transfer between endpoints.

quant-ph

Data Aware Differentiable Neural Architecture Search for Tiny Keyword Spotting Applications

The success of Machine Learning is increasingly tempered by its significant resource footprint, driving interest in efficient paradigms like TinyML. However, the inherent complexity of designing TinyML systems hampers their broad adoption. To reduce this complexity, we introduce "Data Aware Differentiable Neural Architecture Search". Unlike conventional Differentiable Neural Architecture Search, our approach expands the search space to include data configuration parameters alongside architectural choices. This enables Data Aware Differentiable Neural Architecture Search to co-optimize model architecture and input data characteristics, effectively balancing resource usage and system performance for TinyML applications. Initial results on keyword spotting demonstrate that this novel approach to TinyML system design can generate lean but highly accurate systems.

cs.LG

Strong quantum state transfer on graphs via loop edges

We quantify the effect of weighted loops at the source and target nodes of a graph on the strength of quantum state transfer between these vertices. We give lower bounds on loop weights that guarantee strong transfer fidelity that works for any graph where this protocol is feasible. By considering local spectral symmetry, we show that the required weight size depends only on the maximum degree of the graph and, in some less favorable cases, the distance between vertices. Additionally, we explore the duration for which transfer strength remains above a specified threshold.

quant-ph

Quantifying State Transfer Strength on Graphs with Involution

This paper discusses continuous-time quantum walks and asymptotic state transfer in graphs with an involution. By providing quantitative bounds on the eigenvectors of the Hamiltonian, it provides an approach to achieving high-fidelity state transfer by strategically selecting energy potentials based on the maximum degrees of the graphs. The study also involves an analysis of the time necessary for quantum transfer to occur.

quant-ph

Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision

Violence detection has been studied in computer vision for years. However, previous work are either superficial, e.g., classification of short-clips, and the single scenario, or undersupplied, e.g., the single modality, and hand-crafted features based multimodality. To address this problem, in this work we first release a large-scale and multi-scene dataset named XD-Violence with a total duration of 217 hours, containing 4754 untrimmed videos with audio signals and weak labels. Then we propose a neural network containing three parallel branches to capture different relations among video snippets and integrate features, where holistic branch captures long-range dependencies using similarity prior, localized branch captures local positional relation using proximity prior, and score branch dynamically captures the closeness of predicted score. Besides, our method also includes an approximator to meet the needs of online detection. Our method outperforms other state-of-the-art methods on our released dataset and other existing benchmark. Moreover, extensive experimental results also show the positive effect of multimodal (audio-visual) input and modeling relationships. The code and dataset will be released in https://roc-ng.github.io/XD-Violence/.

cs.CV