SearcharxivSearch

arXiv subjects

Junjie Liu

Publications and source records attributed to Junjie Liu.

At least 19 recordsLinked to original sources

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two fundamental limitations. First, most approaches operate exclusively post-vision encoder, leaving the substantial latency of the visual encoding phase unoptimized. Second, under strict token budgets, these methods often fail to jointly preserve holistic visual contexts and fine-grained details, leading to performance degradation. To address these bottlenecks, we propose PACE (Pixel-Adaptive Condense and Extract), a training-free inference framework that accelerates both the vision encoder and the Large Language Model (LLM) via a unified Condense-and-Extract paradigm. During the Condense stage, an Adaptive Pixel Compressor (APC) evaluates visual information density prior to encoding, adaptively downsampling redundant inputs, curtailing encoder computation while preserving global context and essential visual cues. In the Extract stage, a Dynamic Dual-Attention Extractor (DDAE) selectively retains visual tokens via a fusion of internal visual signals from the encoder and semantic signals from the LLM, safeguarding task-critical details. By integrating PACE into Qwen2.5-VL-7B, the model retains 93.8% of its original performance while utilizing only 10% of the visual tokens, yielding a 3.1x speedup in time to first token (TTFT). Our code is available at https://github.com/jjL357/PACE.

cs.CV

Dynamical one-from-many quantum metrology: Sum rule and matrix-free precision bound

Practical quantum single-parameter estimation is rarely a pristine task; it almost invariably involves nuisance parameters, casting it as a one-from-many problem. Existing approaches to this problem rely on multi-parameter metrology, reducing the matrix quantum Cram\'er-Rao bound to obtain scalar quantum precision limits. However, these methods are often hampered by the demanding inversion of the quantum Fisher information (QFI) matrix and the requisite choice of a weight matrix, and they break down when the QFI matrix becomes singular. Here, we show that for dynamical one-from-many estimation, a previously overlooked sum rule connecting the QFI about all model parameters to the QFI about time necessitates including the latter to consider an augmented QFI matrix while simultaneously rendering it inherently singular--precisely the scenario where conventional approaches fail. To meet this challenge, we derive a tight, matrix-free quantum precision bound that involves only scalar quantities, offers broad applicability, and subsumes existing results as special cases. Validated in both unitary and noisy settings, our findings provide a refined operational framework for practical quantum single-parameter metrology.

quant-ph

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.

cs.CV

Coherent electric field manipulation of nuclear spin qudit

Spins in condensed matter, especially well-isolated nuclear spins, offer attractive quantum degrees of freedom for computing, sensing, and networking because of their long coherence times. The possibility of electric-field control is an important feature for practical scalable quantum technologies, but, typically, nuclear spins couple only weakly to electric fields in conventional semiconductor hosts, limiting operation efficiency. Here we show that a choice of a highly polarizable oxide host can overcome this bottleneck. In Mn2+ doped ZnO, electric-field modulation of the spin Hamiltonian is amplified by hyperfine-coupled electron spins, and offers efficient electric-field manipulation of an I = 5/2 nuclear spin qudit, in a manner analogous to the hyperfine enhancement of conventional nuclear magnetic resonance. We demonstrate both resonant and non-resonant coherent manipulation using a single uniaxial electric field applied along the crystalline c-axis, the polarization axis of ZnO. This approach allows universal single-qudit gate operations with efficiencies comparable to or exceeding those of conventional magnetic-field driving. These results support the deployment of doped oxides as active host materials for electrically controllable spin qubits, highlighting the importance of materials design in developing scalable quantum technologies.

quant-ph

NanoBTE: Fast Iterative Solution of the Phonon Boltzmann Transport Equation for Nanoscale Heat Transport

Nanoscale heat dissipation has become a critical challenge in advanced semiconductor devices, where phonon transport can strongly deviate from the classical Fourier description due to the boundary scattering and ballistic effects. In this work, we propose NanoBTE, a deterministic finite-volume solver for the non-gray phonon Boltzmann transport equation under the relaxation-time approximation. The solver supports complex two- and three-dimensional geometries, band-resolved phonon properties, discrete-ordinates angular quadrature, volumetric heat generation, and multiple phonon boundary conditions, including thermalizing, diffuse, and specular reflections. %To improve the efficiency of multiscale simulations, both sequential and synthetic iterative schemes are implemented, where the latter couples the microscopic phonon transport equation with a macroscopic diffusion-type temperature equation to accelerate convergence in near-diffusive regimes. Both sequential and synthetic iterative options are implemented for the steady-state solution. Furthermore, NanoBTE adopts a band-direction task decomposition strategy, enabling efficient MPI-based CPU parallelization and GPU acceleration of the dominant sparse transport operations.

cond-mat.mtrl-sci

Initial-state-dependent dephasing effect in non-Hermitian Su-Schrieffer-Heeger models

Understanding the dynamical evolution of non-Hermitian systems under extra external dissipation is essential. Dephasing, a major realistic dissipation, is conventionally considered detrimental to information processing. However, its impact on non-Hermitian systems remains largely unexplored. Here, we focus on finite-sized non-Hermitian Su-Schrieffer-Heeger (SSH) lattice models with alternating gain and loss in real space and examine the dynamical evolution of the trace distance under pure dephasing. By tuning system parameters, this model supports phases with either parity-time or anti-parity-time symmetries, enabling us to explore the interplay between dephasing and different non-Hermitian symmetries. While the trace distance exhibits distinct dynamical behaviors across the different phases in the absence of dephasing, its response to dephasing is largely symmetry-independent but instead initial-state dependent. By varying initial states, we observe that increasing the dephasing strength can either merely accelerate the decay of the trace distance or stabilize it. Interestingly, we reveal two kinds of dephasing-induced stabilization that differ in the strong dephasing limit: a partial stabilization, where the trace distance approaches a finite value smaller than its initial value in the long-time limit, and a complete stabilization, where the trace distance remains at its initial value throughout the entire evolution. By analyzing the equation of motion, we attribute the initial-state dependent dephasing effect to the alternating gain and loss in the system and confirm its absence in Hermitian counterparts. Furthermore, in the anti-parity-time symmetry unbroken phase, we identify a continuous suppression-upon increasing the dephasing strength-of the otherwise exponential decay of the trace distance seen in the absence of dephasing.

quant-ph

Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers. To address this gap in UAV task settings, we introduce \textbf{FLIGHT}, a \textbf{F}ine-grained \textbf{L}ong-horizon \textbf{I}nstruction-\textbf{G}uided benchmark for \textbf{H}ybrid UAV navigation and reasoning \textbf{T}asks, which combines multi-stage instructions with dense 6-DoF trajectory annotations across two dataset splits: Fine-grained VLN and Long-horizon Flow. To endow the UAV agent with the capability of real-time in-flight reasoning over task execution status and mission planning, while simultaneously accommodating high-frequency, real-time precise control, we further propose \textbf{FLIGHT VLA}, an asynchronous architecture that decouples a low-frequency Streaming Pilot Vision-Language Model (VLM) for task-state reasoning from a high-frequency diffusion action model for continuous control, supervised by explicit \textbf{Pilot Reasoning} texts that summarize the current flight state and anticipate the next subgoal. In closed-loop evaluation, FLIGHT VLA consistently surpasses representative VLN and VLA baselines on our FLIGHT benchmarks, achieving stronger multi-stage completion, subgoal adherence, and terminal control. Its trained Streaming Pilot Reasoning VLM further improves UAV video reasoning, validating the effectiveness of our design.

cs.RO

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization

4-bit quantization reduces the memory footprint and latency of large language model inference, but its aggressive precision reduction can severely degrade accuracy. Prior methods address this by decomposing each weight matrix into two components (e.g., via singular value decomposition) and quantizing them separately, assigning the bulk of values to a low-precision residual component while handling outliers with a high-precision low-rank component. However, such decompositions are designed to minimize the real-valued energy of the residual, rather than the post-quantization error of the residual and low-rank components. We propose TwinQuant, a 4-bit quantization framework that learns quantization-friendly decomposed subspaces and jointly reshapes both the low-rank and residual components. TwinQuant learns component-specific transformations via a joint optimization over the Stiefel and general linear manifolds, flattening their distributions and reducing dynamic-range imbalance. To enable efficient end-to-end execution, we further design a fused dual-component kernel that pipelines the two-stage low-rank computation on-chip and merges both components with a single epilogue, avoiding intermediate global-memory traffic. Across LLaMA3 and Qwen3 models, TwinQuant preserves near-FP16 accuracy and delivers up to $1.8\times$ end-to-end speedup over an FP16 baseline.

cs.DC

RadioFormer3D: Weakly Supervised 3D Radio Map Estimation in Low-Altitude Airspace via Generative Modeling

With the emergence of wireless applications in three-dimensional environments, such as the low-altitude airspace and 3D heterogeneous networks, radio map estimation is increasingly required to characterize signal propagation across both horizontal and vertical dimensions. However, extending radio map estimation from 2D to 3D remains challenging due to increased spatial sparsity and limited supervision across continuous altitudes. In this paper, we propose \textbf{\textit{RadioFormer3D}}, a specialized model for volumetric spectrum reconstruction under weak supervision. Building on the dual-stream, multi-granularity fusion architecture of \textit{RadioFormer}, \textit{RadioFormer3D} introduces a Fourier-based sampling encoder and a volumetric decoder to efficiently process sparse measurements in 3D space. To alleviate the lack of vertical supervision, we propose the \textbf{\textit{Joint Spectrum Integrity Loss}}, which integrates volume-level pseudo-label supervision, map-level geometry-aware radio rendering, and pixel-level localized constraints within a unified optimization scheme. This design enables the model to capture complex vertical structural relationships more effectively under sparse supervision. Extensive experiments across several radio map datasets show that \textit{RadioFormer3D} achieves superior overall performance compared to representative existing methods. In particular, it demonstrates improved reconstruction quality at unlabeled altitudes while maintaining a favorable trade-off between accuracy and inference efficiency, positioning it as a highly promising solution for future 3D environment-aware wireless networks.

cs.CV

Manipulation of information flow and thermodynamic performance in nonreciprocal quantum dot information engines

Quantum information engines leverage information as a thermodynamic resource to facilitate energy conversion. In the operation of such engines, the information flow between the working substance and the controller is pivotal, however, strategies for its efficient manipulation remain largely unexplored. Here, we investigate an autonomous information engine based on a double-quantum-dot setup, where a downstream dot coupled to two reservoirs acts as the working substance, and an upstream dot coupled to a single reservoir serves as the controller. By extending the second law of thermodynamics to incorporate the effects of nonreciprocal couplings between the dots and their electronic reservoirs, we develop a thermodynamic framework that allows us to demonstrate that nonreciprocity can significantly modulate the inter-dot information flow, thereby providing a robust control mechanism. We show that the influence of nonreciprocity can be equivalently understood through a mapping to an effective reciprocal system upon a reparameterization of chemical potentials and the electron-electron coupling strength. We further analyze the impact of nonreciprocity on the engine's performance and operation regime. Our findings establish nonreciprocal coupling as an effective control knob for designing and optimizing quantum dot information engines, surpassing the capabilities of conventional reciprocal configurations.

cond-mat.stat-mech

GTF: Omnidirectional EPI Transformer for Light Field Super-Resolution

Light field (LF) image super-resolution benefits from Epipolar Plane Images (EPIs), whose line slopes explicitly encode disparity. However, existing Transformer-based LF SR methods mainly attend to horizontal and vertical EPIs, leaving diagonal epipolar geometry underexplored. We present GTF, an omnidirectional EPI Transformer that explicitly models horizontal, vertical, 45-degree, and 135-degree EPIs within a unified reconstruction framework. GTF combines directional EPI processing, MacPI-based prior injection, adaptive directional fusion, and a topology-preserving feed-forward network to better exploit LF geometry. For the NTIRE 2026 fidelity tracks, we use GTF as the main model, while a lightweight GTF-Tiny variant targets the efficiency track. On five standard LF SR benchmarks covering both real-captured and synthetic scenes, GTF reaches 32.78 dB without inference-time enhancement, and stronger inference settings with EPSW and test-time augmentation further improve performance. Under the NTIRE 2026 efficiency constraint, GTF-Tiny attains 32.57 dB with only 0.915M parameters and 19.81 GFLOPs. In the NTIRE 2026 Light Field Image Super-Resolution Challenge, our submissions rank 3rd on Track 1 and Track 3 and 4th on Track 2. Architecture-evolution, channel-width, and inference analyses further support the effectiveness of diagonal EPI modeling, directional fusion, and the lightweight design.

cs.CV

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD-FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The challenge received strong community interest, with 128 registered participants and a total of 696 submissions. Among them, 31 teams actively participated, and 19 teams submitted valid final results. Participants explored a wide range of strategies, introducing innovative methods that push the performance frontier under both open-source and closed-source tracks. This report presents a detailed overview of the NTIRE 2026 CD-FSOD Challenge, including a summary of the submitted approaches and an analysis of the final results across all participating teams. Challenge Codes: https://github.com/ohMargin/NTIRE2026_CDFSOD.

cs.CV

Dynamic Whole-Body Dancing with Humanoid Robots -- A Model-Based Control Approach

This paper presents an integrated model-based framework for generating and executing dynamic whole-body dance motions on humanoid robots. The framework operates in two stages: offline motion generation and online motion execution, both leveraging future state prediction to enable robust and dynamic dance motions in real-world environments. In the offline motion generation stage, human dance demonstrations are captured via a motion capture (MoCap) system, retargeted to the robot by solving a Quadratic Programming (QP) problem, and further refined using Trajectory Optimization (TO) to ensure dynamic feasibility. In the online motion execution stage, a centroidal dynamics-based Model Predictive Control (MPC) framework tracks the planned motions in real time and proactively adjusts swing foot placement to adapt to real world disturbances. We validate our framework on the full-size humanoid robot Kuavo 4Pro, demonstrating the dynamic dance motions both in simulation and in a four-minute live public performance with a team of four robots. Experimental results show that longer prediction horizons improve both motion expressiveness in planning and stability in execution.

cs.RO

Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance

Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on task-specific requirements and sample variations. To address these limitations, this paper presents an efficient RGB-D scene understanding model that performs a range of tasks, including semantic segmentation, instance segmentation, orientation estimation, panoptic segmentation, and scene classification. The proposed model incorporates an enhanced fusion encoder, which effectively leverages redundant information from both RGB and depth inputs. For semantic segmentation, we introduce normalized focus channel layers and a context feature interaction layer, designed to mitigate issues such as shallow feature misguidance and insufficient local-global feature representation. The instance segmentation task benefits from a non-bottleneck 1D structure, which achieves superior contour representation with fewer parameters. Additionally, we propose a multi-task adaptive loss function that dynamically adjusts the learning strategy for different tasks based on scene variations. Extensive experiments on the NYUv2, SUN RGB-D, and Cityscapes datasets demonstrate that our approach outperforms existing methods in both segmentation accuracy and processing speed.

cs.CV

Seeking Necessary and Sufficient Information from Multimodal Medical Data

Learning multimodal representations from medical images and other data sources can provide richer information for decision-making. While various multimodal models have been developed for this, they overlook learning features that are both necessary (must be present for the outcome to occur) and sufficient (enough to determine the outcome). We argue learning such features is crucial as they can improve model performance by capturing essential predictive information, and enhance model robustness to missing modalities as each modality can provide adequate predictive signals. Such features can be learned by leveraging the Probability of Necessity and Sufficiency (PNS) as a learning objective, an approach that has proven effective in unimodal settings. However, extending PNS to multimodal scenarios remains underexplored and is non-trivial as key conditions of PNS estimation are violated. We address this by decomposing multimodal representations into modality-invariant and modality-specific components, then deriving tractable PNS objectives for each. Experiments on synthetic and real-world medical datasets demonstrate our method's effectiveness. Code will be available on GitHub.

cs.CV

LIBERO-X: Robustness Litmus for Vision-Language-Action Models

Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often provide limited or misleading assessments due to insufficient evaluation protocols that inadequately capture real-world distribution shifts. This work systematically rethinks VLA benchmarking from both evaluation and data perspectives, introducing LIBERO-X, a benchmark featuring: 1) A hierarchical evaluation protocol with progressive difficulty levels targeting three core capabilities: spatial generalization, object recognition, and task instruction understanding. This design enables fine-grained analysis of performance degradation under increasing environmental and task complexity; 2) A high-diversity training dataset collected via human teleoperation, where each scene supports multiple fine-grained manipulation objectives to bridge the train-evaluation distribution gap. Experiments with representative VLA models reveal significant performance drops under cumulative perturbations, exposing persistent limitations in scene comprehension and instruction grounding. By integrating hierarchical evaluation with diverse training data, LIBERO-X offers a more reliable foundation for assessing and advancing VLA development.

cs.CV

Direct temperature readout in nonequilibrium quantum thermometry

Quantum thermometry aims to measure temperature in nanoscale quantum systems, paralleling classical thermometry. However, temperature is not a quantum observable, and most theoretical studies have therefore concentrated on analyzing fundamental precision limits set by the quantum Fisher information through the quantum Cramer-Rao bound. In contrast, whether a direct temperature readout can be achieved in quantum thermometry remains largely unexplored, particularly under the nonequilibrium conditions prevalent in real-world applications. To address this, we develop a direct temperature readout scheme based on a thermodynamic inference strategy. The scheme integrates two conceptual developments: (i) By applying the maximum entropy principle with the thermometer's mean energy as a constraint, we assign a reference temperature to the nonequilibrium thermometer. We demonstrate that this reference temperature outperforms a commonly used effective temperature defined through equilibrium analogy. (ii) We obtain positive semi-definite error functions that lower-bound the deviation of the reference temperature from the true temperature-in analogy to the quantum Cramer-Rao bound for the mean squared error-and vanish upon thermalization with the sample. Combining the reference temperature with these error functions, we introduce a notion of corrected dynamical temperature which furnishes a postprocessed temperature readout under nonequilibrium conditions. This corrected dynamical temperature can be evaluated adaptively without prior knowledge of the actual temperature. We validate the corrected dynamical temperature in a qubit-based thermometer under a range of nonequilibrium initial states, confirming its capability to estimate the true temperature. Importantly, we find that increasing quantum coherence can enhance the precision of this readout.

quant-ph

Quantum Interaction Between Free Electrons and Light Involving First-order and Second-order Process

Photon-induced Near-field Electron Microscopy (PINEM) effect has revealed the quantum interaction between free electrons and optical near filed, which demonstrated plenty of novel phenomena of manipulating free electron wave packet and detecting/shaping quantum photonic states. However, free electrons generally only absorb/emit one photon at a time, while the physical mechanism and phenomena of free electron-two-photon interaction have not been studied yet. Moreover, the relationship between PINEM and Kapitza-Dirac (KD) effect and nonlinear Compton scattering is still unclear. Here we develop the full quantum theory of electron-photon interaction considering the two-photon process. It is revealed that the emission/absorption of two photons by electrons can be greatly enhanced by manipulating the electric field component of optical near field, and the quantum interference between single-photon and two-photon processes can occur in some circumstances, which affects the photon number state, electron energy states and electron-photon entanglement. Meanwhile, it is found that the KD effect (elastic electron-photon scattering) and nonlinear Compton scattering (inelastic electron-photon scattering) are also a kind of two-photon process and the distribution of electrons can be deduced analytically based on the full quantum theory. Our work uncovers the possible abundant phenomena when free electron interacting with two photons, paves the way for more in-depth studies of nonlinear processes in electron-photon quantum interactions in the future.

quant-ph