SearcharxivSearch

arXiv subjects

Yichi Zhang

Publications and source records attributed to Yichi Zhang.

At least 19 recordsLinked to original sources

SoftRerank: Hierarchical Soft Fusion with Candidate-Label Reranking for Long-Tailed Micro-Action Recognition

Micro-actions are subtle, low-intensity non-verbal behaviors that provide cues to fine-grained human states, including emotions and intentions. Recognizing them remains difficult because they are brief, contain weak visual changes, and often exhibit similar motion patterns across categories. This paper addresses these challenges with a fine-grained micro-action recognition method that combines full fine-tuning of InternVideo2.5, hierarchical soft fusion, and a lightweight candidate-label reranker. For the long-tailed label distribution in MA-52, we use class-balanced sampling and inverse-frequency reweighting to reduce the effect of frequent classes during training. We fine-tune InternVideo2.5 end to end and attach coarse and group-conditional fine-grained classification heads to the shared video representation, improving the consistency between coarse and fine predictions. For ambiguous samples, the candidate-label reranker uses hard samples and video-label matching to focus on easily confused fine-grained actions. Experiments validate the proposed method, which achieves a 79.99% F1-mean on MA-52 and ranks first in the 3rd Micro-Action Analysis Grand Challenge at ACM Multimedia 2026.

cs.CV

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehensive whole-body PET/CT analysis. In this work, we introduce MetaStructAtlas, a large-scale dataset for grounded whole-body PET/CT interpretation that synthesizes multimodal imaging with integrated anatomical, metabolic, and semantic annotations. MetaStructAtlas provides 490 co-registered 3D PET and CT volumes with 50,470 organ-level segmentation masks and grounded radiology reports. To facilitate interactive reasoning, we further developed MetaStructVQA, a standardized 3D grounded visual question-answering benchmark containing 100,565 QA pairs. This framework explicitly links diagnostic queries to visual evidence across modalities, encompassing anatomical, morphological, and metabolic characteristics. Finally, we evaluate state-of-the-art 3D medical VLMs on MetaStructVQA, establishing a robust foundation for multimodal representation learning and integrated whole-body reasoning in nuclear medicine.

cs.CV

Lead-Lag Relationships in Financial Markets: A Comparison of Multiple Clustering Algorithms

Lead-lag relationships are widely used in financial time series, and many clustering algorithms based on them have been developed. The traditional DTW-KMedoids algorithm performs well both on the synthetic dataset and the real financial dataset. However, there are still several limitations to these algorithms: low efficiency caused by high time complexity, poor mathematical properties from DTW distance, the clustering effect is sensitive to the number of clusters. To solve the problems above and improve the performance, this paper introduces three clustering algorithms: MiniRocket-KMeans, KShape, Ensemble algorithm (a combination of KShape and DTW-KMedoids) and compares their performance on synthetic and real stock datasets with DTW-KMedoids algorithm under the same trade strategy. In addition, this paper also finds the best number of clusters by maximizing the silhouette coefficient in each clustering algorithm to improve the stability of the experiment results. Our main conclusions are as follows: MiniRocket-KMeans performs best under the lead strategy, achieving a Sharpe ratio of 0.866 with a maximum drawdown controlled at -63.9\%; the ensemble algorithm exhibits excellent stability; the robustness is significantly improved after finding the best number of clusters; the p-values of the hypothesis test on the Sharpe ratio of all strategies are 0.0, verifying the statistical validity of the lead-lag trading strategy. Finally, future improvement directions such as customized lead-lag matrices and optimized ensemble voting mechanisms are proposed.

q-fin.ST

UR$^{2}$-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

Radiologists generate diagnostic reports through iterative and selective revisiting of suspicious regions to refine their interpretations. Recent multimodal large language models (MLLMs) for radiology report generation (RRG) have shifted from text-only reasoning toward a ``Thinking-with-Images'' paradigm, incorporating visual evidence into the reasoning process. However, existing methods provide static visual evidence without a dynamic revisit mechanism during reasoning, neglecting how radiologists re-examine uncertain observations. To this end, we propose an Uncertainty-aware Revisit Reasoning MLLM (UR$^{2}$-MLLM) framework that dynamically revisits uncertain regions during reasoning for RRG. UR$^{2}$-MLLM is first equipped with uncertainty perception by training on an uncertainty-aware dataset. We then construct a multimodal reasoning trajectory dataset together with a detect-and-copy mechanism, which guides when and where to revisit. Finally, a visual grounding reward refines this behavior through reinforcement learning, aligning the revisited regions with corresponding anatomical structures. Experiments on MIMIC-CXR and IU-Xray show that UR$^{2}$-MLLM achieves state-of-the-art performance, highlighting the value of uncertainty-aware visual revisit reasoning for reliable and clinically aligned report generation.

cs.CV

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

cs.RO

Robust controlled-Z gate for Rydberg atoms based on level-crossing-free echoing rapid adiabatic passage

We propose a controlled-Z gate scheme for Rydberg atoms based on level-crossing-free echoing rapid adiabatic population transfer. We design antisymmetric Rabi frequency pulses and symmetric detuning pulses, enabling the system to completely avoid level-crossing points throughout the evolution, and the dynamical phase is naturally eliminated by the time-reversal symmetry of the double-pulse sequence. We incorporate dissipative effects through the Lindblad master equation. The numerical simulation yields a two-qubit CZ gate fidelity of 0.9999. When the Rabi-frequency fluctuation is within $\pm 2\%$, and the detuning offset is within $\pm 1\%$, the fidelity can still remain above 0.999. Under the same dissipative model, the three-qubit CCZ gate achieves a fidelity of 0.999. When a single-parameter fluctuation does not exceed $\pm 3\%$, the fidelity is always higher than 0.997. Our scheme requires no laser phase jumps or fast switching operations. The zero-area pulse structure suppresses first-order intensity noise, and the symmetric double-pulse sequence avoids spatially resolved laser switching, making it suitable for parallel gate operations in large-scale neutral-atom arrays.

quant-ph

Context Is Not Authority: Structured Runtime Governance for Financial Market Agents

Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present SAGE-Fin, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the object of runtime control. SAGE-Fin compiles proposals into typed, adapter-bound candidates; records missing or stale institutional obligations as coverage debt; contracts authority under current market, account, policy, and dialogue state; and requires an exact-artifact receipt whose nominal type matches the consuming response, execution, or policy adapter. Evidence and workflow progress cannot substitute for effect authority, and prior authorization is rechecked after state changes. Across an authored 616-case catalog, five deterministic specifications yield 3,080 outputs; a label-isolated harness obtains 616/616 binary reference-prototype parity, including 3/3 named response-gate fixtures, while 22 tests cover selected paths. These results establish executable conformance, not independent safety accuracy. Separately, SAGE-Fin's response gate processed real customer-facing production requests at a confidential digital-asset platform. An operational team independent of the implementation team reached a strongly positive post-deployment conclusion on practical usefulness and workflow fit, and end-user feedback was also strongly positive. Disclosure permits only the review's independence, stakeholder classes, assessed dimensions, and directional conclusion, so this is qualitative field corroboration rather than an aggregate effect estimate. Three distinct de-identified predecessor failures, with independently confirmed 0/3 interception, ground repeated-emission drift, stale account evidence, and missing escalation state without estimating prevalence or treatment effect.

cs.AI

A fully integrated dispersion-managed femtosecond mode-locked laser

Femtosecond lasers underpin applications ranging from material processing to corneal surgery, while their regular pulse trains form optical frequency combs that have revolutionized timekeeping, spectroscopy, and metrology. On-chip optical frequency combs, such as Kerr microcombs, have enabled high-repetition-rate applications in optical communications and microwave photonics. However, integrated chip-scale sources operating at low repetition rates (100 MHz to 1 GHz), crucial for high peak intensities, remain elusive, as existing devices typically operate well beyond 10 GHz. Here, we demonstrate a self-starting, photonic integrated mode-locked laser based on a dispersion-managed architecture that accesses this regime. The laser combines erbium-implanted silicon nitride gain waveguides, integrated chirped Bragg gratings, and a semiconductor saturable absorber mirror to generate optical pulses with repetition rates from 0.5 to 1.2 GHz, pulse durations as short as 300 fs, and mode-locking thresholds down to 27.3 mW. The output forms a passively stable optical frequency comb with a comb-line drift below 1% of the repetition rate, surpassing the stability of commercial fiber lasers by two orders of magnitude. Leveraging this ultra-low threshold, we achieve complete hybrid integration by co-packaging the laser with a telecom-grade 980-nm III-V pump diode chip inside a compact photonic module. The resulting electrical-in/optical-out module delivers turnkey, stable mode-locked pulses, providing a compact, low-power, and vibration-insensitive foundry-compatible platform for field-deployable optical metrology and precision sensing.

physics.optics

Low Mach number limit for the Navier--Stokes--Korteweg equations with a stationary force

In this paper, we investigate the low Mach number limit for the three-dimensional compressible Navier--Stokes--Korteweg equations in the whole space under a small stationary external force. We first construct a family of small stationary solutions uniformly with respect to the Mach number $\epsilon$ and prove that both the stationary density fluctuation and the compressible component of the stationary velocity are of order $\epsilon^2$. For ill-prepared non-stationary perturbations around these stationary solutions, we establish the existence and uniqueness of global strong solution by combining uniform high-order energy estimates with a low-frequency Besov estimate and a Kawashima-type compensating functional. The main difficulty is that Korteweg tensor not only changes the elliptic structure of the stationary problem, but also modifies the dispersive mechanism of the acoustic modes. In Korteweg-symmetric variables, the associated spectral projections are uniformly bounded zero-order Fourier multipliers, while the acoustic-capillary phase is wave-like at low frequencies and Schr\"odinger-like at high frequencies. Since the source terms generated by the stationary coefficients are generally not integrable in time, we decompose the Duhamel source according to its time-integrability and frequency behavior. Dyadic dispersive estimates, high-frequency damping estimates, and maximal regularity for the heat equation yield the global-in-time convergence rate $\epsilon^{\min\{1/r,\,1/2-1/p\}}$ in the mixed Besov norms $L^r(0,\infty;\dot B^s_{p,1})$. As a consequence, Besov embeddings also yield quantitative convergence in the mixed Lebesgue norms $L^r(0,\infty;L^p)$.

math.AP

Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behavior for the remainder of a session but are silently dropped during compaction. To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic trajectory, and long-horizon research. Current compactors retain only 17% of injected SCs on average, and most perform worse than running the same task without compaction. Retention varies sharply with compactor, prompt, context length, SC phrasing, and injection location, showing that the loss is systematic rather than tied to any single setting. We propose an SC-aware extractor that runs alongside the compactor as a plug-and-play module, achieving over 90% retention across all three scenarios without modifying the compactor or LLM. The COMPINT evaluation suite and accompanying implementation are available at https://github.com/ZhiqiEliWang/compaction-integrity.

cs.CL

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model's tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.

cs.CL

Decoder-Guided Lossy Contour Coding Via Anchor Refinement

Object contours serve as compact structural priors for many receiver-side vision tasks such as image super-resolution, edge-conditioned generation, and machine vision. When such tasks are deployed over a bandwidth-limited channel, the sender transmits the high-quality object contour as structural side information to guide reconstruction at the receiver, while-to save bandwidth-only a low-quality reference such as a downsampled image or base-layer reconstruction is delivered. As a result, the decoder can already extract a coarse contour from this reference at no transmission cost, creating an encoder-decoder asymmetry: the fine contour must be coded and sent, yet a free coarse version is available at the decoder. This asymmetry is ignored by existing contour codecs such as JBIG2 and chain coding, which are lossless, symmetric, and offer no rate-distortion control, leading to high bitrates. In this paper, we propose a coarse-to-fine contour coding framework that models a high-quality contour as a structured geometric refinement of the decoder-available coarse contour. The encoder extracts ordered anchors along the fine contour and performs adaptive anchor skipping under a distortion constraint. The decoder then reconstructs the contour by using the coarse prior to guide anchor connectivity. This formulation enables lossy contour compression with an explicit rate-distortion trade-off. Experiments show 54.5%-66.9% bitrate reduction over methods without decoder-side guidance, and up to 5 times savings over JBIG2, while preserving high geometric accuracy.

eess.IV

High-fidelity multiqubit gates with Rydberg atoms via level-crossing-free Rapid adiabatic passage

We propose a rapid adiabatic passage (RAP) scheme based on level-crossing-free pulses for deterministic generation of multiqubit entangled states in Rydberg atom systems. Unlike conventional RAP protocols that rely on level crossings, our approach uses an antisymmetric Rabi frequency and an even-symmetric detuning, enabling robust population transfer without passing through any level crossing. By exploiting the Rydberg blockade effect, the protocol prepares entangled states directly from an initial product state. Specifically, two sequential RAP pulses separated by a pi_g pulse generate two-qubit Bell states, three-qubit W states, four-qubit GHZ states, and six-qubit honeycomb W states. Numerical simulations show that the fidelities exceed 0.9997 for the Bell and three-qubit W states, reach 0.997 for the four-qubit GHZ state, and surpass 0.9995 for the six-qubit honeycomb W state. The scheme demonstrates excellent robustness against pulse parameter fluctuations, with fidelities remaining above 0.99 under +/-5% parameter variations. This work provides a simple, efficient, and robust method for entangled-state preparation in neutral-atom quantum information processing.

quant-ph

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptions, failing to diagnose the mastery of specific physical principles. Moreover, the high stochasticity of generative processes causes current prompt optimization methods to suffer from gradient hallucinations, where optimizers are misled by transient visual artifacts rather than systemic flaws. To address these challenges, we introduce OmniPhys, a rigorous benchmark of 1,551 samples grounded in a Physical Knowledge Graph. By aligning PhET simulations with standard curricula, OmniPhys operationalizes a knowledge-to-scenario pipeline that performs diagnostic stress tests via a dual-path verification protocol. We further propose OmniPrompt, an iterative framework that treats physical alignment as a discrete optimization problem. For each query, OmniPrompt aggregates K stochastic images into a per-query feedback buffer. Across training, it further merges feedback from batches of B queries before each meta-policy update, filtering seed and query-local noise. Evaluations across 12 representative text-to-image models reveal universal physical bottlenecks. Results demonstrate that OmniPrompt significantly enhances physical consistency across diverse backbones, proving the transferability and efficacy of our evolved meta-policies. The code and data are available at https://github.com/zjukg/OmniPhys

cs.CV

HijackKV: New Threat in Position-Independent KV Cache Reuse

Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates across inference requests because it requires exact token and position matches. To improve efficiency, recent system optimizations introduce position-independent KV reuse, allowing KV cache to be reused whenever identical text chunks appear, regardless of their position in the sequence. We show this design introduces a new threat, KV Cache Hijacking. Since KV caches are retrieved by token match but encode the context in which they were originally computed, the KV tied to a benign-looking token chunk may encode an attacker-controlled prefix. When later reused in a victim query, this contaminated KV silently hijacks the model's behavior, even if no attacker-controlled text appears in the input. We introduce HIJACKKV, the first attack framework that systematically exploits this vulnerability, demonstrating its severity and practicality. HIJACKKV optimizes an attacker-controlled prefix, so that the KV computed for a subsequent common benign text encodes the attacker's goal, while the text remains unchanged for future cache hits. HIJACKKV achieves an average 94% success rate in a single attempt, remains effective under realistic constraints including low hit rates (10%) and frequent recomputation (50%), persists over multi-turn interactions, and transfers across models in black-box settings. We further provide design insights for building secure KV reuse systems.

cs.CR

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, while proprietary products and APIs increasingly expose real-time translation capabilities to users. However, practical deployment remains challenging for open and reproducible systems, especially in long-form and multi-speaker conversations where partial ASR hypotheses are unstable, turn boundaries are ambiguous, and target speech must be generated with an appropriate speaker prompt. We present X-Translator, a low-cost modular cascaded S2ST system that combines streaming ASR, machine translation, and prompt-conditioned TTS through a session-level runtime controller. The system uses incremental segment commitment to convert unstable ASR streams into translation-ready units, and an online speaker prompt manager to bind source speech spans to speaker-specific voice prompts for synthesis. We evaluate translation, speech quality, and latency with OpenSTBench, compare against proprietary speech translation APIs as behavioral baselines, measure long-form voice stability, evaluate speaker preservation in multi-speaker conversations, and assess multilingual translation quality. X-Translator provides an open platform for understanding the practical trade-offs of deployment-oriented S2ST. Code and demo are available at https://github.com/zhaoyx239/X-Translator.

eess.AS

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we build an auditable framework that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change. Across 1,080 frozen games spanning belief-disabled, active-belief, kernel-ablation, camp-restricted, consumption-policy, and high-load arms, and including a seed-paired A0/A1 comparison, the active-belief condition is associated with substantially better good-side outcomes: in the 200-seed A0/A1 comparison the good-side win rate rises from 0.205 to 0.390 (paired McNemar $\chi^2 = 16.4$, $p < 0.001$), with fewer irreversible witch-poison errors. We do not, however, attribute this shift to belief content. Direct action-belief consistency is low ($\approx 0.21$), and giving belief only to the werewolves helps the good side more than giving it only to the good side, which argues against a simple holder-benefit account; we therefore report the effect as an association and treat its mechanism as unresolved. The contribution is the audit framework itself: it makes the effect measurable, exposes low direct action-belief consistency, rejects an unreliable forced-consumption intervention with evidence, and separates strategy effects from load confounds. We accordingly position external belief in high-noise hidden-information games primarily as an auditable cognitive baseline that also carries decision-relevant signal, turning opaque agent behavior into replayable evidence for safer, controlled iteration.

cs.MA

Design of an Electrically Tunable Microtoroid for Frequency Selection of Polarization-Entangled Photons

Encoding quantum information into discrete optical frequencies, or "frequency bins," uses different colors of light as additional information channels, allowing each photon to carry more information than polarization alone. We present a computational design for an electrically tunable silica microtoroid that selects desired frequency channels after a polarization-entangled photon pair has been generated without disturbing the photons' polarization entanglement. In the proposed architecture, the 750 nm signal photon passes through the microtoroid, while its entangled 880 nm partner bypasses the resonator and serves as a reference for the selected frequency channel. The principal challenge is resonator birefringence: because horizontally and vertically polarized light resonate at slightly different frequencies, the selected frequency can reveal the photon's polarization state and weaken the quantum correlation between the photon pair. We solve this problem by adding a small lithium-niobate tuning element controlled with a single applied voltage. The voltage shifts the resonator so that it responds almost identically to horizontally and vertically polarized light, reducing the remaining mismatch to only 0.286 optical linewidths across nine frequency channels. The photons remain strongly entangled after passing through the device, with a concurrence of C = 0.969, a Bell-state fidelity of F = 0.981, and a Bell parameter of S_max = 2.785. If the relative timing between the frequency channels is also controlled, the same device can generate a nine-channel polarization-frequency hyperentangled state with an effective dimension of K = 8.97. This computational design provides a compact, electrically tunable bridge between polarization-entangled photon sources and future high-capacity quantum photonic systems.

quant-ph