SearcharxivSearch

arXiv subjects

Ke Tian

Publications and source records attributed to Ke Tian.

13 recordsLinked to original sources

QA-ReID: Quality-Aware Query-Adaptive Convolution Leveraging Fused Global and Structural Cues for Clothes-Changing ReID

Unlike conventional person re-identification (ReID), clothes-changing ReID (CC-ReID) presents severe challenges due to substantial appearance variations introduced by clothing changes. In this work, we propose the Quality-Aware Dual-Branch Matching (QA-ReID), which jointly leverages RGB-based features and parsing-based representations to model both global appearance and clothing-invariant structural cues. These heterogeneous features are adaptively fused through a multi-modal attention module. At the matching stage, we further design the Quality-Aware Query Adaptive Convolution (QAConv-QA), which incorporates pixel-level importance weighting and bidirectional consistency constraints to enhance robustness against clothing variations. Extensive experiments demonstrate that QA-ReID achieves state-of-the-art performance on multiple benchmarks, including PRCC, LTCC, and VC-Clothes, and significantly outperforms existing approaches under cross-clothing scenarios.

cs.CV

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured reasoning beyond surface-level alignment. We introduce CrossCheck-Bench, a diagnostic benchmark for evaluating contradiction detection in multimodal inputs. The benchmark adopts a hierarchical task framework covering three levels of reasoning complexity and defines seven atomic capabilities essential for resolving cross-modal inconsistencies. CrossCheck-Bench includes 15k question-answer pairs sourced from real-world artifacts with synthetically injected contradictions. The dataset is constructed through a multi-stage annotation pipeline involving more than 450 expert hours to ensure semantic validity and calibrated difficulty across perception, integration, and reasoning. We evaluate 13 state-of-the-art vision-language models and observe a consistent performance drop as tasks shift from perceptual matching to logical contradiction detection. Most models perform well on isolated entity recognition but fail when multiple clues must be synthesized for conflict reasoning. Capability-level analysis further reveals uneven skill acquisition, especially in tasks requiring multi-step inference or rule-based validation. Additional probing shows that conventional prompting strategies such as Chain-of-Thought and Set-of-Mark yield only marginal gains. By contrast, methods that interleave symbolic reasoning with grounded visual processing achieve more stable improvements. These results highlight a persistent bottleneck in multimodal reasoning and suggest new directions for building models capable of robust cross-modal verification.

cs.CL

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in image understanding, but long-video are constrained by context windows and computational cost. Uniform frame sampling often leads to substantial information loss. Meanwhile existing keyframe selection methods such as text-frame retrieval or RL-based frame optimization typically yield sparse and temporally disjointed frames, overlooking scene continuity and lacking flexibility for multi-scale frame selection. To address these limitations, we introduce K-frames, a novel paradigm for scene-driven keyframe selection that preserves temporal continuity. Instead of selecting individual frames, K-frames predicts semantically coherent, query-relevant clips, which enables any-k keyframes selection to meet diverse user budgets. To achieve this approach, we first introduce PeakClips, a dataset of 200K video highlights conditioned by query. Building on this dataset, K-frames learns clip2frame selection using a three-stage progressive curriculum. It involves two Supervised Fine-Tuning stages for temporal grounding and key-clip perception, followed by a Reinforcement Learning stage that directly optimizes the scene-driven prediction policy for downstream task without further annotations. Extensive experiments on major long-video understanding benchmarks demonstrate that K-frames provides an effective, interpretable, and plug-and-play solution for keyframe selection at various scales. Our dataset and model will be available.

cs.LG

X-ray microcomputed tomography of 3D chaotic microcavities

Chaotic microcavities play a crucial role in several research areas, including the study of unidirectional microlasers, nonlinear optics, sensing, quantum chaos, and non-Hermitian physics. To date, most theoretical and experimental explorations have focused on two-dimensional (2D) chaotic dielectric microcavities, while there have been minimal studies on three-dimensional (3D) ones since precise geometrical information of a 3D microcavity can be difficult to obtain. Here, we image 3D microcavities with submicron resolution using X-ray microcomputed tomography (micro CT), enabling nondestructive imaging that preserves the sample for subsequent use. By analyzing the ray dynamics of a typical deformed microsphere, we demonstrate that a sufficient deformation along all three dimensions can lead to chaotic ray trajectories over extended time scales. Notably, using the X-ray micro CT reconstruction results, the phase space chaotic ray dynamics of a deformed microsphere are accurately established. X-ray micro CT could become a unique platform for the characterization of such deformed 3D microcavities by providing a precise means for determining the degree of deformation necessary for potential applications in ray chaos and quantum chaos.

physics.optics

Evaporation characteristics of Er$^{3+}$ doped silica fiber and its application in the preparation of whispering gallery mode lasers

The fabrication of whispering gallery lasers (WGL) is used to experimentally evaluate the evaporation rate (mol/$μ$m) and ratio (mol/mol) of erbium and silica lost from a doped fiber during heating. Fixed lengths of doped silica fiber are spliced to different lengths of undoped fiber and then evaporated by feeding into the focus of a CO$_{2}$ laser. During evaporation, erbium ions are precipitated in the doped silica fiber to control the erbium concentration in the remaining SiO$_2$, which is melted into a microsphere. By increasing the length of the undoped section, a critical point is reached where effectively no ions remain in the glass microsphere. The critical point is found using the lasing spectra of the whispering gallery modes in microspheres with equal sizes. From the critical point, it is estimated that, for a given CO$_{2}$ laser power, $6.36 \times 10^{-21}$~mol of Er$^{3+}$ is lost during the evaporation process for every cubic micron of silica fiber. This is equivalent to $1.74 \times 10^{-7}$~mol of Er$^{3+}$ lost per mol of SiO$_{2}$ evaporated. This result facilitates the control of the doping concentration in WGLs and provides insight into the kinetics of laser-induced evaporation of doped silica.

physics.optics

Structural characterization of thin-walled microbubble cavities

Whispering gallery mode (WGM) microbubble cavities are a versatile optofluidic sensing platform owing to their hollow core geometry. To increase the light-matter interaction and, thereby, achieve a higher sensitivity, thin-walled microbubbles are desirable. However, a lack of knowledge about the precise geometry of hollow microbubbles prevents us from having an accurate theoretical model to describe the WGMs and their response to external stimuli. In this work, we provide a complete characterization of the wall structure of a microbubble and propose a theoretical model for the WGMs in this thin-walled microcavity based on the optical waveguide approach. Structural characterization of the wavelength-scale wall is enabled by focused ion beam milling and scanning electron microscopy imaging. The proposed theoretical model is verified by finite element method simulations. Our approach can readily be extended to other low-dimensional micro-/nanophotonic structures.

physics.optics

Blue-band frequency comb and photodarkening in silica whispering gallery microresonators

To date there are extensive studies of optical nonlinearities in whispering gallery resonators (WGRs) in the near and mid-infrared wavelengths. Pushing this research into the visible region is equally valuable. Here, we demonstrate a Kerr frequency comb and Raman lasing at 462 nm in an SiO2 WGR. Notably, due to the high optical intensities achieved, photodarkening is unavoidable and can quickly degrade the optical quality of both the coupling optical nanofiber and the microcavity even at very low pump powers. Nonetheless, stable stimulated Raman scattering (SRS) and hyper-parametric oscillation in normally dispersed WGRs is demonstrated in the presence of photodarkening by taking advantage of in-situ thermal bleaching. These observations highlight the challenges of silica-based, short wavelength nonlinear optics in high quality, small mode volume devices. We propose a method to overcome this apparent limitation and demonstrate blue-band nonlinear optical processes in silica WGRs, thus providing a baseline for optics research in the blue region for any optical devices fabricated from SiO2.

physics.optics

Living-Off-The-Land Command Detection Using Active Learning

In recent years, enterprises have been targeted by advanced adversaries who leverage creative ways to infiltrate their systems and move laterally to gain access to critical data. One increasingly common evasive method is to hide the malicious activity behind a benign program by using tools that are already installed on user computers. These programs are usually part of the operating system distribution or another user-installed binary, therefore this type of attack is called "Living-Off-The-Land". Detecting these attacks is challenging, as adversaries may not create malicious files on the victim computers and anti-virus scans fail to detect them. We propose the design of an Active Learning framework called LOLAL for detecting Living-Off-the-Land attacks that iteratively selects a set of uncertain and anomalous samples for labeling by a human analyst. LOLAL is specifically designed to work well when a limited number of labeled samples are available for training machine learning models to detect attacks. We investigate methods to represent command-line text using word-embedding techniques, and design ensemble boosting classifiers to distinguish malicious and benign samples based on the embedding representation. We leverage a large, anonymized dataset collected by an endpoint security product and demonstrate that our ensemble classifiers achieve an average F1 score of 0.96 at classifying different attack classes. We show that our active learning method consistently improves the classifier performance, as more training data is labeled, and converges in less than 30 iterations when starting with a small number of labeled instances.

cs.CR

A packaged whispering gallery resonator device based on an optical nanoantenna coupler

In this work, we present the design and fabrication of a packaged whispering gallery mode (WGM) device based on an optical nanoantenna as the coupler and a glass microsphere as the resonator. The microspheres were fabricated from SiO$_2$ fiber or Er$^{3+}$-doped fiber, the latter creating a WGM laser with a threshold of 93 $μ$W at 1531 nm. The coupler-resonator WGM device is packaged in a glass capillary. The performance of the packaged microlaser is characterized, with lasing emission both excited in and collected from the WGM cavity via the nanoantenna. The packaged system provides isolation from environmental contamination, a small size, and unidirectional coupling while maintaining a high quality (Q-) factor ($\sim$10$^8$). It opens up new possibilities for practical applications of WGM microdevices in a variety of fields such as low threshold lasers, filters, and sensors.

physics.optics

Unsupervised heart abnormality detection based on phonocardiogram analysis with Beta Variational Auto-Encoders

Heart Sound (also known as phonocardiogram (PCG)) analysis is a popular way that detects cardiovascular diseases (CVDs). Most PCG analysis uses supervised way, which demands both normal and abnormal samples. This paper proposes a method of unsupervised PCG analysis that uses beta variational auto-encoder ($β-\text{VAE}$) to model the normal PCG signals. The best performed model reaches an AUC (Area Under Curve) value of 0.91 in ROC (Receiver Operating Characteristic) test for PCG signals collected from the same source. Unlike majority of $β-\text{VAE}$s that are used as generative models, the best-performed $β-\text{VAE}$ has a $β$ value smaller than 1. Further experiments then find that the introduction of a light weighted KL divergence between distribution of latent space and normal distribution improves the performance of anomaly PCG detection based on anomaly scores resulted by reconstruction loss. The fact suggests that anomaly score based on reconstruction loss may be better than anomaly scores based on latent vectors of samples

cs.SD

CryptoGuard: High Precision Detection of Cryptographic Vulnerabilities in Massive-sized Java Projects

Cryptographic API misuses, such as exposed secrets, predictable random numbers, and vulnerable certificate verification, seriously threaten software security. The vision of automatically screening cryptographic API calls in massive-sized (e.g., millions of LoC) Java programs is not new. However, hindered by the practical difficulty of reducing false positives without compromising analysis quality, this goal has not been accomplished. State-of-the-art crypto API screening solutions are not designed to operate on a large scale. Our technical innovation is a set of fast and highly accurate slicing algorithms. Our algorithms refine program slices by identifying language-specific irrelevant elements. The refinements reduce false alerts by 76% to 80% in our experiments. Running our tool, CrytoGuard, on 46 high-impact large-scale Apache projects and 6,181 Android apps generate many security insights. Our findings helped multiple popular Apache projects to harden their code, including Spark, Ranger, and Ofbiz. We also have made substantial progress towards the science of analysis in this space, including: i) manually analyzing 1,295 Apache alerts and confirming 1,277 true positives (98.61% precision), ii) creating a benchmark with 38-unit basic cases and 74-unit advanced cases, iii) performing an in-depth comparison with leading solutions including CrySL, SpotBugs, and Coverity. We are in the process of integrating CryptoGuard with the Software Assurance Marketplace (SWAMP).

cs.CR

Checking is Believing: Event-Aware Program Anomaly Detection in Cyber-Physical Systems

Securing cyber-physical systems (CPS) against malicious attacks is of paramount importance because these attacks may cause irreparable damages to physical systems. Recent studies have revealed that control programs running on CPS devices suffer from both control-oriented attacks (e.g., code-injection or code-reuse attacks) and data-oriented attacks (e.g., non-control data attacks). Unfortunately, existing detection mechanisms are insufficient to detect runtime data-oriented exploits, due to the lack of runtime execution semantics checking. In this work, we propose Orpheus, a new security methodology for defending against data-oriented attacks by enforcing cyber-physical execution semantics. We first present a general method for reasoning cyber-physical execution semantics of a control program (i.e., causal dependencies between the physical context and program control flows), including the event identification and dependence analysis. As an instantiation of Orpheus, we then present a new program behavior model, i.e., the event-aware finite-state automaton (eFSA). eFSA takes advantage of the event-driven nature of CPS control programs and incorporates event checking in anomaly detection. It detects data-oriented exploits if a specific physical event is missing along with the corresponding event dependent state transition. We evaluate our prototype's performance by conducting case studies under data-oriented attacks. Results show that eFSA can successfully detect different runtime attacks. Our prototype on Raspberry Pi incurs a low overhead, taking 0.0001s for each state transition integrity checking, and 0.063s~0.211s for the cyber-physical contextual consistency checking.

cs.CR

Breaking the Target: An Analysis of Target Data Breach and Lessons Learned

This paper investigates and examines the events leading up to the second most devastating data breach in history: the attack on the Target Corporation. It includes a thorough step-by-step analysis of this attack and a comprehensive anatomy of the malware named BlackPOS. Also, this paper provides insight into the legal aspect of cybercrimes, along with a prosecution and sentence example of the well-known TJX case. Furthermore, we point out an urgent need for improving security mechanisms in existing systems of merchants and propose three security guidelines and defenses. Credit card security is discussed at the end of the paper with several best practices given to customers to hide their card information in purchase transactions.

cs.CR