SearcharxivSearch

arXiv subjects

Changqing Li

Publications and source records attributed to Changqing Li.

15 recordsLinked to original sources

Evaluating LLM Coding Agents on SZ-Family Lossy Compression Across Architectures

Large language model (LLM) coding agents are increasingly applied to code translation and optimization, yet their effectiveness in performance-critical high-performance computing (HPC) settings remains poorly characterized. This paper evaluates LLM-based coding workflows on SZ-family error-bounded lossy compression kernels, which combine numerical constraints with memory-intensive and control-flow-heavy implementations. We study two representative CUDA workloads (SZp and SZx) and target two heterogeneous execution platforms: NVIDIA GPUs and Cerebras wafer-scale accelerators. Focusing on single-agent iterative generation, we analyze not only final throughput but also agent runtime behavior, including iteration patterns, sensitivity to prompt specification, and characteristic failure modes. Our results reveal a pronounced cross-architecture divergence. On GPUs, stronger models can achieve substantially higher throughput but exhibit increased sensitivity to prompt precision and optimization guidance, whereas on Cerebras the dominant challenge lies in producing runnable programs under a PE-centric spatial execution model. We further observe that LLM agents are more effective on modular kernels (SZx) than on tightly coupled bit-level pipelines (SZp), where structural dependencies hinder optimization progress. These findings suggest that evaluating LLM coding agents for HPC requires accounting for both performance outcomes and architecture-specific robustness, and that success on thread-based platforms does not directly transfer to spatial accelerators.

cs.DC

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex inter-relationships between images and scattered critical information across image sets. Inspired by human cognitive processes, we propose a Cognition-Inspired Meta-Action Framework (CINEMA), which decomposes multi-image reasoning into five structured meta-actions: Global, Focus, Hint, Think, and Answer, explicitly modeling the sequential cognitive steps humans naturally employ. For cold-start training, we introduce a Retrieval-Based Tree Sampling strategy that generates high-quality meta-action trajectories to bootstrap the model with reasoning patterns. During reinforcement learning, we adopt a two-stage paradigm: an exploration phase with Diversity-Preserving Strategy to avoid entropy collapse, followed by an annealed exploitation phase with DAPO to gradually strengthen exploitation. To train our model, we construct a dataset of 56k cold-start and 58k reinforcement learning instances spanning multi-image, multi-frame, and single-image tasks. We conduct extensive evaluations on multi-image reasoning benchmarks, video understanding benchmarks, and single-image benchmarks, achieving competitive state-of-the-art performance on several key benchmarks. Our model surpasses GPT-4o on the MUIR and MVMath benchmarks and notably outperforms specialized video reasoning models on video understanding benchmarks, demonstrating the effectiveness and generalizability of our human cognition-inspired reasoning framework.

cs.CV

Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition

Recent advances in multimodal large language models (MLLMs) have been primarily evaluated on general-purpose benchmarks, while their applications in domain-specific scenarios, such as intelligent product moderation, remain underexplored. To address this gap, we introduce an open-world logo recognition benchmark, a core challenge in product moderation. Unlike traditional logo recognition methods that rely on memorizing representations of tens of thousands of brands-an impractical approach in real-world settings-our proposed method, Logo-VGR, enables generalization to large-scale brand recognition with supervision from only a small subset of brands. Specifically, we reformulate logo recognition as a comparison-based task, requiring the model to match product images with candidate logos rather than directly generating brand labels. We further observe that existing models tend to overfit by memorizing brand distributions instead of learning robust multimodal reasoning, which results in poor performance on unseen brands. To overcome this limitation, Logo-VGR introduces a new paradigm of domain-specific multimodal reasoning: Logo Perception Grounding injects domain knowledge, and Logo-Guided Visual Grounded Reasoning enhances the model's reasoning capability. Experimental results show that Logo-VGR outperforms strong baselines by nearly 10 points in OOD settings, demonstrating superior generalization.

cs.CV

MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair

Large Language Models (LLMs) face persistent and evolving trustworthiness issues, motivating developers to seek automated and flexible repair methods that enable convenient deployment across diverse scenarios. Existing repair methods like supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) are costly and slow, while prompt engineering lacks robustness and scalability. Representation engineering, which steers model behavior by injecting targeted concept vectors during inference, offers a lightweight, training-free alternative. However, current approaches depend on manually crafted samples and fixed steering strategies, limiting automation and adaptability. To overcome these challenges, we propose MASteer, the first end-to-end framework for trustworthiness repair in LLMs based on representation engineering. MASteer integrates two core components: AutoTester, a multi-agent system that generates diverse, high-quality steer samples tailored to developer needs; and AutoRepairer, which constructs adaptive steering strategies with anchor vectors for automated, context-aware strategy selection during inference. Experiments on standard and customized trustworthiness tasks show MASteer consistently outperforms baselines, improving metrics by 15.36% on LLaMA-3.1-8B-Chat and 4.21% on Qwen-3-8B-Chat, while maintaining general model capabilities. MASteer demonstrates strong robustness, generalization, and practical value for scalable, efficient trustworthiness repair.

cs.AI

Timing-injection locking in a self-starting Mamyshev oscillator induced by the dissipative Faraday instability

Mamyshev oscillators (MOs), a novel class of passively mode-locked fiber lasers, serve as an excellent platform to explore complex nonlinear dynamics, ranging from localized structures to chaos. Despite their versatility, achieving self-starting mode-locking remains a significant challenge. In this study, we unveil the critical role of the dissipative Faraday instability (DFI) in facilitating the self-starting process of MOs, where the DFI triggers the symmetry breaking of the homogeneous solution to overcome the initiation barriers. A panoramic view of several distinct operational regimes with distinct DFI patterns is provided, namely the non-self-starting states, the irregular patterns, the harmonic mode locking regime, the stable single pulse and the stable multi pulse regime. For the lattest case, we uncover the origins of randomness in these pulse sequences through analyzing the causality between the timing of the random pulses and the initial seeding conditions. Building upon these findings, we propose the novel time-injection locking technique to customize the temporal locations of the pulses as well as the pattern timing in MOs, thus demonstrating its potential for applications in all-optical data storage and tunable ultrashort pulse sources.

physics.optics

Inference Performance Optimization for Large Language Models on CPUs

Large language models (LLMs) have shown exceptional performance and vast potential across diverse tasks. However, the deployment of LLMs with high performance in low-resource environments has garnered significant attention in the industry. When GPU hardware resources are limited, we can explore alternative options on CPUs. To mitigate the financial burden and alleviate constraints imposed by hardware resources, optimizing inference performance is necessary. In this paper, we introduce an easily deployable inference performance optimization solution aimed at accelerating LLMs on CPUs. In this solution, we implement an effective way to reduce the KV cache size while ensuring precision. We propose a distributed inference optimization approach and implement it based on oneAPI Collective Communications Library. Furthermore, we propose optimization approaches for LLMs on CPU, and conduct tailored optimizations for the most commonly used models. The code is open-sourced at https://github.com/intel/xFasterTransformer.

cs.AI

Development of Focused X-ray Luminescence Compute Tomography Imaging

X-ray luminescence is produced when contrast agents absorb energy from X-ray photons and release a portion of that energy by emitting photons in the visible and near-infrared range. X-ray luminescence computed tomography (XLCT) was introduced in the past decade as a hybrid molecular imaging modality combining the merits of both X-ray imaging (high spatial resolution) and optical imaging (high sensitivity to tracer nanophosphors).

eess.IV

Distributed Inference Performance Optimization for LLMs on CPUs

Large language models (LLMs) hold tremendous potential for addressing numerous real-world challenges, yet they typically demand significant computational resources and memory. Deploying LLMs onto a resource-limited hardware device with restricted memory capacity presents considerable challenges. Distributed computing emerges as a prevalent strategy to mitigate single-node memory constraints and expedite LLM inference performance. To reduce the hardware limitation burden, we proposed an efficient distributed inference optimization solution for LLMs on CPUs. We conduct experiments with the proposed solution on 5th Gen Intel Xeon Scalable Processors, and the result shows the time per output token for the LLM with 72B parameter is 140 ms/token, much faster than the average human reading speed about 200ms per token.

cs.DC

X-ray projection imaging of metal oxide particles inside gingival tissues

There is increasing recognition that oral health affects overall health and systemic diseases. Nonetheless it remains challenging to rapidly screen patient biopsies for signs of inflammation or the pathogens or foreign materials that elicit the immune response. This is especially true in conditions such as foreign body gingivitis (FBG), where the foreign particles are often difficult to detect. Our long term goal is to establish a method to determine if the inflammation of the gingival tissue is due to the presence of a metal oxide, with emphasis on elements that were previously reported in FBG biopsies, such as silicon dioxide, silica, and titanium dioxide whose persistent presence can be carcinogenic. In this paper, we proposed to use multiple energy X-ray projection imaging to detect and to differentiate different metal oxide particles embedded inside gingival tissues. To simulate the performance of the imaging system, we have used GATE simulation software to mimic the proposed system and to obtain images with different systematic parameters. The simulated parameters include the X-ray tube anode metal, the X-ray spectra bandwidth, the X-ray focal spot size, the X-ray photon number, and the X-ray dector pixel. We have also applied the de-noising algorithm to obtain better Contrast-to-noise ratio (CNR). Our results indicate that it is feasible to detect metal particles as small as 0.5 micrometer in diameter when we use a Chromium anode target with an energy bandwidth of 5 keV, an X-ray photon number of 10^8, and an X-ray detector with a pixel size of 0.5 micrometer and 100 by 100 pixels. We have also found that different metal particles could be differentiated from the CNR at four different X-ray anodes and spectra. These encouraging initial results will guide our future imaging system design.

physics.med-ph

Correlation between X-ray tube current exposure time and X-ray photon number in GATE

The image quality of X-ray imaging relies heavily on the X-ray output number which is dependent on the X-ray tube current and the exposure time. Hybrid X-ray imaging modalities like X-ray luminescence CT (XLCT) and X-ray fluorescence CT (XFCT) rely on the intensity of the X-ray tube to provide an accurate image reconstruction of the nanoprobe distribution in the imaging sample. A limiting factor of good image quality is the radiation dose that will be delivered to the imaging object. To accurately estimate the absorbed dose in an imaging protocol, it is better to simulate the X-ray imaging with a Monte Carlo platform such as GATE (Geant4 Application for Tomographic Emission). However, the input of GATE is a photon number of the simulated X-ray tube. So far, there is no good way to setup the photon number for a desired X-ray tube current. In this work, the accumulated radiation dose of a micro-CT X-ray tube at different current exposure times was recorded with a general-purpose ion chamber. GATE was used to model the total absorbed dose (cGy) in the sensitive volume of the ion chamber with different X-ray output numbers. A linear regression model was generated between the X-ray photon number in the GATE simulations and the tube current exposure time (mAs). The findings of this work provide an approach to correlate the X-ray tube current exposure time (mAs) to the X-ray photon number in GATE simulation of the X-ray tube.

physics.med-ph

A feasibility Study of Time of Flight Computed Tomography for Breast Imaging

Cone beam computed tomography (CBCT) for breast imaging has potential to replace conventional mammograms. However, concerns over dose and image quality prevent CBBCT systems from the clinical trial phase to next stage. The time of flight (TOF) method was recently shown to reduce the x-ray scattering effects by 95% and improve the image CNR by 110% for large volume objects. The advancements in x-ray sources like in compact Free Electron Lasers (FEL) and advancements in detector technology show potential for the TOF method to be feasible in CBCT when imaging large objects. In this study, we investigate the efficacy of this TOF CBCT in improving the breast cancer imaging. The GATE software was used to simulate the cone beam CT imaging of an 8 cm diameter cylindrical water phantom using a modeled 20 keV quasi-energetic FEL source and various detector temporal resolutions ranging from 1 to 1000 ps. An inhomogeneous breast phantom of similar size was also imaged using the same system setup. Results show that a detector temporal resolution of 10 ps improved the image contrast-to-noise ratio (CNR) by 57% and reduced the scatter-to-primary ratio (SPR) by 8.63 for a small cylindrical phantom. For the breast phantom, the image CNR was enhanced by 12% and the SPR was reduced by 1.35 at 5 ps temporal resolution.

physics.med-ph

X-ray fluorescence computed tomography (XFCT) imaging with a superfine pencil beam x-ray source

X-ray fluorescence computed tomography (XFCT) is a molecular imaging technique of x-ray photons, which can be used to sense different elements or nanoparticle (NP) agents inside deep samples or tissues. XFCT has been an active research topic for many years. However, XFCT has not been a popular molecular imaging tool because it has limited molecular sensitivity and spatial resolution. To further investigate XFCT imaging, we present a benchtop XFCT imaging system, in in which a unique pencil beam x-ray source and a ring of x-ray spectrometers were simulated using GATE (Geant4 Application for Tomographic Emission) software. An accelerated majorization minimization (MM) algorithm with an L1 regularization scheme was used to reconstruct the XRF image of Molybdenum (Mo) NP targets from the numerical measurements of GATE simulations. With a low x-ray source output rate, good target localization was achieved with a DICE coefficient of 83.681%. The reconstructed signal intensity of the targets was found to be relatively proportional to the target concentrations if detector number and placement is optimized. The MM algorithm performance was compared with maximum likelihood expectation maximization (ML-EM) and filtered back projection (FBP) algorithms. In the future, the benchtop XFCT imaging system will be tested experimentally

physics.med-ph

Radiation dose estimation in pencil beam x-ray luminescence computed tomography imaging

Pencil x-ray beam imaging provides superior spatial resolution than other imaging geometries like sheet beam and cone beam geometries due to the illumination of a line instead of an area or volume. However, the pencil beam geometry suffers from long scan times and concerns over dose discourage laboratory use of pencil beam x-ray sources. Molecular imaging techniques like XLCT imaging benefit most from pencil beam imaging to accurately localize the distribution of contrast agents embedded in a small animal object. To investigate the dose deposited by pencil beam x-ray imaging in XLCT, dose estimations from one angular projection scan by three different x-ray source energies were performed on a small animal object composed of water, bone, and blood with a Monte Carlo simulation platform, GATE (Geant4 Application for Tomographic Emission). Our results indicate that, with an adequate x-ray benchtop source with high brilliance and quasi-monochromatic properties like the Sigray source, the dose concerns can be reduced. With the Sigray source, the bone marrow was estimated to have a radiation dose of 30 mGy for a typical XLCT imaging, in which we have 6 angular projections, 100 micrometer scan step size, and 10^6 x-ray photons per linear scan.

physics.med-ph

Development of a focused-X-ray luminescence tomography (FXLT) system

Biophotonics is an active research area in molecular imaging, genetic diagnosis and prognosis, with direct applicability in precision medicine. However, long-standing challenges of biophotonics are well known due to low signal-to-noise ratio and poor image quality, mainly due to strong optical scattering especially in deep tissues. Recently, X-ray luminescence computed tomography (XLCT) has emerged as a hybrid molecular imaging modality and shown great promises in overcoming the optical scattering in deep tissues. However, its high spatial resolution capacity has not been fully implemented yet. In this paper, with a superfine focused X-ray beam we design a focused-X-ray luminescence tomography (FXLT) system for spatial resolution better than 150 micrometers and molecular sensitivity of 2 micromolar. First, we describe our system design. Then, we demonstrate that the molecular sensitivity of FXLT is about 5 micromolar considering the emitted visible photons from background. Furthermore, we analyze the measurement time per scan from measured photons numbers with a fiber bundle-PMT setup, report numerical and experimental results. Finally, we specify the imaging system performance based on numerical simulations and physical measurements.

physics.med-ph

X-ray luminescence computed tomography using a focused X-ray beam

Due to the low X-ray photon utilization efficiency and low measurement sensitivity of the electron multiplying charge coupled device (EMCCD) camera setup, the collimator based narrow beam X-ray luminescence computed tomography (XLCT) usually requires a long measurement time. In this paper, we, for the first time, report a focused X-ray beam based XLCT imaging system with measurements by a single optical fiber bundle and a photomultiplier tube (PMT). An X-ray tube with a polycapillary lens was used to generate a focused X-ray beam whose X-ray photon density is 1200 times larger than a collimated X-ray beam. An optical fiber bundle was employed to collect and deliver the emitted photons on the phantom surface to the PMT. The total measurement time was reduced to 12.5 minutes. For numerical simulations of both single and six fiber bundle cases, we were able to reconstruct six targets successfully. For the phantom experiment, two targets with an edge-to-edge distance of 0.4 mm and a center-to-center distance of 0.8 mm were successfully reconstructed by the measurement setup with a single fiber bundle and a PMT.

physics.med-ph