SearcharxivSearch

arXiv subjects

Yongjian Zhao

Publications and source records attributed to Yongjian Zhao.

5 recordsLinked to original sources

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines (i) a declarative configuration interface covering model choice, prompt set, and inference engine, (ii) a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and (iii) a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straight-forward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services.

cs.LG

Enabling Augmented Segmentation and Registration in Ultrasound-Guided Spinal Surgery via Realistic Ultrasound Synthesis from Diagnostic CT Volume

This paper aims to tackle the issues on unavailable or insufficient clinical US data and meaningful annotation to enable bone segmentation and registration for US-guided spinal surgery. While the US is not a standard paradigm for spinal surgery, the scarcity of intra-operative clinical US data is an insurmountable bottleneck in training a neural network. Moreover, due to the characteristics of US imaging, it is difficult to clearly annotate bone surfaces which causes the trained neural network missing its attention to the details. Hence, we propose an In silico bone US simulation framework that synthesizes realistic US images from diagnostic CT volume. Afterward, using these simulated bone US we train a lightweight vision transformer model that can achieve accurate and on-the-fly bone segmentation for spinal sonography. In the validation experiments, the realistic US simulation was conducted by deriving from diagnostic spinal CT volume to facilitate a radiation-free US-guided pedicle screw placement procedure. When it is employed for training bone segmentation task, the Chamfer distance achieves 0.599mm; when it is applied for CT-US registration, the associated bone segmentation accuracy achieves 0.93 in Dice, and the registration accuracy based on the segmented point cloud is 0.13~3.37mm in a complication-free manner. While bone US images exhibit strong echoes at the medium interface, it may enable the model indistinguishable between thin interfaces and bone surfaces by simply relying on small neighborhood information. To overcome these shortcomings, we propose to utilize a Long-range Contrast Learning Module to fully explore the Long-range Contrast between the candidates and their surrounding pixels.

eess.IV

Exact Sparse Orthogonal Dictionary Learning

Over the past decade, learning a dictionary from input images for sparse modeling has been one of the topics which receive most research attention in image processing and compressed sensing. Most existing dictionary learning methods consider an over-complete dictionary, such as the K-SVD method, which may result in high mutual incoherence and therefore has a negative impact in recognition. On the other side, the sparse codes are usually optimized by adding the $\ell_0$ or $\ell_1$-norm penalty, but with no strict sparsity guarantee. In this paper, we propose an orthogonal dictionary learning model which can obtain strictly sparse codes and orthogonal dictionary with global sequence convergence guarantee. We find that our method can result in better denoising results than over-complete dictionary based learning methods, and has the additional advantage of high computation efficiency.

eess.IV

Light-scanning hand-held photoacoustic probe design

Significance: We proposed a new design of hand-held linear-array photoacoustic (PA) probe which can acquire multi images via motor moving. Moreover, images from different locations are utilized via imaging fusion for SNR enhancement. Aim: We devised an adjustable hand-held for the purpose of realizing different images at diverse location for further image fusion. For realizing the light spot which is more matched with the Ultrasonic transducer detection area, we specially design a light adjust unit. Moreover, due to no displacement among the images, there is no need to execute image register process. The program execution time be reduced, greatly. Approach: mechanical design; Montel carol simulation; no-registration image fusion; Spot compression. Results: Multiple PA images with different optical illumination areas were acquired. After image fusion, we obtained fused PA images with higher signal-to-noise-ratio (SNR) and image fidelity than each single PA image. A quantitative comparison shows that the SNR of fused image is improved by 36.06% in agar-milk phantom, and 44.69% in chicken breast phantom, respectively. Conclusions: In this paper, the light scanning adjustable hand-held PA imaging probe is proposed, which can realize PA imaging with different illumination positions via adjusting the optical unit.

physics.med-ph

Portable probe design for photoacoustic imaging in vivo

A low-cost adjustable illumination scheme for hand-held photoacoustic imaging probe is presented, manufactured and tested in this paper. Compared with traditional photoacoustic probe design, it has the following advantages: (1) Different excitation modes can be selected as needed. By tuning control parameters, it can achieve bright-field, dark-field, and hybrid field light illumination schemes. (2) The spot-adjustable unit (SAU) specifically designed for beam expansion, together with a water tank for transmitting ultrasonic waves, enable the device to break through the constraints of the transfer medium and is more widely used. The beam-expansion experiment is conducted to verify the function of SAU. After that, we built a PAT system based on our newly designed apparatus. Phantom and in vivo experimental results show different performance in different illumination schemes

eess.IV