SearcharxivSearch

arXiv subjects

Hailiang Zhang

Publications and source records attributed to Hailiang Zhang.

12 recordsLinked to original sources

Zarr-Based Chunk-Level Cumulative Sums in Reduced Dimensions

Data analysis on massive multi-dimensional data, such as high-resolution large-region time averaging or area averaging for geospatial data, often involves calculations over a significant number of data points. While performing calculations in scalable and flexible distributed or cloud environments is a viable option, a full scan of large data volumes still serves as a computationally intensive bottleneck, leading to significant cost. This paper introduces a generic and comprehensive method to address these computational challenges. This method generates a small, size-tunable supplementary dataset that stores the cumulative sums along specific subset dimensions on top of the raw data. This minor addition unlocks rapid and cheap high-resolution large-region data analysis, making calculations over large numbers of data points feasible with small instances or even microservices in the cloud. This method is general-purpose, but is particularly well-suited for data stored in chunked, cloud-optimized formats and for services running in distributed or cloud environments. We present a Zarr extension proposal to integrate the specifications of this method and facilitate its straightforward implementation in general-purpose software applications. Benchmark tests demonstrate that this method, implemented in Amazon Web services (AWS), significantly outperforms the brute-force approach used in on-premises services. With just 5% supplemental storage, this method achieves a performance that is 3-4 orders of magnitude (~10,000 times) faster than the brute-force approach, while incurring significantly reduced computational costs.

cs.DC

The Solution for the CVPR2023 NICE Image Captioning Challenge

In this paper, we present our solution to the New frontiers for Zero-shot Image Captioning Challenge. Different from the traditional image captioning datasets, this challenge includes a larger new variety of visual concepts from many domains (such as COVID-19) as well as various image types (photographs, illustrations, graphics). For the data level, we collect external training data from Laion-5B, a large-scale CLIP-filtered image-text dataset. For the model level, we use OFA, a large-scale visual-language pre-training model based on handcrafted templates, to perform the image captioning task. In addition, we introduce contrastive learning to align image-text pairs to learn new visual concepts in the pre-training stage. Then, we propose a similarity-bucket strategy and incorporate this strategy into the template to force the model to generate higher quality and more matching captions. Finally, by retrieval-augmented strategy, we construct a content-rich template, containing the most relevant top-k captions from other image-text pairs, to guide the model in generating semantic-rich captions. Our method ranks first on the leaderboard, achieving 105.17 and 325.72 Cider-Score in the validation and test phase, respectively.

cs.CV

First Place Solution of 2023 Global Artificial Intelligence Technology Innovation Competition Track 1

In this paper, we present our champion solution to the Global Artificial Intelligence Technology Innovation Competition Track 1: Medical Imaging Diagnosis Report Generation. We select CPT-BASE as our base model for the text generation task. During the pre-training stage, we delete the mask language modeling task of CPT-BASE and instead reconstruct the vocabulary, adopting a span mask strategy and gradually increasing the number of masking ratios to perform the denoising auto-encoder pre-training task. In the fine-tuning stage, we design iterative retrieval augmentation and noise-aware similarity bucket prompt strategies. The retrieval augmentation constructs a mini-knowledge base, enriching the input information of the model, while the similarity bucket further perceives the noise information within the mini-knowledge base, guiding the model to generate higher-quality diagnostic reports based on the similarity prompts. Surprisingly, our single model has achieved a score of 2.321 on leaderboard A, and the multiple model fusion scores are 2.362 and 2.320 on the A and B leaderboards respectively, securing first place in the rankings.

cs.CL

The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA

In this paper, we introduce a grounded video question-answering solution. Our research reveals that the fixed official baseline method for video question answering involves two main steps: visual grounding and object tracking. However, a significant challenge emerges during the initial step, where selected frames may lack clearly identifiable target objects. Furthermore, single images cannot address questions like "Track the container from which the person pours the first time." To tackle this issue, we propose an alternative two-stage approach:(1) First, we leverage the VALOR model to answer questions based on video information.(2) concatenate the answered questions with their respective answers. Finally, we employ TubeDETR to generate bounding boxes for the targets.

cs.CV

A Spatial Calibration Method for Robust Cooperative Perception

Cooperative perception is a promising technique for intelligent and connected vehicles through vehicle-to-everything (V2X) cooperation, provided that accurate pose information and relative pose transforms are available. Nevertheless, obtaining precise positioning information often entails high costs associated with navigation systems. {Hence, it is required to calibrate relative pose information for multi-agent cooperative perception.} This paper proposes a simple but effective object association approach named context-based matching (CBM), which identifies inter-agent object correspondences using intra-agent geometrical context. In detail, this method constructs contexts using the relative position of the detected bounding boxes, followed by local context matching and global consensus maximization. The optimal relative pose transform is estimated based on the matched correspondences, followed by cooperative perception fusion. Extensive experiments are conducted on both the simulated and real-world datasets. Even with larger inter-agent localization errors, high object association precision and decimeter-level relative pose calibration accuracy are achieved among the cooperating agents.

cs.RO

NICE: CVPR 2023 Challenge on Zero-shot Image Captioning

In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project and share the results and outcomes of 2023 challenge. This project is designed to challenge the computer vision community to develop robust image captioning models that advance the state-of-the-art both in terms of accuracy and fairness. Through the challenge, the image captioning models were tested using a new evaluation dataset that includes a large variety of visual concepts from many domains. There was no specific training data provided for the challenge, and therefore the challenge entries were required to adapt to new types of image descriptions that had not been seen during training. This report includes information on the newly proposed NICE dataset, evaluation methods, challenge results, and technical details of top-ranking entries. We expect that the outcomes of the challenge will contribute to the improvement of AI models on various vision-language tasks.

cs.CV

A Cooperative Perception System Robust to Localization Errors

Cooperative perception is challenging for safety-critical autonomous driving applications.The errors in the shared position and pose cause an inaccurate relative transform estimation and disrupt the robust mapping of the Ego vehicle. We propose a distributed object-level cooperative perception system called OptiMatch, in which the detected 3D bounding boxes and local state information are shared between the connected vehicles. To correct the noisy relative transform, the local measurements of both connected vehicles (bounding boxes) are utilized, and an optimal transport theory-based algorithm is developed to filter out those objects jointly detected by the vehicles along with their correspondence, constructing an associated co-visible set. A correction transform is estimated from the matched object pairs and further applied to the noisy relative transform, followed by global fusion and dynamic mapping. Experiment results show that robust performance is achieved for different levels of location and heading errors, and the proposed framework outperforms the state-of-the-art benchmark fusion schemes, including early, late, and intermediate fusion, on average precision by a large margin when location and/or heading errors occur.

cs.MA

The Underlying Mechanisms of Time Dilation and Doppler Effect in Curved Space-Time

In this paper, we theoretically investigate the time dilation and Doppler effect in curved space-time from the perspective of quantum field theory (QFT). A Coordinate Transformation which Maintains the Period of Clocks is introduced, and such coordinate transformation is named as CTMPC throughout this paper. By analogy with the Lorentz transformation in Minkowski space-time, CTMPC is a correct transformation in curved space-times in a sense that it shows the correct relation between the time measured by the two observers, moreover, Lorentz transformation is just a special case of CTMPC applied in Minkowski space-time. We demonstrate that the Coordinate Transformation which Maintains the Local Metric (CTMLM) is one CTMPC, while the mathematical forms of physics formulas in QFT will be maintained. As applications of CTMLM, the time dilation and Doppler effect with an arbitrary time-dependent relative velocity in curved space-time are analysed. For Minkowski space-time, the time dilation and Doppler effect agree with the clock hypothesis. For curved space-time, we show that even if the emitted wave has a narrow frequency range, the Doppler effect may, in general, broaden the frequency spectrum and, at the meantime, shift the frequencies values. These new findings will deepen our understanding on the nature of space-time and the Doppler effect in curved space-time, they may also provide theoretical guidance in future astronomical observations.

hep-th

Non-line-of-sight polarized single-scatter propagation model for noncoplanar geometries

The classical model of non-line-of-sight (NLOS) single-scatter propagation for coplanar geometries was recently extended to include noncoplanar geometries; the calculation processes in the extended model are partly based on the Cartesian coordinate system and are somewhat complicated. A new NLOS single-scatter propagation model for noncoplanar geometries is presented based only on the prolate spheroidal coordinate system, which can be considered as the simplified version of the extended model mentioned above. Similar to the polarization-extension of the Monte-Carlo-based multiple-scatter model, the new single-scatter model for noncoplanar geometries is also extended to take polarization into account; the polarized single-scatter model is validated by the Monte-Carlo-based polarized model, results show perfect match. The theoretical feasibility of a 2-polarization UV communication link is validated based on the polarized single-scatter model.

physics.optics