SearcharxivSearch

arXiv subjects

Jiankai Li

Publications and source records attributed to Jiankai Li.

10 recordsLinked to original sources

Beyond Benchmarks of IUGC: Rethinking Requirements of Deep Learning Methods for Intrapartum Ultrasound Biometry from Fetal Ultrasound Videos

A substantial proportion (45\%) of maternal deaths, neonatal deaths, and stillbirths occur during the intrapartum phase, with a particularly high burden in low- and middle-income countries. Intrapartum biometry plays a critical role in monitoring labor progression; however, the routine use of ultrasound in resource-limited settings is hindered by a shortage of trained sonographers. To address this challenge, the Intrapartum Ultrasound Grand Challenge (IUGC), co-hosted with MICCAI 2024, was launched. The IUGC introduces a clinically oriented multi-task automatic measurement framework that integrates standard plane classification, fetal head-pubic symphysis segmentation, and biometry, enabling algorithms to exploit complementary task information for more accurate estimation. Furthermore, the challenge releases the largest multi-center intrapartum ultrasound video dataset to date, comprising 774 videos (68,106 frames) collected from three hospitals, providing a robust foundation for model training and evaluation. In this study, we present a comprehensive overview of the challenge design, review the submissions from eight participating teams, and analyze their methods from five perspectives: preprocessing, data augmentation, learning strategy, model architecture, and post-processing. In addition, we perform a systematic analysis of the benchmark results to identify key bottlenecks, explore potential solutions, and highlight open challenges for future research. Although encouraging performance has been achieved, our findings indicate that the field remains at an early stage, and further in-depth investigation is required before large-scale clinical deployment. All benchmark solutions and the complete dataset have been publicly released to facilitate reproducible research and promote continued advances in automatic intrapartum ultrasound biometry.

cs.CV

Multi-Grained Compositional Visual Clue Learning for Image Intent Recognition

In an era where social media platforms abound, individuals frequently share images that offer insights into their intents and interests, impacting individual life quality and societal stability. Traditional computer vision tasks, such as object detection and semantic segmentation, focus on concrete visual representations, while intent recognition relies more on implicit visual clues. This poses challenges due to the wide variation and subjectivity of such clues, compounded by the problem of intra-class variety in conveying abstract concepts, e.g. "enjoy life". Existing methods seek to solve the problem by manually designing representative features or building prototypes for each class from global features. However, these methods still struggle to deal with the large visual diversity of each intent category. In this paper, we introduce a novel approach named Multi-grained Compositional visual Clue Learning (MCCL) to address these challenges for image intent recognition. Our method leverages the systematic compositionality of human cognition by breaking down intent recognition into visual clue composition and integrating multi-grained features. We adopt class-specific prototypes to alleviate data imbalance. We treat intent recognition as a multi-label classification problem, using a graph convolutional network to infuse prior knowledge through label embedding correlations. Demonstrated by a state-of-the-art performance on the Intentonomy and MDID datasets, our approach advances the accuracy of existing methods while also possessing good interpretability. Our work provides an attempt for future explorations in understanding complex and miscellaneous forms of human expression.

cs.CV

Leveraging Predicate and Triplet Learning for Scene Graph Generation

Scene Graph Generation (SGG) aims to identify entities and predict the relationship triplets \textit{\textless subject, predicate, object\textgreater } in visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate, it can be quite challenging to model and refine predicate representations directly across such pairs, which is however a common strategy adopted by most existing SGG methods. We observe that visual variations within the identical triplet are relatively small and certain relation cues are shared in the same type of triplet, which can potentially facilitate the relation learning in SGG. Moreover, for the long-tail problem widely studied in SGG task, it is also crucial to deal with the limited types and quantity of triplets in tail predicates. Accordingly, in this paper, we propose a Dual-granularity Relation Modeling (DRM) network to leverage fine-grained triplet cues besides the coarse-grained predicate ones. DRM utilizes contexts and semantics of predicate and triplet with Dual-granularity Constraints, generating compact and balanced representations from two perspectives to facilitate relation recognition. Furthermore, a Dual-granularity Knowledge Transfer (DKT) strategy is introduced to transfer variation from head predicates/triplets to tail ones, aiming to enrich the pattern diversity of tail classes to alleviate the long-tail problem. Extensive experiments demonstrate the effectiveness of our method, which establishes new state-of-the-art performance on Visual Genome, Open Image, and GQA datasets. Our code is available at \url{https://github.com/jkli1998/DRM}

cs.CV

InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization

Recent strides in the development of diffusion models, exemplified by advancements such as Stable Diffusion, have underscored their remarkable prowess in generating visually compelling images. However, the imperative of achieving a seamless alignment between the generated image and the provided prompt persists as a formidable challenge. This paper traces the root of these difficulties to invalid initial noise, and proposes a solution in the form of Initial Noise Optimization (InitNO), a paradigm that refines this noise. Considering text prompts, not all random noises are effective in synthesizing semantically-faithful images. We design the cross-attention response score and the self-attention conflict score to evaluate the initial noise, bifurcating the initial latent space into valid and invalid sectors. A strategically crafted noise optimization pipeline is developed to guide the initial noise towards valid regions. Our method, validated through rigorous experimentation, shows a commendable proficiency in generating images in strict accordance with text prompts. Our code is available at https://github.com/xiefan-guo/initno.

cs.CV

Zero-Shot Scene Graph Generation via Triplet Calibration and Reduction

Scene Graph Generation (SGG) plays a pivotal role in downstream vision-language tasks. Existing SGG methods typically suffer from poor compositional generalizations on unseen triplets. They are generally trained on incompletely annotated scene graphs that contain dominant triplets and tend to bias toward these seen triplets during inference. To address this issue, we propose a Triplet Calibration and Reduction (T-CAR) framework in this paper. In our framework, a triplet calibration loss is first presented to regularize the representations of diverse triplets and to simultaneously excavate the unseen triplets in incompletely annotated training scene graphs. Moreover, the unseen space of scene graphs is usually several times larger than the seen space since it contains a huge number of unrealistic compositions. Thus, we propose an unseen space reduction loss to shift the attention of excavation to reasonable unseen compositions to facilitate the model training. Finally, we propose a contextual encoder to improve the compositional generalizations of unseen triplets by explicitly modeling the relative spatial relations between subjects and objects. Extensive experiments show that our approach achieves consistent improvements for zero-shot SGG over state-of-the-art methods. The code is available at https://github.com/jkli1998/T-CAR.

cs.CV

Random lasing actions in self-assembled perovskite nanoparticles

Solution-based perovskite nanoparticles have been intensively studied in past few years due to their applications in both photovoltaic and optoelectronic devices. Here, based on the common ground between the solution-based perovskite and random lasers, we have studied the mirrorless lasing actions in self-assembled perovskite nanoparticles. After the synthesis from solution, discrete lasing peaks have been observed from the optically pumped perovskites without any well-defined cavity boundaries. The obtained quality (Q) factors and thresholds of random lasers are around 500 and 60 uJ/cm2, respectively. Both values are comparable to the conventional perovskite microdisk lasers with polygon shaped cavity boundaries. From the corresponding studies on laser spectra and fluorescence microscope images, the lasing actions are considered as random lasers that are generated by strong multiple scattering in random gain media. In additional to conventional single-photon excitation, due to the strong nonlinear effects of perovskites, two-photon pumped random lasers have also been demonstrated for the first time. We believe this research will find its potential applications in low-cost coherent light sources and biomedical detection.

physics.optics

Formation of Single-mode Laser in Perovskite Nanowire via Nano-manipulation

Perovskite based micro- and nano- lasers have attracted considerable research attention in past two years. However, the properties of perovskite devices are mostly fixed once they are synthesized. Here we demonstrate the tailoring of lasing properties of perovskite nanowire lasers via nano-manipulation. By utilizing a tungsten probe, one nanowire has been lifted from the wafer and re-positioned its two ends on two nearby perovskite blocks. Consequently, the conventional Fabry-Perot lasers are completely suppressed and a single laser peak has been observed. The corresponding numerical model reveals that the single-mode lasing operation is formed by the whispering gallery mode in the transverse plane of perovskite nanowire. Our research provides a simple way to tailor the properties of nanowire and it will be essential for the applications of perovskite optoelectronics.

physics.optics

The combination of directional outputs and single-mode operation in circular microdisk with broken PT symmetry

Monochromaticity and directionality are two key characteristics of lasers. However, the combination of directional emission and single-mode operation is quite challenging, especially for the on-chip devices. Here we propose a microdisk laser with single-mode operation and directional emissions by exploiting the recent developments associated with parity-time (PT) symmetry. This is accomplished by introducing one-dimensional periodic gain and loss into a circular microdisk, which induces a coupling between whispering gallery modes with different radial numbers. The lowest threshold mode is selected at the positions with least initial wavelength difference. And the directional emissions are formed by the introduction of additional grating vectors by the periodic distribution of gain and loss regions. We believe this research will impact the practical applications of on-chip microdisk lasers.

physics.optics

Inversed Vernier effect based single-mode laser emission in coupled microdisks

Recently, on-chip single-mode laser emission has attracted considerable research attention due to its wide applications. While most of single-mode lasers in coupled microdisks or microrings have been qualitatively explained by either Vernier effect or inversed Vernier effect, none of them have been experimentally confirmed. Here, we studied the mechanism for single-mode operation in coupled microdisks. We found that the mode numbers had been significantly reduced to nearly single-mode within a large pumping power range from threshold to gain saturation. The detail laser spectra showed that the largest gain and the first lasing peak were mainly generated by one disk and the laser intensity was proportional to the frequency detuning. The corresponding theoretical analysis showed that the experimental observations were dominated by internal coupling within one cavity, which was similar to the recently explored inversed Vernier effect in two coupled microrings. We believe our finding will be important for understanding the previous experimental findings and the development of on-chip single-mode laser.

physics.optics

Quasi-guiding modes in microfibers on high refractive index substrate

Light confinement and amplification in micro- & nano-fiber have been intensively studied and a number of applications have been developed. However, the typical micro- & anno- fibers are usually free-standing or positioned on a substrate with lower refractive index to ensure the light confinement and guiding mode. Here we numerically and experimentally demonstrate the possibility of confining light within a microfiber on a high refractive index substrate. In contrast to the strong leaky to the substrate, we found that the radiation loss was dependent on the radius of microfiber and the refractive index contrast. Consequently, quasi-guiding modes could be formed and the light could propagate and be amplified in such systems. By fabricating tapered silica fiber and dye-doped polymer fiber and placing them on sapphire substrates, the light propagation, amplification, and laser behaviors have been experimentally studied to verify the quasi-guiding modes in microfer with higher index substrate. We believe that our research will be essential for the applications of micro- and nano-fibers.

physics.optics