Searcharxiv⌕ Search

arXiv subjects

Long Xu

Publications and source records attributed to Long Xu.

At least 37 records · Page 2Linked to original sources

When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding

Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video compression. Herein, we introduce a unified paradigm for Cross-Modality Video Coding (CMVC), which is a pioneering approach to explore multimodality representation and video generative models in video coding. Specifically, on the encoder side, we disentangle a video into spatial content and motion components, which are subsequently transformed into distinct modalities to achieve very compact representation by leveraging MLLMs. During decoding, previously encoded components and video generation models are leveraged to create multiple encoding-decoding modes that optimize video reconstruction quality for specific decoding requirements, including Text-Text-to-Video (TT2V) mode to ensure high-quality semantic information and Image-Text-to-Video (IT2V) mode to achieve superb perceptual consistency. In addition, we propose an efficient frame interpolation model for IT2V mode via Low-Rank Adaption (LoRA) tuning to guarantee perceptual quality, which allows the generated motion cues to behave smoothly. Experiments on benchmarks indicate that TT2V achieves effective semantic reconstruction, while IT2V exhibits competitive perceptual consistency. These results highlight potential directions for future research in video coding.

cs.CV↗

"Molecular waveplate" for the control of ultrashort pulses carrying orbital angular momentum

Ultrashort laser pulses carrying orbital angular momentum (OAM) have become essential tools in Atomic, Molecular, and Optical (AMO) studies, particularly for investigating strong-field light-matter interactions. However, controlling and generating ultrashort vortex pulses presents significant challenges, since their broad spectral content complicates manipulation with conventional optical elements, while the high peak power inherent in short-duration pulses risks damaging optical components. Here, we introduce a novel method for generating and controlling broadband ultrashort vortex beams by exploiting the non-adiabatic alignment of linear gas-phase molecules induced by vector beams. The interaction between the vector beam and the gas-phase molecules results in spatially varying polarizability, imparting a phase modulation to a probe laser. This process effectively creates a tunable ``molecular waveplate'' that adapts naturally to a broad spectral range. By leveraging this approach, we can generate ultrashort vortex pulses across a wide range of wavelengths. Under optimized gas pressure and interaction length conditions, this method allows for highly efficient conversion of circularly polarized light into the desired OAM pulse, thus enabling the generation of few-cycle, high-intensity vortex beams. This molecular waveplate, which overcomes the limitations imposed by conventional optical elements, opens up new possibilities for exploring strong-field physics, ultrafast science, and other applications that require high-intensity vortex beams.

physics.optics↗

Learning to Plan Maneuverable and Agile Flight Trajectory with Optimization Embedded Networks

In recent times, an increasing number of researchers have been devoted to utilizing deep neural networks for end-to-end flight navigation. This approach has gained traction due to its ability to bridge the gap between perception and planning that exists in traditional methods, thereby eliminating delays between modules. However, the practice of replacing original modules with neural networks in a black-box manner diminishes the overall system's robustness and stability. It lacks principled explanations and often fails to consistently generate high-quality motion trajectories. Furthermore, such methods often struggle to rigorously account for the robot's kinematic constraints, resulting in the generation of trajectories that cannot be executed satisfactorily. In this work, we combine the advantages of traditional methods and neural networks by proposing an optimization-embedded neural network. This network can learn high-quality trajectories directly from visual inputs without the need of mapping, while ensuring dynamic feasibility. Here, the deep neural network is employed to directly extract environment safety regions from depth images. Subsequently, we employ a model-based approach to represent these regions as safety constraints in trajectory optimization. Leveraging the availability of highly efficient optimization algorithms, our method robustly converges to feasible and optimal solutions that satisfy various user-defined constraints. Moreover, we differentiate the optimization process, allowing it to be trained as a layer within the neural network. This approach facilitates the direct interaction between perception and planning, enabling the network to focus more on the spatial regions where optimal solutions exist. As a result, it further enhances the quality and stability of the generated trajectories.

cs.RO↗

PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image Understanding

The previous advancements in pathology image understanding primarily involved developing models tailored to specific tasks. Recent studies has demonstrated that the large vision-language model can enhance the performance of various downstream tasks in medical image understanding. In this study, we developed a domain-specific large language-vision assistant (PA-LLaVA) for pathology image understanding. Specifically, (1) we first construct a human pathology image-text dataset by cleaning the public medical image-text data for domain-specific alignment; (2) Using the proposed image-text data, we first train a pathology language-image pretraining (PLIP) model as the specialized visual encoder for pathology image, and then we developed scale-invariant connector to avoid the information loss caused by image scaling; (3) We adopt two-stage learning to train PA-LLaVA, first stage for domain alignment, and second stage for end to end visual question \& answering (VQA) task. In experiments, we evaluate our PA-LLaVA on both supervised and zero-shot VQA datasets, our model achieved the best overall performance among multimodal models of similar scale. The ablation experiments also confirmed the effectiveness of our design. We posit that our PA-LLaVA model and the datasets presented in this work can promote research in field of computational pathology. All codes are available at: https://github.com/ddw2AIGROUP2CQUPT/PA-LLaVA}{https://github.com/ddw2AIGROUP2CQUPT/PA-LLaVA

cs.AI↗

ClickAttention: Click Region Similarity Guided Interactive Segmentation

Interactive segmentation algorithms based on click points have garnered significant attention from researchers in recent years. However, existing studies typically use sparse click maps as model inputs to segment specific target objects, which primarily affect local regions and have limited abilities to focus on the whole target object, leading to increased times of clicks. In addition, most existing algorithms can not balance well between high performance and efficiency. To address this issue, we propose a click attention algorithm that expands the influence range of positive clicks based on the similarity between positively-clicked regions and the whole input. We also propose a discriminative affinity loss to reduce the attention coupling between positive and negative click regions to avoid an accuracy decrease caused by mutual interference between positive and negative clicks. Extensive experiments demonstrate that our approach is superior to existing methods and achieves cutting-edge performance in fewer parameters. An interactive demo and all reproducible codes will be released at https://github.com/hahamyt/ClickAttention.

cs.CV↗

LF-3PM: a LiDAR-based Framework for Perception-aware Planning with Perturbation-induced Metric

Just as humans can become disoriented in featureless deserts or thick fogs, not all environments are conducive to the Localization Accuracy and Stability (LAS) of autonomous robots. This paper introduces an efficient framework designed to enhance LiDAR-based LAS through strategic trajectory generation, known as Perception-aware Planning. Unlike vision-based frameworks, the LiDAR-based requires different considerations due to unique sensor attributes. Our approach focuses on two main aspects: firstly, assessing the impact of LiDAR observations on LAS. We introduce a perturbation-induced metric to provide a comprehensive and reliable evaluation of LiDAR observations. Secondly, we aim to improve motion planning efficiency. By creating a Static Observation Loss Map (SOLM) as an intermediary, we logically separate the time-intensive evaluation and motion planning phases, significantly boosting the planning process. In the experimental section, we demonstrate the effectiveness of the proposed metrics across various scenes and the feature of trajectories guided by different metrics. Ultimately, our framework is tested in a real-world scenario, enabling the robot to actively choose topologies and orientations preferable for localization. The source code is accessible at https://github.com/ZJU-FAST-Lab/LF-3PM.

cs.RO↗

Intelligence Preschool Education System based on Multimodal Interaction Systems and AI

Rapid progress in AI technologies has generated considerable interest in their potential to address challenges in every field and education is no exception. Improving learning outcomes and providing relevant education to all have been dominant themes universally, both in the developed and developing world. And they have taken on greater significance in the current era of technology driven personalization.

cs.HC↗

Structured Click Control in Transformer-based Interactive Segmentation

Click-point-based interactive segmentation has received widespread attention due to its efficiency. However, it's hard for existing algorithms to obtain precise and robust responses after multiple clicks. In this case, the segmentation results tend to have little change or are even worse than before. To improve the robustness of the response, we propose a structured click intent model based on graph neural networks, which adaptively obtains graph nodes via the global similarity of user-clicked Transformer tokens. Then the graph nodes will be aggregated to obtain structured interaction features. Finally, the dual cross-attention will be used to inject structured interaction features into vision Transformer features, thereby enhancing the control of clicks over segmentation results. Extensive experiments demonstrated the proposed algorithm can serve as a general structure in improving Transformer-based interactive segmenta?tion performance. The code and data will be released at https://github.com/hahamyt/scc.

cs.CV↗

MST: Adaptive Multi-Scale Tokens Guided Interactive Segmentation

Interactive segmentation has gained significant attention for its application in human-computer interaction and data annotation. To address the target scale variation issue in interactive segmentation, a novel multi-scale token adaptation algorithm is proposed. By performing top-k operations across multi-scale tokens, the computational complexity is greatly simplified while ensuring performance. To enhance the robustness of multi-scale token selection, we also propose a token learning algorithm based on contrastive loss. This algorithm can effectively improve the performance of multi-scale token adaptation. Extensive benchmarking shows that the algorithm achieves state-of-the-art (SOTA) performance, compared to current methods. An interactive demo and all reproducible codes will be released at https://github.com/hahamyt/mst.

cs.CV↗

Enhanced Persistent Orientation of Asymmetric-Top Molecules Induced by Cross-Polarized Terahertz Pulses

We investigate the persistent orientation of asymmetric-top molecules induced by time-delayed THz pulses that are either collinearly or cross polarized. Our theoretical and numerical results demonstrate that the orthogonal configuration outperforms the collinear one, and a significant degree of persistent orientation - approximately 10% at 5 K and nearly 3% at room temperature - may be achieved through parameter optimization. The dependence of the persistent orientation factor on temperature and field parameters is studied in detail. The proposed application of two orthogonally polarized THz pulses is both practical and efficient. Its applicability under standard laboratory conditions lays a solid foundation for future experimental realization of THz-induced persistent molecular orientation.

physics.chem-ph↗

PSF-based Analysis for Detecting Unresolved Wide Binaries

Wide binaries play a crucial role in analyzing the birth environment of stars and the dynamical evolution of clusters. When wide binaries are located at greater distances, their companions may overlap in the observed images, becoming indistinguishable and resulting in unresolved wide binaries, which are difficult to detect using traditional methods. Utilizing deep learning, we present a method to identify unresolved wide binaries by analyzing the point-spread function (PSF) morphology of telescopes. Our trained model demonstrates exceptional performance in differentiating between single stars and unresolved binaries with separations ranging from 0.1 to 2 physical pixels, where the PSF FWHM is ~2 pixels, achieving an accuracy of 97.2% for simulated data from the Chinese Space Station Telescope. We subsequently tested our method on photometric data of NGC 6121 observed by the Hubble Space Telescope. The trained model attained an accuracy of 96.5% and identified 18 wide binary candidates with separations between 7 and 140 au. The majority of these wide binary candidates are situated outside the core radius of NGC 6121, suggesting that they are likely first-generation stars, which is in general agreement with the results of Monte Carlo simulations. Our PSF-based method shows great promise in detecting unresolved wide binaries and is well suited for observations from space-based telescopes with stable PSF. In the future, we aim to apply our PSF-based method to next-generation surveys such as the China Space Station Optical Survey, where a larger-field-of-view telescope will be capable of identifying a greater number of such wide binaries.

astro-ph.SR↗

An Efficient Trajectory Planner for Car-like Robots on Uneven Terrain

Autonomous navigation of ground robots on uneven terrain is being considered in more and more tasks. However, uneven terrain will bring two problems to motion planning: how to assess the traversability of the terrain and how to cope with the dynamics model of the robot associated with the terrain. The trajectories generated by existing methods are often too conservative or cannot be tracked well by the controller since the second problem is not well solved. In this paper, we propose terrain pose mapping to describe the impact of terrain on the robot. With this mapping, we can obtain the SE(3) state of the robot on uneven terrain for a given state in SE(2). Then, based on it, we present a trajectory optimization framework for car-like robots on uneven terrain that can consider both of the above problems. The trajectories generated by our method conform to the dynamics model of the system without being overly conservative and yet able to be tracked well by the controller. We perform simulations and real-world experiments to validate the efficiency and trajectory quality of our algorithm.

cs.RO↗

Decentralized Planning for Car-Like Robotic Swarm in Cluttered Environments

Robot swarm is a hot spot in robotic research community. In this paper, we propose a decentralized framework for car-like robotic swarm which is capable of real-time planning in cluttered environments. In this system, path finding is guided by environmental topology information to avoid frequent topological change, and search-based speed planning is leveraged to escape from infeasible initial value's local minima. Then spatial-temporal optimization is employed to generate a safe, smooth and dynamically feasible trajectory. During optimization, the trajectory is discretized by fixed time steps. Penalty is imposed on the signed distance between agents to realize collision avoidance, and differential flatness cooperated with limitation on front steer angle satisfies the non-holonomic constraints. With trajectories broadcast to the wireless network, agents are able to check and prevent potential collisions. We validate the robustness of our system in simulation and real-world experiments. Code will be released as open-source packages.

cs.RO↗

An Efficient Spatial-Temporal Trajectory Planner for Autonomous Vehicles in Unstructured Environments

As a core part of autonomous driving systems, motion planning has received extensive attention from academia and industry. However, real-time trajectory planning capable of spatial-temporal joint optimization is challenged by nonholonomic dynamics, particularly in the presence of unstructured environments and dynamic obstacles. To bridge the gap, we propose a real-time trajectory optimization method that can generate a high-quality whole-body trajectory under arbitrary environmental constraints. By leveraging the differential flatness property of car-like robots, we simplify the trajectory representation and analytically formulate the planning problem while maintaining the feasibility of the nonholonomic dynamics. Moreover, we achieve efficient obstacle avoidance with a safe driving corridor for unmodelled obstacles and signed distance approximations for dynamic moving objects. We present comprehensive benchmarks with State-of-the-Art methods, demonstrating the significance of the proposed method in terms of efficiency and trajectory quality. Real-world experiments verify the practicality of our algorithm. We will release our codes for the research community

cs.RO↗

Towards Efficient Trajectory Generation for Ground Robots beyond 2D Environment

With the development of robotics, ground robots are no longer limited to planar motion. Passive height variation due to complex terrain and active height control provided by special structures on robots require a more general navigation planning framework beyond 2D. Existing methods rarely considers both simultaneously, limiting the capabilities and applications of ground robots. In this paper, we proposed an optimization-based planning framework for ground robots considering both active and passive height changes on the z-axis. The proposed planner first constructs a penalty field for chassis motion constraints defined in R3 such that the optimal solution space of the trajectory is continuous, resulting in a high-quality smooth chassis trajectory. Also, by constructing custom constraints in the z-axis direction, it is possible to plan trajectories for different types of ground robots which have z-axis degree of freedom. We performed simulations and realworld experiments to verify the efficiency and trajectory quality of our algorithm.

cs.RO↗

Ionization-induced Long-lasting Orientation of Symmetric-top Molecules

We theoretically consider the phenomenon of field-free long-lasting orientation of symmetric-top molecules ionized by two-color laser pulses. The anisotropic ionization produces a significant long-lasting orientation of the surviving neutral molecules. The degree of orientation increases with both the pulse intensity and, counterintuitively, with the rotational temperature. The orientation may be enhanced even further by using multiple delayed two-color pulses. The long-lasting orientation may be probed by even harmonic generation or by Coulomb-explosion-based methods. The effect may enable the study of relaxation processes in dense molecular gases, and may be useful for molecular guiding and trapping by inhomogeneous fields.

physics.chem-ph↗

Contextual Modeling for 3D Dense Captioning on Point Clouds

3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence for describing each located object. However, the existing methods mainly focus on mining inter-object relationship, while ignoring contextual information, especially the non-object details and background environment within the point clouds, thus leading to low-quality descriptions, such as inaccurate relative position information. In this paper, we make the first attempt to utilize the point clouds clustering features as the contextual information to supply the non-object details and background environment of the point clouds and incorporate them into the 3D dense captioning task. We propose two separate modules, namely the Global Context Modeling (GCM) and Local Context Modeling (LCM), in a coarse-to-fine manner to perform the contextual modeling of the point clouds. Specifically, the GCM module captures the inter-object relationship among all objects with global contextual information to obtain more complete scene information of the whole point clouds. The LCM module exploits the influence of the neighboring objects of the target object and local contextual information to enrich the object representations. With such global and local contextual modeling strategies, our proposed model can effectively characterize the object representations and contextual information and thereby generate comprehensive and detailed descriptions of the located objects. Extensive experiments on the ScanRefer and Nr3D datasets demonstrate that our proposed method sets a new record on the 3D dense captioning task, and verify the effectiveness of our raised contextual modeling of point clouds.

cs.CV↗

Echo-enhanced molecular orientation at high temperatures

Ultrashort laser pulses are widely used for transient field-free molecular orientation -- a phenomenon important in chemical reaction dynamics, ultrafast molecular imaging, high harmonics generation, and attosecond science. However, significant molecular orientation usually requires rotationally cold molecules, like in rarified molecular beams, because chaotic thermal motion is detrimental to the orientation process. Here we propose to use the mechanism of the echo phenomenon previously observed in hadron accelerators, free-electron lasers, and laser-excited molecules to overcome the destructive thermal effects and achieve efficient field-free molecular orientation at high temperatures. In our scheme, a linearly polarized short laser pulse transforms a broad thermal distribution in the molecular rotational phase space into many separated narrow filaments due to the nonlinear phase mixing during the post-pulse free evolution. Molecular subgroups belonging to individual filaments have much-reduced dispersion of angular velocities. They are rotationally cold, and a subsequent moderate terahertz (THz) pulse can easily orient them. The overall enhanced orientation of the molecular gas is achieved with some delay, in the course of the echo process combining the contributions of different filaments. Our results demonstrate that the echo-enhanced orientation is an order of magnitude higher than that of the THz pulse alone. The mechanism is robust -- it applies to different types of molecules, and the degree of orientation is relatively insensitive to the temperature. The laser and THz pulses used in the scheme are readily available, allowing quick experimental demonstration and testing in various applications. Breaking the phase space to individual filaments to overcome hindering thermal conditions may find a wide range of applications beyond molecular orientation.

physics.chem-ph↗