SearcharxivSearch

arXiv subjects

Congcong Li

Publications and source records attributed to Congcong Li.

At least 19 recordsLinked to original sources

Higher odd-order nonlinear Hall effect in magnetic topological insulator Mn(Bi1-xSbx)2Te4

The nonlinear Hall effect is a new member of the Hall effect family, which attracts intense research interests, and it is closely related to the quantum geometry of quantum materials. The previous studies primarily concentrate on the second-order and third-order nonlinear Hall effect. However, the experimental study of higher-order nonlinear Hall effect is scarce at present. In this work, we report the observations of the higher odd-order (third-, fifth-, seventh-order) nonlinear Hall effect in magnetic topological insulator Mn(Bi1-xSbx)2Te4 thin flakes. The higher odd-order nonlinear Hall voltage exhibits a twofold angular dependence and exists only below the N\'eel temperature. It reaches its maximum near the charge neutral point and decays exponentially as the order of the nonlinear Hall effect increases. Furthermore, such higher odd-order nonlinear Hall effect is observed in both odd- and even-layer samples with comparable magnitudes. Theoretical analysis indicates that the higher odd-order nonlinear Hall effect responses may arise from the Berry curvature multipoles. Our work paves the way for the study of the higher-order nonlinear transport phenomena.

cond-mat.mes-hall

Room-temperature third-order nonlinear anomalous Hall effect in ferromagnetic metal Fe3GaTe2

Berry curvature, as the imaginary component of quantum geometry, plays a crucial role in condensed matter physics. The spatial distribution of Berry curvature can be characterized by its dipole and multipole moments, which can induce the nonlinear anomalous Hall effect (NLAHE). To date, the NLAHE has been demonstrated in various materials, yet reports on room-temperature NLAHE are still limited. In this work, we report the observation of the third-order NLAHE in ferromagnetic metal Fe3GaTe2. The third-order NLAHE shows hysteretic behavior with the variation of magnetic field, where the coercive field is the same as that of the anomalous Hall effect, and the third-order NLAHE remains observable up to the Curie temperature (~350 K). The scaling analysis suggests that the third-order NLAHE may be attributed to the Berry curvature quadrupole. Our work not only provides an approach to study magnetic materials through nonlinear electric transports, but also opens up possibilities for the future development of room-temperature third-order nonlinear electronic devices.

cond-mat.mtrl-sci

Structured Context Learning for Generic Event Boundary Detection

Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, called Structured Context Learning, which introduces the Structured Partition of Sequence (SPoS) to provide a structured context for learning temporal information. Our approach is end-to-end trainable and flexible, not restricted to specific temporal models like GRU, LSTM, and Transformers. This flexibility enables our method to achieve a better speed-accuracy trade-off. Specifically, we apply SPoS to partition the input frame sequence and provide a structured context for the subsequent temporal model. Notably, SPoS's overall computational complexity is linear with respect to the video length. We next calculate group similarities to capture differences between frames, and a lightweight fully convolutional network is utilized to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, we adapt the Gaussian kernel to preprocess the ground-truth event boundaries. Our proposed method has been extensively evaluated on the challenging Kinetics-GEBD, TAPOS, and shot transition detection datasets, demonstrating its superiority over existing state-of-the-art methods.

cs.CV

Car-GS: Addressing Reflective and Transparent Surface Challenges in 3D Car Reconstruction

3D car modeling is crucial for applications in autonomous driving systems, virtual and augmented reality, and gaming. However, due to the distinctive properties of cars, such as highly reflective and transparent surface materials, existing methods often struggle to achieve accurate 3D car reconstruction.To address these limitations, we propose Car-GS, a novel approach designed to mitigate the effects of specular highlights and the coupling of RGB and geometry in 3D geometric and shading reconstruction (3DGS). Our method incorporates three key innovations: First, we introduce view-dependent Gaussian primitives to effectively model surface reflections. Second, we identify the limitations of using a shared opacity parameter for both image rendering and geometric attributes when modeling transparent objects. To overcome this, we assign a learnable geometry-specific opacity to each 2D Gaussian primitive, dedicated solely to rendering depth and normals. Third, we observe that reconstruction errors are most prominent when the camera view is nearly orthogonal to glass surfaces. To address this issue, we develop a quality-aware supervision module that adaptively leverages normal priors from a pre-trained large-scale normal model.Experimental results demonstrate that Car-GS achieves precise reconstruction of car surfaces and significantly outperforms prior methods. The project page is available at https://lcc815.github.io/Car-GS.

cs.CV

Symbol-based multilevel block $\tau$ preconditioners for multilevel block Toeplitz systems: GLT-based analysis and applications

In recent years, there has been a renewed interest in preconditioning for multilevel Toeplitz systems, a research field that has been extensively explored over the past several decades. This work introduces novel preconditioning strategies using multilevel $\tau$ matrices for both symmetric and nonsymmetric multilevel Toeplitz systems. Our proposals constitute a general framework, as they are constructed solely based on the generating function of the multilevel Toeplitz coefficient matrix, when it can be defined. We begin with nonsymmetric systems, where we employ a symmetrization technique by permuting the coefficient matrix to produce a real symmetric multilevel Hankel structure. We propose a multilevel $\tau$ preconditioner tailored to the symmetrized system and prove that the eigenvalues of the preconditioned matrix sequence cluster at $\pm 1$, leading to rapid convergence when using the preconditioned minimal residual method. The high effectiveness of this approach is demonstrated through its application in solving space fractional diffusion equations. Next, for symmetric systems we introduce another multilevel $\tau$ preconditioner and show that the preconditioned conjugate gradient method can achieve an optimal convergence rate, namely a rate that is independent of the matrix size, when employed for a class of ill-conditioned multilevel Toeplitz systems. Numerical examples are provided to critically assess the effectiveness of our proposed preconditioners compared to several leading existing preconditioned solvers, highlighting their superior performance.

math.NA

Absolute-value based preconditioner for complex-shifted Laplacian systems

The complex-shifted Laplacian systems arising in a wide range of applications. In this work, we propose an absolute-value based preconditioner for solving the complex-shifted Laplacian system. In our approach, the complex-shifted Laplacian system is equivalently rewritten as a $2\times 2$ block real linear system. With the Toeplitz structure of uniform-grid discretization of the constant-coefficient Laplacian operator, the absolute value of the block real matrix is fast invertible by means of fast sine transforms. For more general coefficient function, we then average the coefficient function and take the absolute value of the averaged matrix as our preconditioner. With assumptions on the complex shift, we theoretically prove that the eigenvalues of the preconditioned matrix in absolute value are upper and lower bounded by constants independent of matrix size, indicating a matrix-size independent linear convergence rate of MINRES solver. Interestingly, numerical results show that the proposed preconditioner is still efficient even if the assumptions on the complex shift are not met. The fast invertibility of the proposed preconditioner and the robust convergence rate of the preconditioned MINRES solver lead to a linearithmic (nearly optimal) complexity of the proposed solver. The proposed preconditioner is compared with several state-of-the-art preconditioners via several numerical examples to demonstrate the efficiency of the proposed preconditioner.

math.NA

Multilevel Tau preconditioners for symmetrized multilevel Toeplitz systems with applications to solving space fractional diffusion equations

In this work, we develop a novel multilevel Tau matrix-based preconditioned method for a class of non-symmetric multilevel Toeplitz systems. This method not only accounts for but also improves upon an ideal preconditioner pioneered by [J. Pestana. Preconditioners for symmetrized Toeplitz and multilevel Toeplitz matrices. SIAM J. Matrix Anal. Appl., 40(3):870-887, 2019]. The ideal preconditioning approach was primarily examined numerically in that study, and an effective implementation was not included. To address these issues, we first rigorously show in this study that this ideal preconditioner can indeed achieve optimal convergence when employing the MINRES method, with a convergence rate is that independent of the mesh size. Then, building on this preconditioner, we develop a practical and optimal preconditioned MINRES method. To further illustrate its applicability and develop a fast implementation strategy, we consider solving Riemann-Liouville fractional diffusion equations as an application. Specifically, following standard discretization on the equation, the resultant linear system is a non-symmetric multilevel Toeplitz system, affirming the applicability of our preconditioning method. Through a simple symmetrization strategy, we transform the original linear system into a symmetric multilevel Hankel system. Subsequently, we propose a symmetric positive definite multilevel Tau preconditioner for the symmetrized system, which can be efficiently implemented using discrete sine transforms. Theoretically, we demonstrate that mesh-independent convergence can be achieved. In particular, we prove that the eigenvalues of the preconditioned matrix are bounded within disjoint intervals containing $\pm 1$, without any outliers. Numerical examples are provided to critically discuss the results, showcase the spectral distribution, and support the efficacy of our strategy.

math.NA

STT: Stateful Tracking with Transformers for Autonomous Driving

Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTPS that address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset.

cs.RO

Nonlinear Hall effect and scaling law in Sb-doped topological insulator MnBi4Te7

Nonlinear Hall effect (NLHE), as a new member of Hall effect family, has been realized in many materials, attracting a great deal of attention. Here, we report the observation of NLHE in magnetic topological insulator Sb-doped MnBi4Te7 flakes. The NLHE generation efficiency can reach up to 0.06 V^-1, which is comparable to that observed in MnBi2Te4. Differently, the NLHE can survive up to 200 K, much larger than the magnetic transition temperature. We further study the scaling behavior of the NLHE with longitudinal conductivity. The linear relationship with opposite slope when temperature is below and above the magnetic transition temperature is uncovered. It reveals that the NLHE originates from skew scattering. Our work provides a platform to search NLHE with larger generation efficiency at higher temperatures.

cond-mat.mtrl-sci

Local Compressed Video Stream Learning for Generic Event Boundary Detection

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which contains significant spatio-temporal redundancy and demands considerable computational power and storage space. To remedy these issues, we propose a novel compressed video representation learning method for event boundary detection that is fully end-to-end leveraging rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, without fully decoding the video. Specifically, we use lightweight ConvNets to extract features of the P-frames in the GOPs and spatial-channel attention module (SCAM) is designed to refine the feature representations of the P-frames based on the compressed information with bidirectional information flow. To learn a suitable representation for boundary detection, we construct the local frames bag for each candidate frame and use the long short-term memory (LSTM) module to capture temporal relationships. We then compute frame differences with group similarities in the temporal domain. This module is only applied within a local window, which is critical for event boundary detection. Finally a simple classifier is used to determine the event boundaries of video sequences based on the learned feature representation. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD and TAPOS datasets demonstrate that the proposed method achieves considerable improvements compared to previous end-to-end approach while running at the same speed. The code is available at https://github.com/GX77/LCVSL.

cs.CV

A preconditioned MINRES method for block lower triangular Toeplitz systems

In this study, a novel preconditioner based on the absolute-value block $\alpha$-circulant matrix approximation is developed, specifically designed for nonsymmetric dense block lower triangular Toeplitz (BLTT) systems that emerge from the numerical discretization of evolutionary equations. Our preconditioner is constructed by taking an absolute-value of a block $\alpha$-circulant matrix approximation to the BLTT matrix. To apply our preconditioner, the original BLTT linear system is converted into a symmetric form by applying a time-reversing permutation transformation. Then, with our preconditioner, the preconditioned minimal residual method (MINRES) solver is employed to solve the symmetrized linear system. With properly chosen $\alpha$, the eigenvalues of the preconditioned matrix are proven to be clustered around $\pm1$ without any significant outliers. With the clustered spectrum, we show that the preconditioned MINRES solver for the preconditioned system has a convergence rate independent of system size. To the best of our knowledge, this is the first preconditioned MINRES method with size-independent convergence rate for the dense BLTT system. The efficacy of the proposed preconditioner is corroborated by our numerical experiments, which reveal that it attains optimal convergence.

math.NA

Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory prediction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study.

cs.CV

SC-Transformer++: Structured Context Transformer for Generic Event Boundary Detection

This report presents the algorithm used in the submission of Generic Event Boundary Detection (GEBD) Challenge at CVPR 2022. In this work, we improve the existing Structured Context Transformer (SC-Transformer) method for GEBD. Specifically, a transformer decoder module is added after transformer encoders to extract high quality frame features. The final classification is performed jointly on the results of the original binary classifier and a newly introduced multi-class classifier branch. To enrich motion information, optical flow is introduced as a new modality. Finally, model ensemble is used to further boost performance. The proposed method achieves 86.49% F1 score on Kinetics-GEBD test set. which improves 2.86% F1 score compared to the previous SOTA method.

cs.CV

Depth Estimation Matters Most: Improving Per-Object Depth Estimation for Monocular 3D Detection and Tracking

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior performance when compared to LiDAR-based techniques. Through systematic analysis, we identified that per-object depth estimation accuracy is a major factor bounding the performance. Motivated by this observation, we propose a multi-level fusion method that combines different representations (RGB and pseudo-LiDAR) and temporal information across multiple frames for objects (tracklets) to enhance per-object depth estimation. Our proposed fusion method achieves the state-of-the-art performance of per-object depth estimation on the Waymo Open Dataset, the KITTI detection dataset, and the KITTI MOT dataset. We further demonstrate that by simply replacing estimated depth with fusion-enhanced depth, we can achieve significant improvements in monocular 3D perception tasks, including detection and tracking.

cs.CV

Structured Context Transformer for Generic Event Boundary Detection

Generic Event Boundary Detection (GEBD) aims to detect moments where humans naturally perceive as event boundaries. In this paper, we present Structured Context Transformer (or SC-Transformer) to solve the GEBD task, which can be trained in an end-to-end fashion. Specifically, we use the backbone convolutional neural network (CNN) to extract the features of each video frame. To capture temporal context information of each frame, we design the structure context transformer (SC-Transformer) by re-partitioning input frame sequence. Note that, the overall computation complexity of SC-Transformer is linear to the video length. After that, the group similarities are computed to capture the differences between frames. Then, a lightweight fully convolutional network is used to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, the Gaussian kernel is adopted to preprocess the ground-truth event boundaries to further boost the accuracy. Extensive experiments conducted on the challenging Kinetics-GEBD and TAPOS datasets demonstrate the effectiveness of the proposed method compared to the state-of-the-art methods.

cs.CV

End-to-End Compressed Video Representation Learning for Generic Event Boundary Detection

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which demands considerable computational power and storage space. To that end, we propose a new end-to-end compressed video representation learning for event boundary detection that leverages the rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, without fully decoding the video. Specifically, we first use the ConvNets to extract features of the I-frames in the GOPs. After that, a light-weight spatial-channel compressed encoder is designed to compute the feature representations of the P-frames based on the motion vectors, residuals and representations of their dependent I-frames. A temporal contrastive module is proposed to determine the event boundaries of video sequences. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD dataset demonstrate that the proposed method achieves comparable results to the state-of-the-art methods with $4.5\times$ faster running speed.

cs.CV

Multi-modal 3D Human Pose Estimation with 2D Weak Supervision in Autonomous Driving

3D human pose estimation (HPE) in autonomous vehicles (AV) differs from other use cases in many factors, including the 3D resolution and range of data, absence of dense depth maps, failure modes for LiDAR, relative location between the camera and LiDAR, and a high bar for estimation accuracy. Data collected for other use cases (such as virtual reality, gaming, and animation) may therefore not be usable for AV applications. This necessitates the collection and annotation of a large amount of 3D data for HPE in AV, which is time-consuming and expensive. In this paper, we propose one of the first approaches to alleviate this problem in the AV setting. Specifically, we propose a multi-modal approach which uses 2D labels on RGB images as weak supervision to perform 3D HPE. The proposed multi-modal architecture incorporates LiDAR and camera inputs with an auxiliary segmentation branch. On the Waymo Open Dataset, our approach achieves a 22% relative improvement over camera-only 2D HPE baseline, and 6% improvement over LiDAR-only model. Finally, careful ablation studies and parts based analysis illustrate the advantages of each of our contributions.

cs.CV

Generic Event Boundary Detection Challenge at CVPR 2021 Technical Report: Cascaded Temporal Attention Network (CASTANET)

This report presents the approach used in the submission of Generic Event Boundary Detection (GEBD) Challenge at CVPR21. In this work, we design a Cascaded Temporal Attention Network (CASTANET) for GEBD, which is formed by three parts, the backbone network, the temporal attention module, and the classification module. Specifically, the Channel-Separated Convolutional Network (CSN) is used as the backbone network to extract features, and the temporal attention module is designed to enforce the network to focus on the discriminative features. After that, the cascaded architecture is used in the classification module to generate more accurate boundaries. In addition, the ensemble strategy is used to further improve the performance of the proposed method. The proposed method achieves 83.30% F1 score on Kinetics-GEBD test set, which improves 20.5% F1 score compared to the baseline method. Code is available at https://github.com/DexiangHong/Cascade-PC.

cs.CV