SearcharxivSearch

arXiv subjects

Jinxiang Wang

Publications and source records attributed to Jinxiang Wang.

16 recordsLinked to original sources

On the well-posedness and efficient approximation for the classical Melan equation in suspension bridges

The classical Melan equation modeling suspension bridges is considered. We first study the explicit expression and global properties of the analytical solution for the simplified ``less stiff'' model, based on which we derive an a priori estimate for the original classical Melan equation and establish, under an explicit load condition, that it has a unique solution with nonnegative integral under downward live loads, thereby showing the uniqueness of the corresponding deflection curve in the engineering setting. We also develop an efficient iterative approximation method by taking the solution of the simplified ``less stiff'' model as the first iterate, and prove its geometric convergence with explicit error estimates. The applicability and computational efficiency of the method are demonstrated through calculations for two actual bridges, which also quantify the influence of the nonlinear nonlocal term on the solution and clarify the relationship between the simplified and original models. Several engineering observations are verified and explained, and some related open problems are suggested.

math.NA

PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding

Implicit neural representation (INR) has achieved remarkable progress in novel view synthesis and image/video coding in recent years.Compared to conventional end-to-end image codecs, INR-based compressors demonstrate significant advantages in decoding complexity. However, their practical application has been hindered by the inferior encoding speed and underutilized decoding efficiency.In this work, we propose a feedforward INR image coding architecture, Practical INR Image Codec (PIC), that computes all the necessary information for INR network in a single forward pass, achieving an encoding speed of 20 FPS. Additionally, we implement a highly optimized decoder that reaches 2000 FPS decoding speed, significantly surpassing JPEG's performance at comparable rate-distortion (RD) performance. To the best of our knowledge, this work presents the first learning-based image codec that simultaneously outperforms or is comparable with JPEG in both RD performance and decoding speed while maintaining practical encoding speed. Code is available at https://github.com/actcwlf/PIC.

cs.CV

Splatwizard: A Benchmark Toolkit for 3D Gaussian Splatting Compression

The recent advent of 3D Gaussian Splatting (3DGS) has marked a significant breakthrough in real-time novel view synthesis. However, the rapid proliferation of 3DGS-based algorithms has created a pressing need for standardized and comprehensive evaluation tools, especially for compression task. Existing benchmarks often lack the specific metrics necessary to holistically assess the unique characteristics of different methods, such as rendering speed, rate distortion trade-offs memory efficiency, and geometric accuracy. To address this gap, we introduce Splatwizard, a unified benchmark toolkit designed specifically for benchmarking 3DGS compression models. Splatwizard provides an easy-to-use framework to implement new 3DGS compression model and utilize state-of-the-art techniques proposed by previous work. Besides, an integrated pipeline that automates the calculation of key performance indicators, including image-based quality metrics, chamfer distance of reconstruct mesh, rendering frame rates, and computational resource consumption is included in the framework as well. Code is available at https://github.com/splatwizard/splatwizard

cs.CV

Existence, uniqueness, and approximability of solutions to the classical Melan equation in suspension bridges

The classical Melan equation modeling suspension bridges is considered. We first study the explicit expression and the uniform positivity of the analytical solution for the simplified ``less stiff'' model, based on which we develop a monotone iterative technique of lower and upper solutions to investigate the existence, uniqueness and approximability of the solution for the original classical Melan equation.The applicability and the efficiency of the monotone iterative technique for engineering design calculations are discussed by verifying some examples of actual bridges. Some open problems are suggested.

math.CA

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs

Large language models (LLMs) have made a profound impact across various fields due to their advanced capabilities. However, training these models at unprecedented scales requires extensive AI accelerator clusters and sophisticated parallelism strategies, which pose significant challenges in maintaining system reliability over prolonged training periods. A major concern is the substantial loss of training time caused by inevitable hardware and software failures. To address these challenges, we present FlashRecovery, a fast and low-cost failure recovery system comprising three core modules: (1) Active and real-time failure detection. This module performs continuous training state monitoring, enabling immediate identification of hardware and software failures within seconds, thus ensuring rapid incident response; (2) Scale-independent task restart. By employing different recovery strategies for normal and faulty nodes, combined with an optimized communication group reconstruction protocol, our approach ensures that the recovery time remains nearly constant, regardless of cluster scale; (3) Checkpoint-free recovery within one step. Our novel recovery mechanism enables single-step restoration, completely eliminating dependence on traditional checkpointing methods and their associated overhead. Collectively, these innovations enable FlashRecovery to achieve optimal Recovery Time Objective (RTO) and Recovery Point Objective (RPO), substantially improving the reliability and efficiency of long-duration LLM training. Experimental results demonstrate that FlashRecovery system can achieve training restoration on training cluster with 4, 800 devices in 150 seconds. We also verify that the time required for failure recovery is nearly consistent for different scales of training tasks.

cs.DC

GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View

Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Although there are various tasks for evaluating multimodal reasoning ability, they still have limitations. Reasoning on hierarchical visual clues at different levels of granularity, i.e., local details and global context, is of little discussion, despite its frequent involvement in human reasoning. To bridge the gap, we introduce a challenging dataset, namely GeoExplain, which evaluates explainable geo-localization. Given a street view image, the task is to predict its location and provide a detailed explanation. GeoExplain consists of 40350 panoramas-location-explanation tuples. Each instance contains a set of street-view panoramas, a location on street level, and human-expert explanations describing how the location can be inferred from the visual content of panoramas. Additionally, we present a multimodal and multilevel reasoning method, namely SightSense which can make predictions and generate a comprehensive explanation. Our analysis and experiments demonstrate its outstanding performance in GeoExplain.

cs.CL

ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving

This paper presents ARTEMIS, an end-to-end autonomous driving framework that combines autoregressive trajectory planning with Mixture-of-Experts (MoE). Traditional modular methods suffer from error propagation, while existing end-to-end models typically employ static one-shot inference paradigms that inadequately capture the dynamic changes of the environment. ARTEMIS takes a different method by generating trajectory waypoints sequentially, preserves critical temporal dependencies while dynamically routing scene-specific queries to specialized expert networks. It effectively relieves trajectory quality degradation issues encountered when guidance information is ambiguous, and overcomes the inherent representational limitations of singular network architectures when processing diverse driving scenarios. Additionally, we use a lightweight batch reallocation strategy that significantly improves the training speed of the Mixture-of-Experts model. Through experiments on the NAVSIM dataset, ARTEMIS exhibits superior competitive performance, achieving 87.0 PDMS and 83.1 EPDMS with ResNet-34 backbone, demonstrates state-of-the-art performance on multiple metrics.

cs.RO

CHARMS: A Cognitive Hierarchical Agent for Reasoning and Motion Stylization in Autonomous Driving

To address the challenge of insufficient interactivity and behavioral diversity in autonomous driving decision-making, this paper proposes a Cognitive Hierarchical Agent for Reasoning and Motion Stylization (CHARMS). By leveraging Level-k game theory, CHARMS captures human-like reasoning patterns through a two-stage training pipeline comprising reinforcement learning pretraining and supervised fine-tuning. This enables the resulting models to exhibit diverse and human-like behaviors, enhancing their decision-making capacity and interaction fidelity in complex traffic environments. Building upon this capability, we further develop a scenario generation framework that utilizes the Poisson cognitive hierarchy theory to control the distribution of vehicles with different driving styles through Poisson and binomial sampling. Experimental results demonstrate that CHARMS is capable of both making intelligent driving decisions as an ego vehicle and generating diverse, realistic driving scenarios as environment vehicles. The code for CHARMS is released at https://github.com/chuduanfeng/CHARMS.

cs.RO

EPN: An Ego Vehicle Planning-Informed Network for Target Trajectory Prediction

Trajectory prediction plays a crucial role in improving the safety of autonomous vehicles. However, due to the highly dynamic and multimodal nature of the task, accurately predicting the future trajectory of a target vehicle remains a significant challenge. To address this challenge, we propose an Ego vehicle Planning-informed Network (EPN) for multimodal trajectory prediction. In real-world driving, the future trajectory of a vehicle is influenced not only by its own historical trajectory, but also by the behavior of other vehicles. So, we incorporate the future planned trajectory of the ego vehicle as an additional input to simulate the mutual influence between vehicles. Furthermore, to tackle the challenges of intention ambiguity and large prediction errors often encountered in methods based on driving intentions, we propose an endpoint prediction module for the target vehicle. This module predicts the target vehicle endpoints, refines them using a correction mechanism, and generates a multimodal predicted trajectory. Experimental results demonstrate that EPN achieves an average reduction of 34.9%, 30.7%, and 30.4% in RMSE, ADE, and FDE on the NGSIM dataset, and an average reduction of 64.6%, 64.5%, and 64.3% in RMSE, ADE, and FDE on the HighD dataset. The code will be open sourced after the letter is accepted.

cs.RO

MambaBEV: An EV-based 3D detection model with Mamba2

Accurate 3D object detection in autonomous driving relies on Bird's Eye View (BEV) perception and effective temporal fusion. However, existing fusion strategies based on convolutional layers or deformable self-attention struggle to model global context in BEV space, leading to reduced accuracy for large objects.To address this limitation, we propose MambaBEV, a novel BEV-based 3D object detection model that leverages Mamba2, an advanced state-space model (SSM) optimized for long-sequence processing. Our key contribution is TemporalMamba, a temporal fusion module that enhances global context modeling through a BEV feature discrete rearrangement mechanism tailored for sequential processing. In addition, we introduce a Mamba-based DETR head to improve multi-object representation. Evaluations on the nuScenes dataset demonstrate that MambaBEV-base achieves 51.7% NDS and an 42.7% mAP. Furthermore, evaluation within an end-to-end autonomous driving paradigm validates its effectiveness in motion forecasting and planning.These results highlight the potential of state-space models for improving global context understanding and large-object detection in autonomous driving perception systems.

cs.CV

SigDLA: A Deep Learning Accelerator Extension for Signal Processing

Deep learning and signal processing are closely correlated in many IoT scenarios such as anomaly detection to empower intelligence of things. Many IoT processors utilize digital signal processors (DSPs) for signal processing and build deep learning frameworks on this basis. While deep learning is usually much more computing-intensive than signal processing, the computing efficiency of deep learning on DSPs is limited due to the lack of native hardware support. In this case, we present a contrary strategy and propose to enable signal processing on top of a classical deep learning accelerator (DLA). With the observation that irregular data patterns such as butterfly operations in FFT are the major barrier that hinders the deployment of signal processing on DLAs, we propose a programmable data shuffling fabric and have it inserted between the input buffer and computing array of DLAs such that the irregular data is reorganized and the processing is converted to be regular. With the online data shuffling, the proposed architecture, SigDLA, can adapt to various signal processing tasks without affecting the deep learning processing. Moreover, we build a reconfigurable computing array to suit the various data width requirements of both signal processing and deep learning. According to our experiments, SigDLA achieves an average performance speedup of 4.4$\times$, 1.4$\times$, and 1.52$\times$, and average energy reduction of 4.82$\times$, 3.27$\times$, and 2.15$\times$ compared to an embedded ARM processor with customized DSP instructions, a DSP processor, and an independent DSP-DLA architecture respectively with 17% more chip area over the original DLAs.

cs.AR

Enhancing High-Speed Cruising Performance of Autonomous Vehicles through Integrated Deep Reinforcement Learning Framework

High-speed cruising scenarios with mixed traffic greatly challenge the road safety of autonomous vehicles (AVs). Unlike existing works that only look at fundamental modules in isolation, this work enhances AV safety in mixed-traffic high-speed cruising scenarios by proposing an integrated framework that synthesizes three fundamental modules, i.e., behavioral decision-making, path-planning, and motion-control modules. Considering that the integrated framework would increase the system complexity, a bootstrapped deep Q-Network (DQN) is employed to enhance the deep exploration of the reinforcement learning method and achieve adaptive decision making of AVs. Moreover, to make AV behavior understandable by surrounding HDVs to prevent unexpected operations caused by misinterpretations, we derive an inverse reinforcement learning (IRL) approach to learn the reward function of skilled drivers for the path planning of lane-changing maneuvers. Such a design enables AVs to achieve a human-like tradeoff between multi-performance requirements. Simulations demonstrate that the proposed integrated framework can guide AVs to take safe actions while guaranteeing high-speed cruising performance.

eess.SY

Monotone iterative technique for nonlinear fourth order integro-differential equations

In this paper, we consider the solvability of a class of nonlinear fourth order integro-differential equations with Navier boundary condition. We first deal with a corresponding linear problem and establish a maximum principle. Using the maximum principle, we develop a monotone iterative technique in the presence of lower and upper solutions to solve the nonlinear problem under certain conditions. Some examples are presented to illustrate the main results.

math.CA

Existence and uniqueness of positive solutions for Kirchhoff type beam equations

This paper is concerned with the existence and uniqueness of positive solution for the fourth order Kirchhoff type problem $$\left\{\begin{array}{ll} u''''(x)-(a+b\int_0^1(u'(x))^2dx)u''(x)=λf(u(x)),\ \ \ \ x\in(0,1),\\ u(0)=u(1)=u''(0)=u''(1)=0,\\ \end{array} \right. $$ where $a>0, b\geq 0$ are constants, $λ\in \mathbb{R}$ is a parameter. For the case $f(u)\equiv u$, we use an argument based on the linear eigenvalue problems of fourth order equations and their properties to show that there exists a unique positive solution for all $λ>λ_{1,a}$, here $λ_{1,a}$ is the first eigenvalue of the above problem with $b=0$; For the case $f$ is sublinear, we prove that there exists a unique positive solution for all $λ>0$ and no positive solution for $λ<0$ by using bifurcation method.

math.CA

Heavy Ion Induced SEU Sensitivity Evaluation of 3D Integrated SRAMs

Heavy ions induced single event upset (SEU) sensitivity of three-dimensional integrated SRAMs are evaluated by using Monte Carlo sumulation methods based on Geant4. The cross sections of SEUs and Multi Cell Upsets (MCUs) for 3D SRAM are simulated by using heavy ions with different energies and LETs. The results show that the sensitivity of different die of 3D SRAM has obvious discrepancies at low LET. Average percentage of MCUs of 3D SRAMs rises from 17.2% to 32.95% when LET increases from 42.19 MeV cm2/mg to 58.57MeV cm2/mg. As for a certain LET, the percentage of MCUs shows a notable distinction between face-to-face structure and back-to-face structure. For back-to-face structure, the percentage of MCUs increases with the deeper die. However, the face-to-face die presents the relatively low percentage of MCUs. The comparison of SEU cross sections for planar SRAMs and experiment data are conducted to indicate the effectiveness of our simulation method. Finally, we compare the upset cross sections of planar process and 3D integrated SRAMs. Results demonstrate that the sensitivity of 3D SRAMs is not more than that of planar SRAMs and the 3D structure can be become a great potential application for aerospace and military domain.

physics.ins-det