SearcharxivSearch

arXiv subjects

Chen Song

Publications and source records attributed to Chen Song.

At least 19 recordsLinked to original sources

Z-Fold Wing Aeroelasticity: Compositional Modeling, 1:2 Double-Hopf Dynamics, and Nonlinear Stiffness Effects

Folding changes the linearized aeroelastic spectrum and can switch which mode becomes unstable first, with consequences for local postflutter interactions. This study formulates a three-component Z-fold wing by assigning every aerodynamic station to a structural component and a material coordinate. The same attachment map generates surface motion and returns pressure loads through virtual work. Geometrically exact component dynamics, an explicit-wake unsteady vortex-lattice model, and block-structured descriptor assembly preserve the physical paths of configuration actions. A two-parameter flutter analysis shows that a smooth flutter-speed envelope conceals a high-low-high sequence of controlling neutral branches, expressed as a critical-frequency valley and a redistribution of component deformation. Numerical continuation locates a near-1:2 double-Hopf point. Within a local model retaining quadratic and cubic structural restoring forces with aerodynamic and inertial operators fixed at the scheduling point, the cubic normal form captures selected 26-state observations and admits high-frequency-dominant and mixed phase-locked periodic solutions. The formulation links configuration-dependent flutter-mode identity to local resonant dynamics and supports blockwise sensitivity and design reasoning.

physics.flu-dyn

Compositional Aeroelastic Operators for Morphing Flexible Multibody Aircraft: A Geometric Framework with Structural Verification

Morphing flexible multibody aircraft require structural strain, aerodynamic geometry, surface velocity, and generalized loading to remain compatible as joints and flexible components change configuration. A compositional formulation is developed around an assumed material attachment between each lifting surface and a geometrically exact beam. Separating the component root pose from the section field shows that the body strain and elastic potential of a component depend on its own elastic coordinates, while upstream motion enters kinetic terms and external-load pullbacks. At element level, an exact relative logarithm $d$ supplies strain and potential energy, whereas a reference-anchored section coordinate $\sigma$ supplies deformed section geometry. Finite-order expansions retain the finite reference geometry exactly and truncate only endpoint perturbations. The attachment map then generates surface points, tangents, normals, velocities, and force Jacobians from common section kinematics. Euler--Poincare beam balance, moving-surface potential-flow relations, graph cotangent assembly, and the associated semidiscrete power identity are stated in a common twist--wrench convention. Collocation, pressure, equivalent-load, and structural-station sites are distinguished to expose their approximation errors. Verification gives the expected $N_d+1$ convergence order for degree-$N_d$ relative-log expansions. In a geometrically nonlinear cantilever comparison, a cubic static-manifold correction reduces mean full-record displacement error from $0.479$ to $0.255$ over four completed load cases. These results provide structural and interface-level evidence rather than validation of a complete aircraft aeroelastic prediction.

cs.CE

Scene Reconstruction as Mapping Priors for 3D Detection

In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust structural priors of the static environment, helping resolve ambiguities and correct for sensor data sparsity or noise, especially for distant objects or under adverse weather conditions. However, conventional High-Definition (HD) maps are resource-intensive to obtain and maintain, which presents a challenge for efficient, large-scale deployment. In this paper, we propose a scalable solution to systematically leverage mapping to improve 3D detection by overcoming two primary challenges. First, we introduce a pipeline to automatically build dense mapping priors from aggregated sensor data, eliminating the need for human labeling. Second, we design a novel Mapping Priors Augmented 3D Detection (MPA3D) framework to effectively integrate mapping priors with different sensor modalities. Extensive experiments on the Waymo Open Dataset demonstrate that our approach achieves new state-of-the-art results, proving the effectiveness of scalable reconstructed scene priors for enhancing 3D detection.

cs.CV

STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

Model scaling has demonstrated remarkable success through large-scale training on diverse datasets. It remains an open question whether the same paradigm would apply to autonomous driving perception systems due to unique challenges, such as fusing heterogeneous sensor data and the need for sophisticated 3D spatial understanding. To bridge this gap, we present a comprehensive study on systematically analyzing the impact of scale on these systems. We develop our STELLAR model based on Sparse Window Transformer, by extending the input modalities to include LiDAR, radar, camera, and map prior. We train the model on a large-scale dataset of 50 million driving examples with up to 500 million parameters. Our large-scale experiments reveal empirical scaling trends that connect model performance to model size, data, and compute. The resulting model establishes a new state-of-the-art on the Waymo Open Dataset challenge, outperforming prior arts by a large margin. Our work demonstrates that large-scale training is a highly promising path for advancing the capabilities of perception models for autonomous driving.

cs.CV

Impact of Attitude and Bounded Rationality on Collective Behavioral Transitions

The theory of planned behavior (TPB) is one of the most influential frameworks in social psychology, stating that a person's behavior is driven by intention, which is primarily shaped by attitude, subjective norms, and perceived behavioral control. Despite its strong empirical support, TPB remains a static conceptual framework without explicit mathematical formulations that capture the temporal evolution of its components. To address this gap, we develop a dynamic agent-based modeling framework that integrates the core principles of TPB with a behavior-to-attitude feedback mechanism. Specifically, we define behaviors based on their feedback effects on attitude and examine when the population undergoes collective transitions by either adopting a beneficial behavior or rejecting a harmful one. Results from our model demonstrate that collective transitions can be effectively controlled by adjusting two key behavioral parameters that reflect agents' attitude influence and decision rationality. These findings provide quantitative insights on TPB, highlighting the key factors that drive collective behavioral transitions and the need for further socio-psychological case studies.

cs.SI

RG-Based Local Hopf Reduction and Slow-Manifold Reconstruction for Nonlinear Aeroelastic Systems

Self-excited limit-cycle oscillations (LCOs) from Hopf bifurcations are a key feature of nonlinear aeroelasticity and depend sensitively on structural and aerodynamic parameters. Classical center-manifold and normal-form theory describe this local behavior, but can be cumbersome to apply in large discretized models and standard reduced-order modeling (ROM) workflows. A renormalization-group (RG)-based reduction is developed that directly yields a Hopf-type amplitude equation on a local invariant manifold, specialized for polynomial nonlinearities in tensor-based discretizations and compatible with finite-element-type settings. The method provides explicit coefficients governing the Hopf threshold, criticality, and leading LCO amplitude/frequency trends, and admits a companion slow-manifold approximation with selected stable modes retained as static coordinates. Representative nonlinear-aeroelastic examples illustrate how the proposed framework supplies compact, parameter-aware Hopf/LCO descriptors suitable for local ROM construction near flutter.

physics.flu-dyn

On the Convergence of an Opinion-Action Coevolution Model with Bounded Confidence

This paper presents a theoretical convergence analysis for an opinion-action coevolution model that integrates the opinion updating rule of the Hegselmann-Krause model with a utility-based decision-making mechanism. The model is reformulated into an augmented state-space representation, where the state matrix induces a time-varying social interaction digraph. The convergence analysis is grounded on two existing theoretical findings that establish convergence for the Hegselmann-Krause type of models and containment control systems with multiple stationary leaders, respectively. Results indicate that, if the structure of the interaction digraph stabilizes within finite time, the model either converges to consensus, where all agents' opinions and actions reach an identical state, or exhibits clustering, where some opinion nodes act as stationary leaders while the remaining nodes approach the convex hull formed by the leaders. Numerical simulations are then provided to validate the theoretical results.

eess.SY

HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical bottleneck, suffering from excessive latency and memory consumption when processing high-resolution inputs.To address these challenges, we introduce HyperVL, an efficient multimodal large language model tailored for on-device inference. HyperVL adopts an image-tiling strategy to cap peak memory usage and incorporates two novel techniques: (1) a Visual Resolution Compressor (VRC) that adaptively predicts optimal encoding resolutions to eliminate redundant computation, and (2) Dual Consistency Learning (DCL), which aligns multi-scale ViT encoders within a unified framework, enabling dynamic switching between visual branches under a shared LLM. Extensive experiments demonstrate that HyperVL achieves state-of-the-art performance among models of comparable size across multiple benchmarks. Furthermore, it significantly significantly reduces latency and power consumption on real mobile devices, demonstrating its practicality for on-device multimodal inference.

cs.CV

GPx4 is bound to peroxidized membranes by a hydrophobic anchor

Ferroptosis is a form of cell death discovered in recent years, induced by excessive peroxidation of phospholipids. Glutathione peroxidase 4 (GPx4) is an intracellular enzyme that can repair the peroxidized phospholipids on membranes, thus regulating ferroptosis. By combining multiscale molecular dynamics (MD) simulations and experimental assays, we investigate the binding mechanisms of GPx4 on membranes. Using coarse-grained MD simulations, we found that L130 and its adjacent residues on GPx4 can form a stable and unique binding interface with PE/PS-rich and peroxidized membranes. Subsequent all-atom MD simulations verified the stability of the binding interface. The critical residue on the interface, L130, was inserted deeply into the membrane as a hydrophobic anchor and guided the reaction center toward the membrane surface. Enzyme activity assays and in vitro cell experiments showed that mutations of L130 resulted in weaker activities of the enzyme, probably caused by non-functional binding modes of GPx4 on membranes, as revealed by in silico simulations. This study highlights the crucial role of the hydrophobic residue, L130, in the proper anchoring of GPx4 on membranes, the first step of its membrane-repairing function.

q-bio.BM

Why do Opinions and Actions Diverge? A Dynamic Framework to Explore the Impact of Subjective Norms

Socio-psychological studies have identified a common phenomenon where an individual's public actions do not necessarily coincide with their private opinions, yet most existing models fail to capture the dynamic interplay between these two aspects. To bridge this gap, we propose a novel agent-based modeling framework that integrates opinion dynamics with a decision-making mechanism. More precisely, our framework generalizes the classical Hegselmann-Krause model by combining it with a utility maximization problem. Preliminary results from our model demonstrate that the degree of opinion-action divergence within a population can be effectively controlled by adjusting two key parameters that reflect agents' personality traits, while the presence of social network amplifies the divergence. In addition, we study the social diffusion process by introducing a small number of committed agents into the model, and identify three key outcomes: adoption of innovation, rejection of innovation, and the enforcement of unpopular norms, consistent with findings in socio-psychological literature. The strong relevance of the results to real-world phenomena highlights our framework's potential for future applications in understanding and predicting complex social behaviors.

cs.SI

PPLNs: Parametric Piecewise Linear Networks for Event-Based Temporal Modeling and Beyond

We present Parametric Piecewise Linear Networks (PPLNs) for temporal vision inference. Motivated by the neuromorphic principles that regulate biological neural behaviors, PPLNs are ideal for processing data captured by event cameras, which are built to simulate neural activities in the human retina. We discuss how to represent the membrane potential of an artificial neuron by a parametric piecewise linear function with learnable coefficients. This design echoes the idea of building deep models from learnable parametric functions recently popularized by Kolmogorov-Arnold Networks (KANs). Experiments demonstrate the state-of-the-art performance of PPLNs in event-based and image-based vision applications, including steering prediction, human pose estimation, and motion deblurring. The source code of our implementation is available at https://github.com/chensong1995/PPLN.

cs.CV

The Syzygy Matrix and the Differential for Rational Curves in Projective Space

In this paper, we study whether a given morphism $f$ from the tangent bundle of $\mathbb{P}^1$ to a balanced vector bundle of degree $(n+1)d$ is induced by the restriction of the tangent bundle $T_{\mathbb{P}^n}$ to a rational curve of degree $d$ in $\mathbb{P}^n$. We propose a conjecture on this problem based on Mathematica computations of some examples and provide computer-assisted proof of the conjecture for certain values of $n$ and $d$.

math.AG

Machine Learning with Physics Knowledge for Prediction: A Survey

This survey examines the broad suite of methods and models for combining machine learning with physics knowledge for prediction and forecast, with a focus on partial differential equations. These methods have attracted significant interest due to their potential impact on advancing scientific research and industrial practices by improving predictive models with small- or large-scale datasets and expressive predictive models with useful inductive biases. The survey has two parts. The first considers incorporating physics knowledge on an architectural level through objective functions, structured predictive models, and data augmentation. The second considers data as physics knowledge, which motivates looking at multi-task, meta, and contextual learning as an alternative approach to incorporating physics knowledge in a data-driven fashion. Finally, we also provide an industrial perspective on the application of these methods and a survey of the open-source ecosystem for physics-informed machine learning.

cs.LG

TutteNet: Injective 3D Deformations by Composition of 2D Mesh Deformations

This work proposes a novel representation of injective deformations of 3D space, which overcomes existing limitations of injective methods: inaccuracy, lack of robustness, and incompatibility with general learning and optimization frameworks. The core idea is to reduce the problem to a deep composition of multiple 2D mesh-based piecewise-linear maps. Namely, we build differentiable layers that produce mesh deformations through Tutte's embedding (guaranteed to be injective in 2D), and compose these layers over different planes to create complex 3D injective deformations of the 3D volume. We show our method provides the ability to efficiently and accurately optimize and learn complex deformations, outperforming other injective approaches. As a main application, we produce complex and artifact-free NeRF and SDF deformations.

cs.CV

An Optimization Framework to Enforce Multi-View Consistency for Texturing 3D Meshes

A fundamental problem in the texturing of 3D meshes using pre-trained text-to-image models is to ensure multi-view consistency. State-of-the-art approaches typically use diffusion models to aggregate multi-view inputs, where common issues are the blurriness caused by the averaging operation in the aggregation step or inconsistencies in local features. This paper introduces an optimization framework that proceeds in four stages to achieve multi-view consistency. Specifically, the first stage generates an over-complete set of 2D textures from a predefined set of viewpoints using an MV-consistent diffusion process. The second stage selects a subset of views that are mutually consistent while covering the underlying 3D model. We show how to achieve this goal by solving semi-definite programs. The third stage performs non-rigid alignment to align the selected views across overlapping regions. The fourth stage solves an MRF problem to associate each mesh face with a selected view. In particular, the third and fourth stages are iterated, with the cuts obtained in the fourth stage encouraging non-rigid alignment in the third stage to focus on regions close to the cuts. Experimental results show that our approach significantly outperforms baseline approaches both qualitatively and quantitatively. Project page: https://aigc3d.github.io/ConsistenTex.

cs.CV

Stability of Kernel Bundles

In this paper, we study the stability of general kernel bundles on $\mathbb{P}^n$. Let $a,b,d>0$ be integers. A kernel bundle $E_{a,b}$ on $\mathbb{P}^n$ is defined as the kernel of a surjective map $\phi:\mathcal{O}_{\mathbb{P}^n}(-d)^a\rightarrow \mathcal{O}_{\mathbb{P}^n}^b$. Here $\phi$ is represented by a $b\times a$ matrix $(f_{ij})$ where the entries $f_{ij}$ are polynomials of degree $d$. We give sufficient conditions for semistability of a general kernel bundle on $\mathbb{P}^n$, in terms of its Chern class.

math.AG

Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting $ $ triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as \emph{\textbf{UniHOI}}. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (\emph{i.e.} GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights are available at: \url{https://github.com/Caoyichao/UniHOI}.

cs.CV

Multi-View Representation is What You Need for Point-Cloud Pre-Training

A promising direction for pre-training 3D point clouds is to leverage the massive amount of data in 2D, whereas the domain gap between 2D and 3D creates a fundamental challenge. This paper proposes a novel approach to point-cloud pre-training that learns 3D representations by leveraging pre-trained 2D networks. Different from the popular practice of predicting 2D features first and then obtaining 3D features through dimensionality lifting, our approach directly uses a 3D network for feature extraction. We train the 3D feature extraction network with the help of the novel 2D knowledge transfer loss, which enforces the 2D projections of the 3D feature to be consistent with the output of pre-trained 2D networks. To prevent the feature from discarding 3D signals, we introduce the multi-view consistency loss that additionally encourages the projected 2D feature representations to capture pixel-wise correspondences across different views. Such correspondences induce 3D geometry and effectively retain 3D features in the projected 2D features. Experimental results demonstrate that our pre-trained model can be successfully transferred to various downstream tasks, including 3D shape classification, part segmentation, 3D object detection, and semantic segmentation, achieving state-of-the-art performance.

cs.CV