SearcharxivSearch

arXiv subjects

Qifeng Li

Publications and source records attributed to Qifeng Li.

At least 19 recordsLinked to original sources

Boundary-Enhanced Segmentation of Pig Point Clouds in Commercial Housing Environments

In real pigsty environments, pig point clouds often come into close contact with background structures, resulting in blurred target boundaries, local adhesion, and background mis-segmentation. This reduces the accuracy of subsequent point cloud completion and body size measurement. To address these challenges, this study proposes a pig point cloud segmentation method based on boundary feature analysis. The proposed method adopts Octree Transformer as the backbone network and integrates local geometric details with global semantic context through octree convolution, self-attention encoding, and multi-scale feature fusion. Furthermore, soft-distance boundary pseudo-labels are generated to provide continuous boundary supervision, and a bidirectional cross-boundary semantic module is designed to enable explicit interaction between boundary and semantic features. Experiments conducted on a comprehensive dataset demonstrate that the proposed method significantly outperforms various state-of-the-art models in terms of segmentation accuracy, mean intersection over union, and boundary delineation. The results indicate that the method effectively alleviates boundary adhesion, providing reliable point cloud inputs for downstream precision livestock farming tasks.

cs.CV

Fundamental forms and infinitesimal symmetries of projective varieties

We give a bound on the dimension of the linear automorphism group of a projective variety $Z \subset \mathbb{P} V$ in terms of its fundamental forms at a general point. Moreover, we show that the bound is achieved precisely when $Z \subset \mathbb{P} V$ is projectively equivalent to an Euler-symmetric variety. As a by-product, we determine the Lie algebra of infinitesimal automorphisms of an Euler-symmetric variety and also obtain a rigidity result on the specialization of an Euler-symmetric variety preserving the isomorphism type of the fundamental forms.

math.AG

Homological rigidity and Schur rigidity of Schubert varieties in rational homogeneous spaces

A Schubert variety $X_0$ on a rational homogenous space $X=G/P$ is said to be homologically rigid, if any subvariety $Z$ on $X$ representing the same homology class with $X_0$ must satisfy $Z=g\cdot X_0$ for some $g\in{\rm Aut_0}(X)$. We say $X_0$ is Schur rigid, if furthermore any subvariety $Z$ on $X$ whose homology class is a multiple $r$ of that of $X_0$ must satisfy $Z=g_1\cdot X_0+\cdots+g_r\cdot X_0$ for some $g_1,\cdots ,g_r\in{\rm Aut_0}(X)$. Homological rigidity and Schur rigidity of Schubert varieties in rational homogeneous spaces of Picard number one have been well studied in extensive literature. In this paper, we study both rigidity problems of Schubert varieties in rational homogeneous spaces of higher Picard numbers. We show that in the long root cases, including all cases when $G$ is of type $ADE$, smooth Schubert varieties have homological rigidity. Besides, we give the complete list of Schubert varieties of subdiagram type with/without homological rigidity. Furthermore, for a Schubert variety $X_0$ of subdiagram type, we show that it has Schur rigidity in long root cases unless $X_0$ admits a fiber bundle structure over the projective space.

math.AG

ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving

Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log. However, existing behavior simulation benchmarks do not directly measure reactive capability. They often let the simulator jointly control the AV and surrounding agents and evaluate realism through log similarity or open-loop prediction metrics. In this work, we introduce ReactSim-Bench for evaluating the reactive capability of behavior world model simulation in autonomous driving. We decouple the control of agents and the AV, using AV behaviors that differ from the log and require agents to respond as independent AV inputs. To obtain these AV behaviors, we construct a pipeline that uses an AV planner model to generate candidate behaviors and filters the data using rules and manual verification. Collision metrics, map-based metrics, and kinematic feasibility metrics are used to evaluate the safety and rule compliance of reactive responses. We construct 2,636 test scenarios with three categories and conduct a systematic evaluation of state-of-the-art models across multiple architectures, including Transformer-based, diffusion-based, and next-token-prediction-based models. We further analyze how replan frequency affects performance and provide insights for future studies.

cs.RO

Effect of startup modes on cold start performance of PEM fuel cells with different cathode flow fields

Proton Exchange Membrane Fuel Cell (PEMFC) is widely recognized for its cleanliness and high efficiency, but is still facing challenges in cold environments. At low temperatures, the formation of ice and repeated freezing/thawing cycles may cause cell performance reduction and irreversible degradation. The cathode flow field of PEMFCs has a significant effect on the performance. In contrast to the conventional ``channel-ridge'' flow field, the metal foam has the advantages of excellent pre-distribution of gases and water drainage, which make it a promising candidate for the cold start. This paper examines the cold start of PEMFCs with metal foam flow field (MFFF) and serpentine flow field (SFF), and the influence of constant current mode, constant voltage mode, and ramping current mode is investigated experimentally through performance test and electrochemical characterization. The results show that lowering the voltage and increasing the current can enhance the cold-start performance of fuel cells. The MFFF fuel cell has superior cold start performance compared to the SFF fuel cell under the constant voltage mode of 0.3 V. Furthermore, the variable current mode is developed by considering the distinct properties of heat and water production during various phases, and the results indicate that increasing the current density at the unsaturated stage leads to an elevated rate of heat production and a reduced rate of water production, which can improve the cold start of PEMFCs.

physics.flu-dyn

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn task-relevant features. However, without explicit guidance, these models often overfit to spurious correlations, such as visual shortcuts or environmental noise, limiting their generalization. In this paper, we introduce GuidedVLA, a framework designed to manually guide the action generation to focus on task-relevant factors. Our core insight is to treat the action decoder not as a monolithic learner, but as an assembly of functional components. Individual attention heads are supervised by manually defined auxiliary signals to capture distinct factors. As an initial study, we instantiate this paradigm with three specialized heads: object grounding, spatial geometry, and temporal skill logic. Across simulation and real-robot experiments, GuidedVLA improves success rates in both in-domain and out-of-domain settings compared to strong VLA baselines. Finally, we show that the quality of these specialized factors correlates positively with task performance and that our mechanism yields decoupled, high-quality features. Our results suggest that explicitly guiding action-decoder learning is a promising direction for building more robust and general VLA models.

cs.RO

A Data-embedded Solution Paradigm for Nonconvex Probable Event Constrained Optimization

This paper introduces a new modeling framework for optimization under uncertainty, called Probable Event Constrained Optimization (PECO). Unlike conventional chance-constrained formulations, which only limit the probability of constraint violation, PECO also explicitly requires feasibility for all events whose probability exceeds a prescribed threshold. This guarantees that solutions remain valid across all high-probability realizations of uncertainty. To solve PECO, we proposed a data-embedded program (DEP) which directly incorporates historical measurements of the uncertain parameters to obtain a deterministic approximation for PECO. While existing solution methods for optimization problems under uncertainty rely heavily on convexity or linearity assumptions, the proposed data-embedded solution paradigm provides a unique opportunity for solving nonlinear and nonconvex PECOs. The effectiveness of this approach depends on properly estimating the number of elements in the family of solution-determining data sets. As we enter the era of big data, this information can be properly estimated by leveraging the power of machine learning.

math.OC

PROMETHEE-based Modeling of Endogenous Behavioral Uncertainty of EV Owners

The electric vehicle (EV) charging demands (CD) are jointly determined by the EV owners' behavior (i.e., human factor) and the electricity prices (i.e., decisions of distribution system operators (DSO)). However, most existing studies either neglect the decision-dependent nature of EVCD uncertainty or idealistically treat EV owners as perfect decision-makers. This paper formulates the optimal operation of power distribution systems (PDS) as a distributionally robust chance-constrained (DRCC) problem considering EVCDs as endogenous uncertainty (i.e., decision-dependent uncertainty). The Preference Ranking Organization Method for Enrichment Evaluation (PROMETHEE) is introduced to capture the human factor of EV owners in the proposed ambiguity set. Case studies on IEEE test systems demonstrate that the proposed method achieves superior performance compared to deterministic and conventional DRCC approaches, thereby enhancing resilience and security in PDS operations.

eess.SY

Effects of gas diffusion layer thickness on PEM fuel cells with composite foam-rib flow fields

Gas diffusion layers (GDLs) play a crucial role for the performance of proton exchange membrane fuel cells (PEMFCs). The utilization of composite foam-rib flow fields (CFRFFs) can alter the reactant gas transfer pattern, hence improving the efficiency of under-rib reactant gas transfer and water drainage. The impact of the cathode and anode GDL thicknesses (h_{c,GDL} and h_{a,GDL}) on the performance of CFRFF design is investigated by three-dimensional multiphase non-isothermal numerical simulation in this study. The results indicate that for the conventional rib flow field (CRFF) design, there is an optimal h_{c,GDL} for optimal cell performance, while for the CFRFF design, as h_{c,GDL} becomes thinner, the cell performance increases, and the trend is dominated by the variation of the oxygen concentration. Under a thin GDL, the rib width of the CRFF design should be as small as possible to minimize concentration polarization loss, while the rib width of the CFRFF design can be slightly larger. Furthermore, by decreasing the thickness of h_{a,GDL} in both the CRFF and CFRFF designs, there is an increase in the dissolved water content in the ionomer of the cathode CL and a subsequent decrease in the Ohmic polarization loss.

physics.flu-dyn

Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models

With the rise of vision-language models (VLM), their application for autonomous driving (VLM4AD) has gained significant attention. Meanwhile, in autonomous driving, closed-loop evaluation has become widely recognized as a more reliable validation method than open-loop evaluation, as it can evaluate the performance of the model under cumulative errors and out-of-distribution inputs. However, existing VLM4AD benchmarks evaluate the model`s scene understanding ability under open-loop, i.e., via static question-answer (QA) dataset. This kind of evaluation fails to assess the VLMs performance under out-of-distribution states rarely appeared in the human collected datasets.To this end, we present Bench2Drive-VL, an extension of Bench2Drive that brings closed-loop evaluation to VLM-based driving, which introduces: (1) DriveCommenter, a closed-loop generator that automatically generates diverse, behavior-grounded question-answer pairs for all driving situations in CARLA,including severe off-route and off-road deviations previously unassessable in simulation. (2) A unified protocol and interface that allows modern VLMs to be directly plugged into the Bench2Drive closed-loop environment to compare with traditional agents. (3) A flexible reasoning and control framework, supporting multi-format visual inputs and configurable graph-based chain-of-thought execution. (4) A complete development ecosystem. Together, these components form a comprehensive closed-loop benchmark for VLM4AD. All codes and annotated datasets are open sourced.

cs.RO

Minimal rational curves on equivariant compactifications of symmetric spaces

Let $G/H$ be a symmetric space of a complex linear algebraic group $G$ and let $X$ be a nonsingular equivariant compactification of $G/H$. We investigate the question: when are minimal rational curves on $X$ orbit-closures of 1-parameter subgroups of $G$? We show that this is the case if the variety of minimal rational tangents (VMRT) at a base point in $G/H \subset X$ is Gauss-nondegenerate. Our method combines algebraic geometry of minimal rational curves with differential geometry of symmetric spaces: orbits of 1-parameter subgroups arise as holomorphic geodesics of an invariant torsion-free affine connection on $G/H$. We prove furthermore that the Gauss-nondegeneracy of VMRT holds for nonsingular equivariant compactifications of simple algebraic groups regarded as symmetric spaces. In this case, we also show that the VMRT is the closure of an adjoint orbit, which generalizes a result of Brion and Fu's on wonderful compactifications to arbitrary equivariant compactifications.

math.AG

Mass transfer and water management in proton exchange membrane fuel cells with a composite foam-rib flow field

Mass transfer capability of reactants and hydrothermal management is important for the performance and durability of proton exchange membrane fuel cells. In the conventional rib flow field, the oxygen transport is affected by the accumulation of under-rib liquid water which causes excessive concentration loss and limits cell performance. To improve the cell performance, a composite foam-rib flow field structure is proposed by combining the metal foam flow field and the conventional rib flow field. The proposed design is simulated by using a three-dimensional homogeneous non-isothermal numerical model. The results show that the composite foam-rib flow field, by improving the oxygen transfer and water removal capabilities under the ribs, can improve the oxygen concentration and current density without increasing the pumping power, thus improving the cell performance under different conditions. The key parameters of the composite foam-rib flow field are optimized. With the optimal metal foam filling ratio of 0.75 and porosity of 0.85, the peak power density and the limiting current density for the composite foam-rib flow field are higher than the conventional rib flow field by 5.20% and 22.68%.

physics.flu-dyn

TrajTok: Technical Report for 2025 Waymo Open Sim Agents Challenge

In this technical report, we introduce TrajTok, a trajectory tokenizer for discrete next-token-prediction based behavior generation models, which combines data-driven and rule-based methods with better coverage, symmetry and robustness, along with a spatial-aware label smoothing method for cross-entropy loss. We adopt the tokenizer and loss for the SMART model and reach a superior performance with realism score of 0.7852 on the Waymo Open Sim Agents Challenge 2025. We will open-source the code in the future.

cs.CL

DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving

End-to-end autonomous driving (E2E-AD) demands effective processing of multi-view sensory data and robust handling of diverse and complex driving scenarios, particularly rare maneuvers such as aggressive turns. Recent success of Mixture-of-Experts (MoE) architecture in Large Language Models (LLMs) demonstrates that specialization of parameters enables strong scalability. In this work, we propose DriveMoE, a novel MoE-based E2E-AD framework, with a Scene-Specialized Vision MoE and a Skill-Specialized Action MoE. DriveMoE is built upon our $\pi_0$ Vision-Language-Action (VLA) baseline (originally from the embodied AI field), called Drive-$\pi_0$. Specifically, we add Vision MoE to Drive-$\pi_0$ by training a router to select relevant cameras according to the driving context dynamically. This design mirrors human driving cognition, where drivers selectively attend to crucial visual cues rather than exhaustively processing all visual information. In addition, we add Action MoE by training another router to activate specialized expert modules for different driving behaviors. Through explicit behavioral specialization, DriveMoE is able to handle diverse scenarios without suffering from modes averaging like existing models. In Bench2Drive closed-loop evaluation experiments, DriveMoE achieves state-of-the-art (SOTA) performance, demonstrating the effectiveness of combining vision and action MoE in autonomous driving tasks. We will release our code and models of DriveMoE and Drive-$\pi_0$.

cs.CV

Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)

Reinforcement Learning (RL) can mitigate the causal confusion and distribution shift inherent to imitation learning (IL). However, applying RL to end-to-end autonomous driving (E2E-AD) remains an open problem for its training difficulty, and IL is still the mainstream paradigm in both academia and industry. Recently Model-based Reinforcement Learning (MBRL) have demonstrated promising results in neural planning; however, these methods typically require privileged information as input rather than raw sensor data. We fill this gap by designing Raw2Drive, a dual-stream MBRL approach. Initially, we efficiently train an auxiliary privileged world model paired with a neural planner that uses privileged information as input. Subsequently, we introduce a raw sensor world model trained via our proposed Guidance Mechanism, which ensures consistency between the raw sensor world model and the privileged world model during rollouts. Finally, the raw sensor world model combines the prior knowledge embedded in the heads of the privileged world model to effectively guide the training of the raw sensor policy. Raw2Drive is so far the only RL based end-to-end method on CARLA Leaderboard 2.0, and Bench2Drive and it achieves state-of-the-art performance.

cs.RO

Real-time Optimization for Wind-to-H2 Driven Critical Infrastructures Based on Active Constraints Identification and Integer Variables Prediction

This paper proposes a concept of wind-to-hydrogen-driven critical infrastructure (W2H-CI) as an engineering solution for decarbonizing the power generation sector where it utilizes wind power to produce hydrogen through electrolysis and combines it with the carbon captured from fossil fuel power plants. First, a convex mathematical model of W2H-CI is developed. Then, an optimization model for optimal operation of W2H-CI, which is a large-scale mixed-integer convex program (MICP), is proposed. Moreover, we propose to solve this problem in real-time in order to hedge against the uncertainty of wind power. For this purpose, a novel solution method based on active constraints identification and integer variable prediction is introduced. This method can solve MICP problems very fast since it uses historical optimization data to predict the values of binary variables and a limited number of constraints which most likely contain all active constraints. We validate the effectiveness of the proposed fast solution method using two W2H-CI case studies.

eess.SY

Automatically Planning Optimal Parallel Strategy for Large Language Models

The number of parameters in large-scale language models based on transformers is gradually increasing, and the scale of computing clusters is also growing. The technology of quickly mobilizing large amounts of computing resources for parallel computing is becoming increasingly important. In this paper, we propose an automatic parallel algorithm that automatically plans the parallel strategy with maximum throughput based on model and hardware information. By decoupling the training time into computation, communication, and overlap, we established a training duration simulation model. Based on this simulation model, we prune the parallel solution space to shorten the search time required. The multi-node experiment results show that the algorithm can estimate the parallel training duration in real time with an average accuracy of 96%. In our test, the recommendation strategy provided by the algorithm is always globally optimal.

cs.AI

Field-free current-induced magnetization switching of a room temperature van der Waals magnet for neuromorphic computing

Spin orbit torque (SOT) has become a promising approach to efficiently manipulate the magnetization switching in spintronic devices. As a main factor to impact the device performance, the high quality interface is essentially desired, which can be readily acquired by using the two-dimensional (2D) van der Waals (vdW) materials. Recently, a 2D ferromagnetic material Fe3GaTe2 has been discovered to possess the above-room-temperature Curie temperature and strong perpendicular magnetic anisotropy (PMA), providing an excellent candidate to build spintronic devices. On the other hand, an external magnetic field is necessary for the SOT-driven deterministic switching of perpendicular magnetization, which has become a block for the real applications. Here, we realize the field-free SOT switching of Fe3GaTe2 at room temperature based on the Fe3GaTe2/MnPt heterostructure. In addition, inspired by the superiority of 2D materials in 3D heterogeneous integration, we explore the potential of our device in the computing in memory (CIM). With the application of the current pulses, the gradual switching of our device at zero field imitates the function of artificial synapse in the convolutional neural network (CNN), achieving a high accuracy (~92.8%) pattern recognition. Our work proposes a feasible solution for field-free SOT switching in 2D vdW spintronic devices, which paves the way for applications in magnetic memory and neuromorphic computing.

physics.app-ph