SearcharxivSearch

arXiv subjects

Shumin Li

Publications and source records attributed to Shumin Li.

6 recordsLinked to original sources

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning

End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However, they inherit two key limitations from 2D foundation models: 1) the reliance on 2D RGB inputs that ignores the intrinsically 3D nature of manipulation; and 2) the lack of spatial 3D alignment between input-output spaces as well as across diverse robot embodiments, camera setups, and trajectory datasets. In this paper, we present a series of contributions to address these issues. First, we introduce aligned vertex map and vertex spectrum -- a pixel-wise 3D representation that elevates 2D visual inputs to 3D, using camera calibration and optional depth. This novel input representation marries 3D awareness with the generalization of 2D large VLMs. Then, we propose to align the inputs and outputs of manipulation policies by expressing per-pixel 3D information of each camera view and robot actions to a shared coordinate. Based on this, we designate a canonical Bird's-Eye-View (BEV) alignment frame and innovatively propose to construct BEV images, producing a view-invariant representation robust to camera pose variations. To enable training and evaluation at scale, we develop a comprehensive data processing pipeline to perform such alignments; we also introduce a novel temporal alignment scheme for trajectories across diverse robots, human operators, and datasets. These contributions collectively mitigate input and output spatial-temporal misalignments, improving the consistency and generalization for real-world manipulation. Pretrained checkpoint, source code and data processing pipeline are available in https://hnuzhy.github.io/projects/Dex-BEV.

cs.RO

Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset

The development of accessible screening tools for early cancer detection in dogs represents a significant challenge in veterinary medicine. Routine laboratory data offer a promising, low-cost source for such tools, but their utility is hampered by the non-specificity of individual biomarkers and the severe class imbalance inherent in screening populations. This study assesses the feasibility of cancer risk classification using the Golden Retriever Lifetime Study (GRLS) cohort under real-world constraints, including the grouping of diverse cancer types and the inclusion of post-diagnosis samples. A comprehensive benchmark evaluation was conducted, systematically comparing 126 analytical pipelines that comprised various machine learning models, feature selection methods, and data balancing techniques. Data were partitioned at the patient level to prevent leakage. The optimal model, a Logistic Regression classifier with class weighting and recursive feature elimination, demonstrated moderate ranking ability (AUROC = 0.815; 95% CI: 0.793-0.836) but poor clinical classification performance (F1-score = 0.25, Positive Predictive Value = 0.15). While a high Negative Predictive Value (0.98) was achieved, insufficient recall (0.79) precludes its use as a reliable rule-out test. Interpretability analysis with SHapley Additive exPlanations (SHAP) revealed that predictions were driven by non-specific features like age and markers of inflammation and anemia. It is concluded that while a statistically detectable cancer signal exists in routine lab data, it is too weak and confounded for clinically reliable discrimination from normal aging or other inflammatory conditions. This work establishes a critical performance ceiling for this data modality in isolation and underscores that meaningful progress in computational veterinary oncology will require integration of multi-modal data sources.

cs.LG

The Adoption Paradox for Veterinary Professionals in China: High Use of Artificial Intelligence Despite Low Familiarity

While the global integration of artificial intelligence (AI) into veterinary medicine is accelerating, its adoption dynamics in major markets such as China remain uncharacterized. This paper presents the first exploratory analysis of AI perception and adoption among veterinary professionals in China, based on a cross-sectional survey of 455 practitioners conducted in mid-2025. We identify a distinct "adoption paradox": although 71.0% of respondents have incorporated AI into their workflows, 44.6% of these active users report low familiarity with the technology. In contrast to the administrative-focused patterns observed in North America, adoption in China is practitioner-driven and centers on core clinical tasks, such as disease diagnosis (50.1%) and prescription calculation (44.8%). However, concerns regarding reliability and accuracy remain the primary barrier (54.3%), coexisting with a strong consensus (93.8%) for regulatory oversight. These findings suggest a unique "inside-out" integration model in China, characterized by high clinical utility but restricted by an "interpretability gap," underscoring the need for specialized tools and robust regulatory frameworks to safely harness AI's potential in this expanding market.

cs.CY

Carleman estimates and some inverse problems for the coupled quantitative thermoacoustic equations by boundary data. Part I: Carleman estimates

In this paper, we consider Carleman estimates and inverse problems for the coupled quantitative thermoacoustic equations. In Part I, we establish Carleman estimates for the coupled quantitative thermoacoustic equations by assuming that the coefficients satisfy suitable conditions and taking the usual weight function $\varphi(x,t)={\rm e}^{\lambda\psi(x,t)}$, $\psi(x,t)=\left|x-x_{0}\right|^{2}-\beta\left(t-t_0\right)^{2}+\beta t_0^{2}$ for $x$ in a bounded domain in $\mathbb{R}^{n}$ with $C^{3}$-boundary and $t\in(0, T)$, where $t_0=T/2$. We will discuss applications of the Carleman estimates to some inverse problems for the coupled quantitative thermoacoustic equations in the succeeding Part II paper \cite{part II}.

math.AP

Observability and Control Property for a Singular Heat Equation with Variable Coefficients

The goal of this paper is to analyze control properties of the parabolic equation with variable coefficients in the principal part and with a singular inverse-square potential:\,$\partial_tu(x,t)-{\rm div}(p(x)\nabla u(x,t))-({\mu}/{|x|^2})u(x,t)=f(x,t).$ Here $\mu$ is a real constant . It was proved in the paper of Goldstein and Zhang (2003) that the equation is well-posedness when $0\leq{\mu\leq p_1(n-2)^2/4}$, and in this paper, we mainly consider the case $0\leq\mu<({ p_1^2}/{ p_2})(n-2)^2/4$ , where $ p_1,p_2$ are two positive constants which satisfy:\, $ 0< p_1\leq p(x)\leq p_2 , \forall x\in\overline{\Omega}$. We extend the specific Carleman estimates in the paper of Ervedoza (2008) and Vancostenoble (2011) to the equation we consider and apply it to deduce an observability inequality for the system. By this inequality and the classical HUM method, we obtain that we can control the equation from any non-empty open subset as for the heat equation. Moreover, we will study the case $\mu>p_2(n-2)^2/4$. We consider a sequence of regularized potentials $\mu/(|x|^2+\epsilon^2),$ and prove that we cannot stabilize the corresponding systems uniformly with respect to $\epsilon>0,$ due to the presence of explosive modes which concentrate around the singularity.

math.AP