SearcharxivSearch

arXiv subjects

Lina Zhang

Publications and source records attributed to Lina Zhang.

At least 19 recordsLinked to original sources

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.4% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

cs.RO

Curvature-induced scalarization of charged AdS black holes

We investigate how a negative cosmological constant affects the Gauss-Bonnet (GB) scalarization in the Einstein-Maxwell-scalar-Gauss-Bonnet theory with a scalar coupling constant $\eta$ to GB term. We focus on the instability of Reissner-Nordstr\"om-AdS (RN-AdS) black holes under a scalar perturbation governed by an effective mass $\mu^2_{\text{eff}}$ sourced by the GB term. Unlike the asymptotically flat spacetime case, the onset of scalarization is not merely determined by $\mu^2_{\text{eff}} < 0$, but it is constrained by the Breitenlohner-Freedman (BF) bound. In case that the BF bound is violated ($\eta>2.25$ with $\Lambda=-0.5$), one may find AdS-tachyonic instability. We find that for $0<\eta<2.25$, the GB$^+$ scalarization may be performed through spontaneous scalarization, while for $\eta<0$ the GB$^-$ scalarization is found to give the single branch of scalarized AdS black holes. For the GB$^+$ scalarization in $\eta_{th}\le\eta<2.25$ with $\eta_{th}$ threshold instability, we obtain the single branch ($n=0$ fundamental branch) of scalarized AdS black holes, in contrast to the infinite branches in asymptotically flat spacetime. A bulk fixed-charge thermodynamic analysis is performed thoroughly for GB$^\pm$ scalarizations.

gr-qc

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporally evolving pathologic motor behaviors such as seizure semiology remains largely untested. To address this gap, we introduce Seizure-Semiology-Suite, a clinically grounded dataset and benchmark for fine-grained, structured seizure semiology understanding. The dataset includes 438 seizure videos annotated with over 35,000 dense labels covering 20 ILAE-defined semiological features. Building on this dataset, we propose a seven-task hierarchical benchmark that systematically evaluates MLLMs from low-level visual perception to temporal sequencing, narrative report generation, and seizure diagnosis. To enable clinically meaningful evaluation of generated reports, we further introduce the Report Quality Index for Seizure Semiology (Seizure-RQI). Extensive baselines across 11 open-weight MLLMs reveal systematic weaknesses in laterality reasoning, temporal localization, symptom sequencing, and clinically faithful reporting. We show that seizure-specific fine-tuning substantially improves performance across tasks, and that a two-stage neuro-symbolic framework achieves an F1 score of 0.96 on epileptic versus non-epileptic seizure classification. Seizure-Semiology-Suite establishes a rigorous benchmark for evaluating multimodal models in safety-critical medical video understanding and guides the development of clinically reliable, domain-adaptive multimodal intelligence.

cs.CV

Charge-dependent scalarization of Einstein- Euler-Heisenberg black holes

Charge-dependent scalarization of the Einstein-Euler-Heisenberg (EEH) black hole is carried out in the EEH-scalar theory by introducing an exponential scalar coupling with $\alpha$ coupling constant to the Maxwell and nonlinear electrodynamic terms. The bald black hole (EEHBH) is described by mass $M$ and arbitrary magnetic charge $q$ and has a single horizon when choosing the action parameter $\mu=0.3$. The spontaneous scalarization ($\alpha^+$) of this black hole is available for charge $0 q_c$ and negative $\alpha$. The former case of $q=0.5$ implies infinite branches of scalarized EEHBHs but its fundamental branch ($n=0$) is stable against radial perturbations, while the latter cases of $q=2,20$ show two stable single branches of scalarized EEHBHs.

gr-qc

Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant involuntary movements in neurological disorders remains largely unexplored. This pilot study evaluates the capability of MLLMs for automated recognition of pathological movements in seizure videos. We assessed the zero-shot performance of state-of-the-art MLLMs on 20 ILAE-defined semiological features across 90 clinical seizure recordings. MLLMs outperformed fine-tuned Convolutional Neural Network (CNN) and Vision Transformer (ViT) baseline models on 13 of 18 features without task-specific training, demonstrating particular strength in recognizing salient postural and contextual features while struggling with subtle, high-frequency movements. Feature-targeted signal enhancement (facial cropping, pose estimation, audio denoising) improved performance on 10 of 20 features. Expert evaluation showed that 94.3 percent of MLLM-generated explanations for correctly predicted cases achieved at least 60 percent faithfulness scores, aligning with epileptologist reasoning. These findings demonstrate the potential of adapting general-purpose MLLMs for specialized clinical video analysis through targeted preprocessing strategies, offering a path toward interpretable, efficient diagnostic assistance. Our code is publicly available at https://github.com/LinaZhangUCLA/PathMotionMLLM.

cs.CV

Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration

Multimodal retrieval systems typically employ Vision Language Models (VLMs) that encode images and text independently into vectors within a shared embedding space. Despite incorporating text encoders, VLMs consistently underperform specialized text models on text-only retrieval tasks. Moreover, introducing additional text encoders increases storage, inference overhead, and exacerbates retrieval inefficiencies, especially in multilingual settings. To address these limitations, we propose a multi-task learning framework that unifies the feature representation across images, long and short texts, and intent-rich queries. To our knowledge, this is the first work to jointly optimize multilingual image retrieval, text retrieval, and natural language understanding (NLU) tasks within a single framework. Our approach integrates image and text retrieval with a shared text encoder that is enhanced by NLU features for intent understanding and retrieval accuracy.

cs.IR

Chaotic motion of particles around a dyonic Kerr-Newman black hole immersed in the Melvin-swirling universe

We employ the Poincar\'{e} section, fast Lyapunov indicator, recurrence analysis, bifurcation diagram and basins of attraction to investigate the dynamical behaviors of the motion of particles around a new dyonic Kerr-Newman black hole immersed in the Melvin-swirling universe presented in [A. Di Pinto, S. Klemm, and A. Vigan\`o, J. High Energy Phys. {\bf 06}, 150 (2025)]. We note that the swirling parameter $j$ and magnetic field strength $B$ make the equations of motion for particles nonseparable, and confirm the presence of chaotic behavior in the motion in this dyonic Kerr-Newman-Melvin-swirling spacetime and its sub-cases by removing the conical singularities and removing both the conical singularities and the Dirac strings. We observe that both the number of chaotic orbits and the chaotic region increase with the increase of the parameters $j$ and $B$, but decrease as the electric charge $Q$, magnetic charge $H$ or spin parameter $a$ increases. Moreover, we find that the presence of $j$ changes the ranges of $B$, $Q$, $H$ and $a$ where the chaotic motion appears for particles. The swirling parameter together with the magnetic field strength, electric charge, magnetic charge and spin parameter yields richer physics in the motion of particles for the spacetime of a dyonic Kerr-Newman black hole immersed in the Melvin-swirling universe.

gr-qc

OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU

The analysis of character appearance frequency is essential for understanding narrative structure, character prominence, and story progression in anime. In this work, we introduce OregairuChar, a benchmark dataset designed for appearance frequency analysis in the anime series My Teen Romantic Comedy SNAFU. The dataset comprises 1600 manually selected frames from the third season, annotated with 2860 bounding boxes across 11 main characters. OregairuChar captures diverse visual challenges, including occlusion, pose variation, and inter-character similarity, providing a realistic basis for appearance-based studies. To enable quantitative research, we benchmark several object detection models on the dataset and leverage their predictions for fine-grained, episode-level analysis of character presence over time. This approach reveals patterns of character prominence and their evolution within the narrative. By emphasizing appearance frequency, OregairuChar serves as a valuable resource for exploring computational narrative dynamics and character-centric storytelling in stylized media.

cs.CV

Spontaneous scalarization of regular Hayward black holes in Einstein-nonlinear electromagnetic-scalar gravity

Regular Hayward black holes provide a useful setting for investigating scalarization in theories with nonminimally coupled matter sectors. Within the framework of Einstein-nonlinear electromagnetic-scalar gravity, we identify the tachyonic threshold that signals the bifurcation from the bald Hayward background and then obtain scalarized charged black holes for both quadratic $(1-\alpha\phi^2)$ and exponential $(e^{-\alpha \phi^2})$ couplings. These configurations form a discrete set of branches classified by the number of nodes in the scalar field. The branch with $n=0$ is the fundamental branch, whereas solutions with $n\geq 1$ are excited branches. By studying radial perturbations, we find that the fundamental branch is stable for both coupling choices, which makes it the most relevant branch for future phenomenological and observational studies.

gr-qc

OCTOPUS: A Versatile, User-Friendly, and Extensible Public Code for General-Relativistic Ray-Tracing in Spherically Symmetric and Static Spacetimes

This paper presents OCTOPUS, a relativistic ray-tracing algorithm developed within a Fortran-based, OpenMP-accelerated framework, designed for asymptotically flat, spherically symmetric curved spacetimes. The code efficiently and accurately computes key relativistic features -- including the black hole event horizon, photon rings, critical curves, and innermost stable circular orbits -- and simulates black hole shadows, redshift factor distributions, accretion disk images, toroidal images, as well as gravitational lensing, light curves, and gravitational radiation from hot-spots. OCTOPUS provides an automated, modular solution for qualitative studies of black hole observables and multi-messenger correlations between electromagnetic and gravitational signals in curved spacetime. Its implementation requires only the metric potential and its first-, second-, and third-order radial derivatives as input, ensuring low user barriers while remaining highly extensible and adaptable. Using a Schwarzschild black hole surrounded by a Dehnen-type dark matter halo, we thoroughly validate the algorithm's precision, efficiency, and functionality, and investigate how dark matter halo parameters affect observational signatures. Our results demonstrate that increasing the scale and density of the dark matter halo strengthens the spacetime's gravitational field, an effect clearly reflected in black hole images and supported by hot-spot light curve signatures. A future version of OCTOPUS, with expanded capabilities for axisymmetric spacetimes, is planned for release.

gr-qc

Newly scalarization of the Einstein-Euler-Heisenberg black hole

Th spontaneous scalarization of the Einstein-Euler-Heisenberg (EEH) black hole is performed in the EEH-scalar theory by introducing an exponential scalar coupling (with $\alpha$ coupling constant) to the Maxwell term.Here, the EEH black hole as a blad black hole is described by mass $M$ and magnetic charge $q$ with an action parameter $\mu$. A choice of $\mu=0.3$ gurantees a single horizon with unrestricted magnetic charge $q$. The onset scalarization of this black hole appears for a positive coupling $\alpha$ with an unlimited magnetic charge $q$. However, there exists a difference between $q\le1$ and $q>1$ onset scalarizations. We notify the presence of infinite branches labeled by the number of $n=0,1,2,\cdots$ of scalarized charged black holes by taking into account the scalar seeds around the EEH black hole. We find that the $n=0$ fundamental branch of all scalarized black holes is stable against the radial perturbations, while the $n=1$ excited branch is unstable.

gr-qc

A Descriptor Is All You Need: Accurate Machine Learning of Nonadiabatic Coupling Vectors

Nonadiabatic couplings (NACs) play a crucial role in modeling photochemical and photophysical processes with methods such as the widely used fewest-switches surface hopping (FSSH). There is therefore a strong incentive to machine learn NACs for accelerating simulations. However, this is challenging due to NACs' vectorial, double-valued character and the singularity near a conical intersection seam. For the first time, we design NAC-specific descriptors based on our domain expertise and show that they allow learning NACs with never-before-reported accuracy of $R^2$ exceeding 0.99. The key to success is also our new ML phase-correction procedure. We demonstrate the efficiency and robustness of our approach on a prototypical example of fully ML-driven FSSH simulations of fulvene targeting the SA-2-CASSCF(6,6) electronic structure level. This ML-FSSH dynamics leads to an accurate description of $S_1$ decay while reducing error bars by allowing the execution of a large ensemble of trajectories. Our implementations are available in open-source MLatom.

physics.comp-ph

Policy Relevant Treatment Effects with Multidimensional Unobserved Heterogeneity

This paper provides a unified framework for bounding policy relevant treatment effects using instrumental variables. In this framework, the treatment selection may depend on multidimensional unobserved heterogeneity. We derive bilinear constraints on the target parameter by extracting information from identifiable estimands. We apply a convex relaxation method to these bilinear constraints and provide conservative yet computationally simple bounds. Our convex-relaxation bounds extend and robustify the bounds by Mogstad, Santos, and Torgovitsky (2018) which require the threshold-crossing structure for the treatment: if this condition holds, our bounds are simplified to theirs for a large class of target parameters; even if it does not, our bounds include the true parameter value whereas theirs may not and are sometimes empty. Linear shape restrictions can be easily incorporated to narrow the proposed bounds. Numerical and simulation results illustrate the informativeness of our convex-relaxation bounds.

econ.EM

Partial Identification of Distributional Treatment Effects in Panel Data using Copula Equality Assumptions

This paper aims to partially identify the distributional treatment effects (DTEs) that depend on the unknown joint distribution of treated and untreated potential outcomes. We construct the DTE bounds using panel data and allow individuals to switch between the treated and untreated states more than once over time. Individuals are grouped based on their past treatment history, and DTEs are allowed to be heterogeneous across different groups. We provide two alternative group-wise copula equality assumptions to bound the unknown joint and the DTEs, both of which leverage information from the past observations. Testability of these two assumptions are also discussed, and test results are presented. We apply this method to study the treatment effect heterogeneity of exercising on the adults' body weight. These results demonstrate that our method improves the identification power of the DTE bounds compared to the existing methods.

econ.EM

Zeolitic Imidazolate Framework-8 offers an anti-inflammatory and antifungal method in the treatment of Aspergillus fungus keratitis in vitro and in vivo

Background: Fungal keratitis is a serious blinding eye disease. Traditional drugs used to treat fungal keratitis commonly have the disadvantages of low bioavailability, poor dispersion, and limited permeability. Purpose: To develop a new method for the treatment of fungal keratitis with improved bioavailability, dispersion, and permeability. Purpose: To develop a new method for the treatment of fungal keratitis with improved bioavailability, dispersion, and permeability. Methods: Zeolitic Imidazolate Framework-8 (ZIF-8) was formed by zinc ions and 2-methylimidazole linked by coordination bonds and characterized by Scanning electron microscopy (SEM), X-ray diffraction (XRD), and Zeta potential. The safety of ZIF-8 on HCECs and RAW 264.7 cells was detected by Cell Counting Kit-8 (CCK-8). The anti-inflammatory effects of ZIF-8 on RAW 246.7 cells were evaluated by Quantitative Real-Time PCR Experiments (qPCR) and Enzyme-linked immunosorbent assay (ELISA). Clinical score, Colony-Forming Units (CFU). In vivo, treatment with ZIF-8 reduced corneal fungal load and mitigated neutrophil infiltration in fungal keratitis, which effectively reduced the severity of keratitis in mice and alleviated the infiltration of inflammatory factors in the mouse cornea. In addition, ZIF-8 reduces the inflammatory response by downregulating the expression of pro-inflammatory cytokines TNF-α, IL-6, and IL-1\b{eta} after Aspergillus fumigatus infection in vivo and in vitro. Conclusion: ZIF-8 has a significant anti-inflammatory and antifungal effect, which provides a new solution for the treatment of fungal keratitis.

q-bio.TO

ZIF-90 treats fungal keratitis by promoting macrophage apoptosis and inhibiting inflammatory response

Fungal keratitis is a severe vision-threatening corneal infection with a prognosis influenced by fungal virulence and the host's immune defense mechanisms. The immune system, through its regulation of the inflammatory response, ensures cells and tissues can effectively activate defense mechanisms in response to infection and injury. However, there is still a lack of effective drugs that attenuate fungal virulence while relieving the inflammatory response caused by fungal keratitis. Therefore, finding effective treatments to solve these problems is particularly important. We synthesized ZIF-90 by water-based synthesis and characterized by SEM, XRD etc. In vitro experiments included CCK-8 and ELISA. These evaluations verified the disruptive effects of ZIF-90 on Aspergillus. fumigatus spore adhesion, morphology, cell membrane, and the effect of ZIF-90 on apoptosis. In addition, to investigate whether the metal-ligand zinc and the organic ligand imidazole act as essential factors in ZIF-90, we investigated the in vitro antimicrobial and anti-inflammatory effects of ZIF-8, ZIF-67, and MOF-74 (Zn) by MIC and ELISA experiments. ZIF-90 has therapeutic effects on fungal keratitis, which could break the protective organelles of Aspergillus. fumigatus, such as the cell wall. In addition, ZIF-90 can avoid excessive inflammatory response by promoting apoptosis of inflammatory cells. The results demonstrated that both zinc ions and imidazole possessed antimicrobial and anti-inflammatory effects. In addition, ZIF-90 exhibited better biocompatibility compared to ZIF-8, ZIF-67, and MOF-74 (Zn). ZIF-90 has anti-inflammatory and antifungal effects and preferable biocompatibility, and has great potential for the treatment of fungal keratitis.

q-bio.SC

HiFiSeg: High-Frequency Information Enhanced Polyp Segmentation with Global-Local Vision Transformer

Numerous studies have demonstrated the strong performance of Vision Transformer (ViT)-based methods across various computer vision tasks. However, ViT models often struggle to effectively capture high-frequency components in images, which are crucial for detecting small targets and preserving edge details, especially in complex scenarios. This limitation is particularly challenging in colon polyp segmentation, where polyps exhibit significant variability in structure, texture, and shape. High-frequency information, such as boundary details, is essential for achieving precise semantic segmentation in this context. To address these challenges, we propose HiFiSeg, a novel network for colon polyp segmentation that enhances high-frequency information processing through a global-local vision transformer framework. HiFiSeg leverages the pyramid vision transformer (PVT) as its encoder and introduces two key modules: the global-local interaction module (GLIM) and the selective aggregation module (SAM). GLIM employs a parallel structure to fuse global and local information at multiple scales, effectively capturing fine-grained features. SAM selectively integrates boundary details from low-level features with semantic information from high-level features, significantly improving the model's ability to accurately detect and segment polyps. Extensive experiments on five widely recognized benchmark datasets demonstrate the effectiveness of HiFiSeg for polyp segmentation. Notably, the mDice scores on the challenging CVC-ColonDB and ETIS datasets reached 0.826 and 0.822, respectively, underscoring the superior performance of HiFiSeg in handling the specific complexities of this task.

cs.CV

Chaotic motion of particles in the spacetime of a Kerr black hole immersed in swirling universes

We investigate the motion of particles in the spacetime of a Kerr black hole immersed in swirling universes. Using the Poincaré section, fast Lyapunov exponent indicator, bifurcation diagram and basins of attraction, we present the effects of the swirling parameter and the spin parameter on the dynamical behaviors of the motion of particles, and confirm the presence of chaos in the motion of particles in this background spacetime. We find that the swirling parameter can change the range of the spin parameter where the chaos occurs, and vice versa. Moreover, we observe clearly that, regardless of the spin parameter, there exist some self-similar fractal fine structures in the basins boundaries of attractors for the spacetime of a black hole immersed in swirling universes. The combination the swirling parameter and the spin parameter provides richer physics in the motion of particles.

gr-qc