SearcharxivSearch

arXiv subjects

Chuan Li

Publications and source records attributed to Chuan Li.

At least 19 recordsLinked to original sources

The Role of Preceding CMEs and SIRs in Enhancing Shock Acceleration of Electrons

The role of large-scale pre-existing interplanetary structures, including preceding coronal mass ejections (CMEs) and stream interaction regions (SIRs), in shaping the shock acceleration environment for energetic electrons remains not fully understood. In this study, we investigate nine interplanetary shocks observed by the Solar Terrestrial Relations Observatory (STEREO) that are associated with significant MeV electron enhancements, as such enhancements are rarely observed at interplanetary shocks. We combine remote-sensing observations, drag-based modeling, and in-situ measurements to analyze the shock propagation through pre-existing interplanetary structures. Eight of the nine events are associated with a preceding slow or intermediate-speed CME, while six shocks are in-situ observed propagating within preceding ICMEs, indicating that large-scale upstream trapping structures are a common feature of these events. Further analysis identifies three distinct scenarios associated with enhanced electron acceleration: shocks propagating through preceding ICMEs, shock-SIR interactions, and direct injection of flare-accelerated electrons into SIRs. As a representative shock-in-ICME event, the 2012 January 29 low-$\beta$, quasi-perpendicular shock ($\theta_{Bn}\sim87^\circ$) is further investigated using observations together with one-dimensional Monte Carlo test-particle simulations of a shock propagating into a large-scale upstream magnetic loop. The simulation suggests that the upstream loop prolongs electron residence near the shock and substantially enhances acceleration efficiency. These results demonstrate that large-scale interplanetary structures can precondition the upstream magnetic environment, providing favorable conditions for prolonged electron residence and efficient shock acceleration.

astro-ph.SR

Context-Aware Intelligent Vehicles

Intelligent vehicles increasingly support adaptive applications beyond driving themselves, ranging from context-aware ADAS and automated driving to in-cabin monitoring and fleet management, all under tight requirements on accuracy, latency, cost, and reliability. Meeting these requirements is challenging because vehicles operate in complex, uncertain, and rapidly changing environments while running on resource-constrained computing platforms. This paper argues that context-situational factors that give meaning to sensor signals and constrain decisions-should be treated as a first-class principle for next-generation vehicle systems, and operationalized as a unified, shared state for learning, risk assessment, and closed-loop control across the software stack. We systematically review state-of-the- art (SOTA) context-aware methods spanning (i) environment understanding, (ii) planning and control, (iii) safety and security, and (iv) connected vehicles. Based on a trend analysis of context-aware design, we identify four key technical challenges in building a general contextual engine for future intelligent vehicles: multi-modal context fusion, temporal context modeling, handling rare events, and collaborative context sharing. We hope this survey will motivate the development of robust and efficient context-aware vehicle applications.

cs.RO

Particle production in $p$-O collisions at LHC energy

We compare our previous predictions for charged-hadron production in $p$-O collisions at an energy of $\sqrt{s_\mathrm{NN}}=9.618$ TeV as published in 2025 with preliminary Run 3 data from ALICE and adapt the model parameters to the data. Our three-source model comprises a gluon-gluon central source and two valence-quark soft-gluon fragmentation sources. We describe the initial conditions using color-glass-condensate states and calculate the time-dependent partial thermalization with a relativistic diffusion model. An inversion of the maximum production amplitude at larger pseudorapidities from backward to forward towards peripheral collisions is predicted.

hep-ph

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.

cs.CV

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields $0$-$1\%$ attack success rate (ASR); allowing 15 rounds of adaptive attack yields $5.4$-$14.0\%$. Pooling three frontier attacker LLMs uncovers $1.4$-$2.2\times$ as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity ($0.02$-$0.14$) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are tied in aggregate ($5.4\%$ each; overlapping $95\%$ CIs), but their weaknesses differ sharply: on one scenario Opus reaches $60\%$ ASR ($95\%$ CI $36$--$80\%$) while GPT-5.4 and Gemini each stay at $7\%$ (CI $1$-$30\%$; the gap is preserved in a higher-$N$ replication). $13$ of $21$ scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's $W = 0.19$). We release the benchmark -- 21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI -- plus 945 transcripts from the 3$\times$3 frontier matrix, an attack-replay dataset, and 18{,}422 gpt-oss-20b battles from an open competition's final scoring rounds.

cs.CR

Three-dimensional evolution of a solar filament with multipoint observations

In this paper, we first devise a geometrical model, featuring a torus-like flux rope based on the shape of 3DCORE model. The global shape of the torus is an ellipse, while the cross sections are circular along the torus. The thinnest point is located between the Sun center and photosphere. Deflections and inclination are considered as well. Using multiwavelength observations from perspectives of Earth, Ahead-STEREO (STA), and Solar Orbiter, we apply the model to three-dimensional (3D) reconstructions and tracking of the filament eruption, which was associated with a flare and a coronal mass ejection (CME) on 2024 October 8. The morphology, direction, and true velocity ($\sim$433 km/s) of the eruptive filament are obtained. It is found that the filament propagates nonradially, deflecting slightly eastward by $\sim$10 degrees and significantly southward by $\sim$40 degrees. Trajectory of the filament in the ecliptic plane reveals that the filament moves toward STA. The true direction of the eruptive filament using imaging and spectral observations is mutually verified by 3D reconstructions. The heliocentric distance of the filament increases from $\sim$1.68 to $\sim$2.94 solar radii within 35 minutes. Based on the results of 3D reconstructions, the true speed of the CME leading edge is evaluated to be 1046$-$1145 km/s.

astro-ph.SR

Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks

$\mathrm{E}(3)$-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the $O(L^6)$ complexity of the Clebsch-Gordan Tensor Product (CGTP). The recently proposed Gaunt Tensor Product (GTP) reduces the complexity but is unable to capture the antisymmetric paths, resulting in incomplete expressivity. In this work, we present SpinGTP, an approach to overcome the GTP incompleteness by generalizing from scalar functions to Spin-Weighted Spherical Harmonics (SWSH). By relying on the algebraic properties of SWSH, SpinGTP recovers the missing antisymmetric interactions while maintaining the asymptotic efficiency of GTP. It also allows for a more expressive equivariant basis that naturally accounts for the parity-odd components of tensor products. We evaluate SpinGTP across diverse benchmarks, including Tetris, 3BPA, SPICE-MACE-OFF, and OC20. Our results show that SpinGTP achieves accuracies comparable to full CGTP. Notably, by explicitly capturing antisymmetric paths, SpinGTP exhibits superior performance in tasks involving chiral materials and non-centrosymmetric geometries. This work provides a complete, scalable, and mathematically rigorous path toward high-order equivariance in large-scale 3D atomistic system simulations.

cs.LG

Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition

Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged database. While recent methods achieve impressive performance, they overlook a serious long-tailed problem hidden in urban-scale datasets, which biases the model towards locations with abundant images and ignores less-visited areas, causing models to systematically favor frequently photographed locations while failing in sparsely covered areas. In this paper, we systematically characterize this imbalance challenge and propose Distribution-Aware Place Recognition (DAPR), a model-agnostic plug-in framework that rebalances gradient contributions across head and tail classes. Additionally, within classification-retrieval pipelines, DAPR applies a multi-scale distance search mechanism to compute per-class distributional compactness, providing complementary gains at the retrieval stage. On the large-scale SF-XL benchmark, our framework outperforms the previous classification-retrieval baseline by 18.3% on test set v1, and 6.7% on test set v2. As a plug-in module, it achieves consistent improvements across representative VPR methods on SF-XL, MSLS, and Pitts30k, demonstrating broad generalizability across different methods and benchmarks.

cs.CV

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often depend on brittle geometric procedures and hard decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition across spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN remains stable and outperforms strong baselines. On the real-world MindCube benchmark, SATURN achieves 78.57% overall accuracy, outperforming the strongest baseline by 14 pp.

cs.CV

Semi-Automatic Correction of 3D Tubular Structure Skeletons via Component-Wise MST and Filtered Delaunay Triangulation

Skeletonization of tubular structures from 3D imaging is essential for tasks such as morphometric analysis, transport or flow simulation, and procedural planning in domains including vascular networks, plant root systems, and neural connectomes. However, automatic skeleton extraction often introduces topological artifacts, such as erroneous connections between nearby branches and fragmented centerlines caused by noise or missing data. Correcting these artifacts manually can be time-consuming and error-prone, especially when precise interaction is required. We present a semi-automatic correction method that reconstructs a plausible centerline connection from minimal user input. Given a user-selected source and target point, our method traces a path by combining (i) component-wise minimum spanning trees for stable local propagation and (ii) a filtered 3D Delaunay edge graph for bridging gaps and handling ambiguous junctions. Candidate steps are ranked using a score that accounts for direction continuity, spatial proximity, component consistency, and target-directed progress. The output is an ordered polyline (or edge sequence) that can be used as a suggested correction and integrated into downstream skeleton post-processing workflows. We implement the system in C++ with an interactive viewer based on Libigl and demonstrate representative qualitative results on brain vessel datasets, including correction of typical "crossing" and "dotted" artifacts. While our current validation is qualitative, the method is lightweight and serves as a practical building block toward more comprehensive interactive correction pipelines in biomedical imaging and related domains.

cs.CG

3D-DLP: Self-Supervised 3D Object-Centric Scene Representation Learning

We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint position, bounding box dimensions, and appearance features, and represents a distinct entity in the scene. The model learns interpretable per-particle segmentation maps through an end-to-end self-supervised reconstruction objective. We demonstrate on both simulated and real-world datasets that the learned latent space is interpretable and controllable: by manipulating particle positions and decoding, we can generate novel scene configurations. Furthermore, we show that leveraging these compact 3D latent particles for downstream robotic manipulation improves performance over baselines that either lack explicit 3D information or rely on memory-intensive dense 3D inputs without object-centric structure. Code and videos are available at https://eubooks3003.github.io/3d-dlp.

cs.LG

Interpretability Transfer from Language to Vision via Sparse Autoencoders

Recent advances in language model interpretability using sparse autoencoders (SAEs) have yet to effectively translate to the visual domain, mainly due to the difficulty and ambiguity of labeling visual concepts. In this paper, we introduce Visual Interpretability via SAE Transfer Alignment (VISTA), a framework that transfers interpretability from language to vision in a LLaVA-style vision-language model by constraining a visual projector to map visual tokens into an LLM's pre-existing, labeled textual SAE space. This approach enables visual interpretability without training dedicated vision SAEs. By regularizing the projector using the LLM's SAE reconstruction loss, VISTA achieves a threefold increase in the matching rate, which measures how accurately the most activating textual concepts in the SAE space correspond to semantic elements in the image. Using this framework, we further analyze spatial localization properties of different vision encoders and show that DINOv2 features have stronger localization abilities than other encoders. Leveraging this precision, we validate VISTA's cross-modal alignment through fine-grained, localized concept interventions, where specific objects are removed or replaced in the model's perception while preserving the surrounding scene. This results in improvements of 35% in object removal and 47% in object replacement tasks over vision-only baselines, providing causal evidence that visual tokens inhabit the text SAE manifold. These contributions are validated across multiple LLM architectures.

cs.CV

OpenJarvis: Personal AI, On Personal Devices

Personal AI stacks, like OpenClaw and Hermes Agent, are becoming central to daily work, yet they route nearly every query (often over sensitive local data) to cloud-hosted frontier models. Replacing frontier models with local models inside existing stacks does not work: swapping Claude Opus 4.6 for Qwen3.5-9B drops accuracy by 25-39 pp across personal AI tasks like PinchBench and GAIA. Existing stacks bundle agentic prompts, tool descriptions, memory configuration, and runtime settings around a specific cloud model. Only the prompts can be tuned, and state-of-the-art prompt optimizers close just 5 pp of the local-cloud gap on their own. This motivates a decomposed personal AI stack: one that exposes individual primitives which can be optimized individually or jointly to close the local-cloud gap. We present OpenJarvis, an architecture that represents a personal AI system as a typed spec over five primitives: Intelligence, Engine, Agents, Tools & Memory, and Learning. Each primitive is an independently editable field, making the stack end-to-end optimizable and measurable against accuracy, cost, and latency. Towards closing the local-cloud gap without surrendering local-model properties, OpenJarvis introduces LLM-guided spec search, a local-cloud collaboration in which frontier cloud models propose edits across the spec at search time, only non-regressing edits are accepted, and the resulting spec runs entirely on-device at inference time. With LLM-guided spec search, on-device specs match or exceed cloud accuracy on 4 of 8 benchmarks and land within 3.2 pp of the best cloud baseline on average. They also reduce marginal API cost by ~800x and end-to-end latency by 4x.

cs.LG

Can We Distinguish the Source Region Location of Filament/Prominence Eruptions from the Sun-as-a-star H$\alpha$ Spectrum?

Solar filament/prominence eruptions can significantly perturb geospace when originating from favorable source locations and directions. While stellar analogs have been recently reported, the disk locations and magnetic environments of their source regions remain spatially unresolved on other stars. To bridge this gap, we investigate the typical Sun-as-a-star H$\alpha$ temporal spectral characteristics of solar filament/prominence eruptions with different source region locations (on-disk vs. limb, active region vs. quiet-Sun region). It is revealed that limb eruptions are characterized by blueshifted/redshifted emission caused by the bright off-limb erupting structures, whereas on-disk eruptions may show blueshifted absorptions due to the dark erupting filaments. Among the limb eruptions, front-side limb eruptions usually display line center emission before the blueshifted/redshifted emission, while far-side limb eruptions show the opposite sequence. Moreover, the magnetic environment at source also shapes the spectral characteristics. On-disk filament eruptions from active region exhibit much more intense flare-ribbon-dominated line center emission features compared with those from quiet-Sun region. Limb active region eruptions often show single-wing emissions, whereas large-scale quiet-Sun region (quiescent) prominence eruptions frequently display expansion-induced emission in both wings followed by line center absorption due to the disappearance of bright prominence. These distinct Sun-as-a-star H$\alpha$ spectral characteristics, dependent on eruption location, provide a diagnostic basis for inferring source regions of stellar filament/prominence eruptions from spatially unresolved H$\alpha$ spectra.

astro-ph.SR

Magnetic Evolution of Highly-Sheared Region in Active Region 13842 Producing Large X9.0 Flare

Shearing motion and magnetic flux cancellation around the polarity inversion line (PIL) play significant roles in the build-up of free magnetic energy and magnetic flux rope (MFR) in source region of major solar flares. Here we investigate the magnetic evolution of a highly-sheared PIL in active region (AR) 13842, hosting the largest X9.0 flare of Solar Cycle 25. Since 2024 September 29, a positive-polarity pore persistently drifted northward along the western side of the AR's main negative-polarity sunspot. The main sunspot remained stationary until negative-polarity patches successively emerged to its east and approached. Rear-ended by these same-polarity patches, the sunspot then began moving westward toward the opposite-polarity pore around October 1, forming a collisional PIL. Meanwhile, on the PIL's other side, the pore was also rear-ended by same-polarity patches sequentially emerging behind it, accelerating the shearing motion around the PIL, where frequent flux cancellations were also observed. Synchronous rapid accumulation of free magnetic energy and formation of MFR were then observed in the PIL, where multiple major flares successively occurred within two days. Before these large flares, the area and total free energy of the high-free-energy-density PIL region gradually decreased in the photosphere, which could be caused by the initial ascent of MFR before eruption and serve as a precursor of solar eruptions. These results suggest that persistent flux emergences with cross separation directions facilitates rapid formation of collisional shearing PIL and frequent flux cancellations, leading to repeated MFR formations and multiple large flares in a relatively short time.

astro-ph.SR

Tailored Speckle Illumination Microscopy with Enhanced Sectioning and Image Quality

Optical speckle patterns have been widely used for illumination in computational imaging, optical sectioning microscopy, and super-resolution imaging. However, commonly used speckles satisfy Rayleigh statistics, which are not ideal for diverse imaging applications. Here we tailor three-dimensional speckle intensity statistics for dynamic speckle illumination microscopy based on linear fluorescence. Optical sectioning is enhanced by axially varying speckle contrast, and image reconstruction noise is minimized with in-focus speckles of binary intensities. The customized speckle statistics are shown to tolerate sample-induced aberration and scattering. We apply tailored speckle illumination to mouse brain vascular imaging and demonstrate much improved image quality than optical-sectioning structured illumination. These results establish customization of speckle intensity statistics as a promising strategy for robust, high-throughput fluorescence imaging in thick, scattering biological specimens.

physics.optics

Solar Energetic Particle Events and Associated Type II Radio Bursts from Different Source Regions

Large solar energetic particle (SEP) events are thought to originate from the shocks driven by fast coronal mass ejections (CMEs) and thus generally accompanied by type II radio bursts. However, a significant proportion of type II radio bursts is not accompanied by SEP events. To study the relationship between SEPs and type II radio bursts and the associated physical mechanisms, we statistically analyze 43 SEP halo-CMEs and 131 non-SEP halo-CMEs observed from 2010 to 2024, and check the related properties of type II radio bursts and solar source region. We find nearly all SEP events and approximately two-thirds of non-SEP events are accompanied by type II radio bursts. Type II radio bursts associated with SEP events usually have longer duration and lower ending frequencies. The starting frequency exhibits a clear source region dependence, being highest for ''single active region (AR)'', intermediate for ''multiple ARs'', and lowest for ''outside of ARs''. Furthermore, the spectra of both protons and electrons exhibit a similar softening trend in the three types of source regions. Joint analysis of spectra and type II radio bursts reveals that the proton spectra index has a good anti-correlation with the starting frequency of the type II radio bursts. Our statistical results have important implications for the mechanisms behind SEP acceleration

astro-ph.SR

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, where decision making relies on a focused readout of relevant memories rather than replaying all pathways, we introduce RISE (Readout Influence Sketching Estimator). Instead of computing and indexing gradients across the entire LLM, RISE focuses on influence hotspots at the output layer, where influence signals concentrate, and the gradient admits a decomposed outer-product form. This enables a dual-channel representation combining a lexical residual channel (RH) and a semantic projected-error channel (GH). Applying CountSketch projections to these channels achieves strong compression while maintaining accurate attribution. Across the OLMo (1B-32B) and Pythia (14M-6.9B) families, RISE reduces index storage by up to 112$\times$ compared to RapidIn and scales to 32B parameters LLM, where gradient-based baselines such as RapidIn and ZO-Inf become memory-infeasible. We evaluate RISE on two paradigms: (1) retrospective attribution, retrieving influential training examples for specific predictions, and (2) prospective valuation, scoring candidate data utility zero-shot. We validate RISE on three tasks: Howdy backdoor data detection, Finance-Medical domain separation, and Brain Rot high-quality data selection. In a closed-loop Brain Rot study, continued pretraining on RISE-selected data yields consistent downstream improvements. Overall, RISE provides a practical and scalable primitive for influence analysis and training-data selection in modern large language models.

cs.LG