SearcharxivSearch

arXiv subjects

Jian Gao

Publications and source records attributed to Jian Gao.

At least 19 recordsLinked to original sources

LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction

Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric accuracy, temporal consistency, and computational efficiency. Motivated by the principles of Simultaneous Localization and Mapping (SLAM), we introduce LingBot-Map, a feed-forward 3D foundation model for reconstructing scenes from streaming data, built upon a geometric context transformer (GCT) architecture. A defining aspect of LingBot-Map lies in its carefully designed attention mechanism, which integrates an anchor context, a pose-reference window, and a trajectory memory to address coordinate grounding, dense geometric cues, and long-range drift correction, respectively. This design keeps the streaming state compact while retaining rich geometric context, enabling stable efficient inference at around 20 FPS on 518 x 378 resolution inputs over long sequences exceeding 10,000 frames. Extensive evaluations across a variety of benchmarks demonstrate that our approach achieves superior performance compared to both existing streaming and iterative optimization-based approaches.

cs.CV

Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training

Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as learning signals during training. We propose Prior-Guided Tuning (PGT), a training perspective that incorporates natural-language priors as auxiliary learning signals for low-resource LLM training. Under this perspective, we introduce Contrastive Prior Steering (CPS), which keeps the original supervised objective intact while adding positive and negative prior-conditioned auxiliary losses to encourage task-consistent learning and discourage plausible but misleading alternatives. Experiments on AmbiMath, Jigsaw, and MNLI/HANS show that CPS consistently improves over plain and prompt fine-tuning. On AmbiMath, CPS achieves 97.6% average exact-match accuracy. On Jigsaw, CPS improves average Macro F1 by 9.5 percentage points over standard fine-tuning, and with 1/10 of the experimental training data slightly exceeds full-data plain fine-tuning. On HANS, CPS improves non-entailment accuracy by 8.3 and 5.2 percentage points for LLaMA 3.1 8B and Qwen 2.5 7B, respectively, while maintaining comparable in-domain MNLI accuracy. These results support our central claim: task-level natural-language priors can provide useful guidance as auxiliary learning signals for low-resource LLM training. Our code and data will be publicly available.

cs.AI

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

Audio reasoning requires multi-step, evidence-grounded inference over temporally dynamic and acoustically mixed signals, exceeding conventional perception tasks such as ASR or captioning. We present VISA, our submission to the Interspeech 2026 Audio Reasoning Challenge (Agent Track), evaluated via the MMAR Rubrics for correctness and reasoning quality. Under a "LALM as a Tool" paradigm, VISA strengthens large audio language models with auxiliary multi-modal evidence while avoiding heavy orchestration. The system integrates three components: multi-modal feature extraction for complementary audio and acoustic-visual clues, model-voting inference with consistency checking for stable predictions, and fine-grained category-aware routing to resolve disagreements and select rubric-aligned reasoning chains. On the official Agent Track leaderboard, VISA ranks 2nd overall with a 66.23% Rubrics score. It also achieves 77.40% Accuracy, the highest among all systems listed across both the Single Model and Agent tracks.

eess.AS

Grounded Normative Rule Generation with Structured Search

Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks primarily reward surface fluency or schema compliance, leaving operational grounding weakly tested. This creates a critical vulnerability where standard language models generate plausible-sounding policies that fail during enforcement because they rely on unavailable data logs or misaligned scopes. To address this challenge, we formalize the problem as Grounded Normative Rule Synthesis (GNRS) and introduce GNRS-Search, a framework that utilizes Markov Chain Monte Carlo (MCMC) sampling to optimize a discrete, five-slot And-Or Graph (AOG). By explicitly decoupling intermediate operational structure from final prose generation, this method isolates executable feasibility from writing style and allows rule failures to be localized prior to surface realization. We evaluate our approach on GNRS-Bench, a benchmark spanning 116 controlled goals across eight scene families, and RealCharter-Bench, which evaluates transfer to 53 real-derived policy tasks with hidden source clauses. GNRS-Search raises average rubric quality from 68.8% to 81.0% and ranks first under a disclosed executable composite metric, while systematic slot interventions confirm that performance gains stem from robust operational logic rather than rhetorical tuning. Ultimately, by transforming automated rule drafting into an inspectable search problem, this work provides a foundational paradigm for deploying verifiable and compliance-ready personal agents within regulated environments.

cs.CL

EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry

Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits require deeper corruption of the very rhythm, transients, timbre, and long-range form that should remain unchanged. Here, we introduce EchoEdit, a training-free and inversion-free framework for real-audio editing that directly constructs an editing field by differencing the drifts conditioned on the source and target prompts. This construction avoids explicit source inversion, paired edit data, and test-time optimization, but its stochastic source marginals introduce uncertainty drift, where small random deviations accumulate along the editing trajectory and can move the edited latent away from the audio data manifold. To address this limitation, we further propose EchoEdit+, an optimal-transport-regularized extension that stabilizes the direct editing path by minimizing the transportation cost between edited variables and the source-conditioned audio manifold. The resulting OT coupling contracts the stochastic displacement at noisy states, keeps model queries closer to the training distribution, and preserves structural information while allowing semantic change. Experiments on sound-effect and music editing demonstrate that EchoEdit+ improves target-prompt alignment and source preservation over inversion-based baselines and the unregularized direct editor. Code and dataset will be released.

cs.SD

Plasma-state metasurfaces for ultra-intensive field manipulation

High-power lasers offer ultrahigh intensities for plasma interactions, but they lack advanced techniques to control the properties of the fields, because no optical elements could withstand their high intensities. The vibrant field of metasurfaces has transformed modern optics by enabling unprecedented control over light at subwavelength through deliberate design. However, metasurfaces have traditionally been limited to solid-state materials and low light intensities. Extending the sophisticated capabilities of metasurfaces from solids into the plasma realm would open new horizons for high-field science. Here, we experimentally demonstrate plasma-state metasurfaces (PSMs) through the photonic spin Hall effect and stable-propagating vortex beam generation irradiated by intense light. Time-resolved pump-probe measurements reveal that the functionality of PSMs can persist for several picoseconds, making them suitable for controlling ultra-intense femtosecond lasers, even in state-of-the-art multi-petawatt systems. Harnessing the powerful toolkit of metasurfaces, this approach holds the promise to revolutionize our ability to manipulate the amplitude, phase, polarization, and wavefront of high-power lasers during their pulse duration. It also opens new possibilities for innovative applications in laser-plasma interactions such as compact particle acceleration and novel radiation sources.

physics.plasm-ph

Infinite Worlds with Versatile Interactions

We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our model achieves an unbounded interaction horizon while maintaining consistent output quality, benefiting from a carefully crafted causal pretraining paradigm. (2) Through distilling a real-time variant from the base model, our system guarantees rapid response time, sufficient to drive 720p video streams at 60 fps. (3) Compared to the previous version, this update introduces highly diverse interactive elements, comprising a broader spectrum of actions (e.g., attacking, archery, spell-casting, and shooting) alongside a richer variety of text-driven events. (4) We pioneer the integration of an agentic harness within the domain of world modeling, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses. Additionally, to facilitate a shared experience, we develop an interface that permits multiple players to simultaneously immerse themselves in this vivid world simulator. We pair our primary 14B model with a lightweight 1.3B counterpart, which supports effortless deployment on a single GPU.

cs.CV

Evidence of triggered star formation in the Pillars of Creation from JWST observations

Stars form in molecular clouds under the influence of their local environments, yet the role of massive stellar feedback in either triggering or suppressing star formation remains a fundamental question in astrophysics. The Pillars of Creation in the Eagle Nebula, sculpted by ionizing radiation and stellar winds from massive stars in NGC 6611, offer a natural laboratory for investigating this question. Here we present high-resolution observations of the Pillars of Creation using the JWST Near Infrared Camera and Mid-Infrared Instrument, revealing 253 young stellar object (YSO) candidates. These YSO candidates show spatial correlations with the edges of feedback-driven structures, with overdensities along the boundaries. A weak trend of decreasing stellar age with increasing distance from the ionizing source was tentatively observed. There also appears to be an enhancement in the star formation rate within the past 1 Myr in this region. Such age and spatial associations suggest that while the bulk of the YSOs may have formed contemporaneously with the central cluster, a subset could be associated with triggered star formation. The JWST image of intricate structures, including a spiral-like disk and bi-reflection nebulae at the tips of Pillar I and Pillar II, further highlights the complexity of star formation processes.

astro-ph.SR

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO

Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried information, leading to the loss of task critical cues. To address this challenge, we propose FaVChat, the first VLLM designed for reasoning over subtle visual and dynamic facial cues. FaVChat introduces a hierarchical, prompt guided visual feature extraction framework that emphasizes question relevant information at three complementary levels. These multi level features are dynamically fused and injected into the LLM, enabling more accurate facial details reasoning To further improve learning efficiency under data scarcity, we propose Data Efficient GRPO, a reinforcement learning strategy that iteratively identifies high utility samples and maximizes the contribution of each instance via per instance utility estimation, substantially enhancing performance gains under limited supervision. We construct a large scale benchmark dataset FaVChat 170K, comprising approximately 60K high quality facial videos and 170K question answer pairs focusing on fine grained facial details. Extensive experiments, including zero shot evaluations on four facial understanding tasks, demonstrate that FaVChat consistently outperforms existing VLLMs.

cs.CV

Random Walk on Point Clouds for Feature Detection

The points on the point clouds that can entirely outline the shape of the model are of critical importance, as they serve as the foundation for numerous point cloud processing tasks and are widely utilized in computer graphics and computer-aided design. This study introduces a novel method, RWoDSN, for extracting such feature points, incorporating considerations of sharp-to-smooth transitions, large-to-small scales, and textural-to-detailed features. We approach feature extraction as a two-stage context-dependent analysis problem. In the first stage, we propose a novel neighborhood descriptor, termed the Disk Sampling Neighborhood (DSN), which, unlike traditional spatially and geometrically invariant approaches, preserves a matrix structure while maintaining normal neighborhood relationships. In the second stage, a random walk is performed on the DSN (RWoDSN), yielding a graph-based DSN that simultaneously accounts for the spatial distribution, topological properties, and geometric characteristics of the local surface surrounding each point. This enables the effective extraction of feature points. Experimental results demonstrate that the proposed RWoDSN method achieves a recall of 0.769-22% higher than the current state-of-the-art-alongside a precision of 0.784. Furthermore, it significantly outperforms several traditional and deep-learning techniques across eight evaluation metrics.

cs.CV

Coarse Graining Reveals a Fluctuation-theorem-like Asymmetry in Financial Markets

Fluctuation theorems show how coarse graining transforms microscopic symmetry into observable irreversibility. Here we ask whether an analogous symmetrybased diagnostic can be constructed for financial markets. At the microscopic level, each transaction pairs a buyer and a seller, whereas trading decisions are typically made from coarse-grained price histories. Using symmetric takeprofit and stop-loss rules, we compare the holding-time distributions of long and short trading ensembles generated from the same price series. Across equityindices, individual stocks and cryptocurrencies, the log-ratio of the two distributions shows a robust crossover. It remains nearly constant at short durations but becomes linear in the tail, implying an exponential directional asymmetry. The tail slope defines an effective market temperature, an operational measure of fluctuation intensity on the chosen observation scale. A Bachelier first-passage benchmark captures the exponential tails but not the asymmetry, because long and short positions share the same leading decay rate. By contrast, short-time correlations between overlapping positions provide a minimal mechanism for the asymmetry by generating direction-dependent subleading relaxation spectra in a coarse-grained Markov description. Together, these results establish a fluctuation-theorem-like diagnostic of irreversibility in financial markets and, more broadly, in complex systems accessible only through coarse-grained observables.

cond-mat.stat-mech

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian languages like Lao. To fill this gap, we introduce \textbf{LaoBench}, the first large-scale, high-quality, and multidimensional benchmark for assessing LLM language understanding and reasoning in Lao. LaoBench contains \textbf{17,000+} expert-curated samples across three dimensions: culturally grounded knowledge application, curriculum-aligned K12 education, and bilingual translation among Lao, Chinese, and English. It includes open-source and held-out subsets, where the held-out portion enables secure black-box evaluation via a controlled service to improve fairness and data security. We construct LaoBench with a hybrid pipeline that combines expert authoring with agent-assisted verification, ensuring linguistic accuracy, cultural relevance, and educational validity. We evaluate diverse state-of-the-art open-source and closed-source LLMs, and find that even strong multilingual models lag behind human experts, particularly in culturally grounded reasoning and translation fidelity. We hope LaoBench will catalyze research on Lao and other underrepresented Southeast Asian languages for more inclusive multilingual evaluation.

cs.CL

Data-driven characterization of spatiotemporal chaos using ensemble reservoir computing

Spatiotemporal chaotic systems are difficult to characterize in a model-free manner because of their high dimensionality, strong nonlinearity, and sensitivity to initial conditions. Coupled map lattices, as a representative class of extended nonlinear systems, exhibit diverse regimes such as frozen random pattern, defect chaotic diffusion, and fully developed turbulence. In this work, we propose an ensemble version of multiplexing local reservoir computing for the data-driven characterization of spatiotemporal chaos. By constructing multiple base learners with randomized hyperparameters and combining their outputs, the method improves prediction robustness and quantifies predictive uncertainty through ensemble spread. More importantly, we show that this uncertainty contains direct dynamical information. It identifies frozen positions in frozen random pattern, supports the estimation of defect diffusion coefficients in defect chaotic diffusion, and provides an effective indicator of chaotic intensity in fully developed turbulence. Analyses of the spatial power spectrum and Lyapunov exponent spectrum further support the consistency between the uncertainty field and the intrinsic dynamical properties of the system. These results show that ensemble reservoir computing can serve not only as a prediction tool but also as a data-driven framework for the dynamical characterization of high-dimensional nonlinear systems.

nlin.CD

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and most fundamental tasks integrated into this paradigm, medical image classification reveals a sobering reality: state-of-the-art medical MLLMs consistently underperform compared to traditional deep learning models, despite their overwhelming advantages in pre-training data and model parameters. This paradox prompts a critical rethinking: where exactly does the performance degradation originate? In this paper, we conduct extensive experiments on 14 open-source medical MLLMs across three representative image classification datasets. Moving beyond superficial performance benchmarking, we employ feature probing to track the information flow of visual features module-by-module and layer-by-layer throughout the entire MLLM pipeline, enabling explicit visualization of where and how classification signals are distorted, diluted, or overridden. As the first attempt to dissect classification performance degradation in medical MLLMs, our findings reveal four failure modes: 1) quality limitation in visual representation, 2) fidelity loss in connector projection, 3) comprehension deficit in LLM reasoning, and 4) misalignment of semantic mapping. Meanwhile, we introduce quantitative scores that characterize the healthiness of feature evolution, enabling principled comparisons across diverse MLLMs and datasets. Furthermore, we provide insightful discussions centered on the critical barriers that prevent current medical MLLMs from fulfilling their promised clinical potential. We hope that our work provokes rethinking within the community-highlighting that the road from high expectations to clinically deployable MLLMs remains long and winding.

cs.CV

Extinction Distributions in Nearby Star-resolved Galaxies. II. M33

Extinction maps are essential for tracing interstellar dust and enabling accurate stellar population studies in galaxies. Here, a high-resolution extinction distribution of nearby galaxy M33 is constructed by fitting multiband color indexes of the individually resolved red giant branch (RGB) stars from the Panchromatic Hubble Andromeda Treasury: Triangulum Extended Region (PHATTER) survey. Achieving an angular resolution of approximately 6$^{\prime\prime}$ ($\sim$ 24.4 pc), the extinction map reveals the intricate and heterogeneous distribution of dust throughout the entire disk of M33, with distinct delineation of spiral arms, inter-arm regions, and compact dust clouds. In addition, it exhibits strong spatial correspondence with the distributions of total hydrogen, H I, and CO, underscoring the reliability of the extinction map for tracing both diffuse and dense components of the interstellar medium. The derived $V$-band extinction reaches up to 2.5 mag per pixel, with a mean value of about 1.05 mag. Beyond providing new insights into the dust structure of M33, the extinction map offers a robust foundation for accurate extinction corrections and will support future studies, including upcoming observations with the Chinese Space Station Telescope.

astro-ph.GA

PowerGenie: Analytically-Guided Evolutionary Discovery of Superior Reconfigurable Power Converters

Discovering superior circuit topologies requires navigating an exponentially large design space-a challenge traditionally reserved for human experts. Existing AI methods either select from predefined templates or generate novel topologies at a limited scale without rigorous verification, leaving large-scale performance-driven discovery underexplored. We present PowerGenie, a framework for automated discovery of higher-performance reconfigurable power converters at scale. PowerGenie introduces: (1) an automated analytical framework that determines converter functionality and theoretical performance limits without component sizing or SPICE simulation, and (2) an evolutionary finetuning method that co-evolves a generative model with its training distribution through fitness selection and uniqueness verification. Unlike existing methods that suffer from mode collapse and overfitting, our approach achieves higher syntax validity, function validity, novelty rate, and figure-of-merit (FoM). PowerGenie discovers a novel 8-mode reconfigurable converter with 23% higher FoM than the best training topology. SPICE simulations confirm average absolute efficiency gains of 10% across 8 modes and up to 17% at a single mode. Code will be released upon publication.

cs.LG

PA-SFM: Tracker-free differentiable acoustic radiation for freehand 3D photoacoustic imaging

Three-dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifacts, which severely limits its clinical flexibility and accessibility. To address this challenge, we present PA-SFM, a tracker-free framework that leverages exclusively single-modality photoacoustic data for both sensor pose recovery and high-fidelity 3D reconstruction via differentiable acoustic radiation modeling. Unlike traditional structure-from-motion (SFM) methods based on visual features, PA-SFM integrates the acoustic wave equation into a differentiable programming pipeline. By leveraging a high-performance, GPU-accelerated acoustic radiation kernel, the framework simultaneously optimizes the 3D photoacoustic source distribution and the sensor array pose via gradient descent. To ensure robust convergence in freehand scenarios, we introduce a coarse-to-fine optimization strategy that incorporates geometric consistency checks and rigid-body constraints to eliminate motion outliers. We validated the proposed method through both numerical simulations and in-vivo rat experiments. The results demonstrate that PA-SFM achieves sub-millimeter positioning accuracy and restores high-resolution 3D vascular structures comparable to ground-truth benchmarks, offering a low-cost, software-defined solution for clinical freehand photoacoustic imaging. The source code is publicly available at \href{https://github.com/JaegerCQ/PA-SFM}{https://github.com/JaegerCQ/PA-SFM}.

cs.CV

Radiative Compression of Dense Cores in the Pillars of Creation as Revealed by JWST Extinction Mapping

The Pillars of Creation in M16 represent an iconic star-forming region where stellar feedback shapes molecular cloud evolution. We present a detailed investigation of dust extinction and density structure in the Pillars of Creation using multiband photometric observations from \emph{JWST} NIRCam. A high-resolution (2\arcsec) extinction map reaching depths of $A_V\sim 100$ mag has been constructed using NIRCam filters F090W, F200W, F335M, and F444W. This map clearly reveals the intricate structure of dense gas within the molecular cloud in the Pillars of Creation region. Analysis of the column density probability distribution function (N-PDF) exhibits a characteristic lognormal distribution at intermediate extinctions ($A_V\approx10-30$\,mag), which transitions to a power-law tail at high extinctions ($A_V\gtrsim$ 30\,mag) where star-forming cores reside. The power-law slope $α$ displays significant spatial variation, steepening from $α\approx 2.0$ at the pillar tips facing the NGC 6611 cluster to $α\approx$4.0 in regions distant from the cluster. This systematic gradient demonstrates that stellar feedback not only disperses molecular clouds but can also locally enhance the formation of dense, self-gravitating structures through radiative compression.

astro-ph.GA