SearcharxivSearch

arXiv subjects

Meng Tian

Publications and source records attributed to Meng Tian.

At least 19 recordsLinked to original sources

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visually plausible yet 4D inconsistent future predictions that mislead downstream planning. To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Specifically, we feed WAM-predicted future frames into a geometric foundation model, and use 4D-aware responses to define a 4D consistency loss. This loss encourages the model to understand, represent, and predict physically consistent 4D scenes during training, without additional inference cost. Moreover, we identify an early-decision phenomenon in WAMs and propose a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages, where driving decisions are primarily formed. By propagating 4D supervision to this critical decision-formation phase, the proposed strategy further improves trajectory planning. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.

cs.CV

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.

cs.CV

Lithium niobate quadratic integrated nonlinear photonics: enabling ultra-wide bandwidth and ultrafast photonic engines

Integrated photonic coherent light sources capable of generating emission with broad spectral coverage and ultrashort pulse durations are critical for both fundamental science and emerging technologies. In this Perspective, we start by discussing emerging quantum and classical photonic applications from the standpoint of operating wavelength and timescale, highlighting the technological gaps that persist in current integrated photonic light sources. Next, we introduce the unique properties of lithium niobate-based integrated quadratic nonlinear photonics, and discuss several promising strategies that exploit this platform to realize wavelength-tunable continuous wave light sources and broadband, ultra-short light pulse generation. We also assessed their advantages and limitations while discussing potential solutions. Finally, we outline future prospects and challenges that need to be addressed, aiming at inspiring continued research and innovation in this rapidly evolving field.

physics.optics

On-chip electrically reconfigurable octave-bandwidth optical amplification from visible to near-infrared

Achieving broadband on-chip optical amplification spanning the visible and near-infrared (NIR) can enable diverse quantum sensing, metrology, and classical communication applications within a single unified device. However, conventional semiconductor and ion-doped amplifiers suffer from limited gain bandwidths set by fixed energy levels, while optical parametric amplifiers (OPAs) operating continuously from the visible to the NIR have remained elusive due to dispersion-limited bandwidth and the high pump powers required in the visible or ultraviolet (UV). Here, we overcome these limitations by introducing an electrically reconfigurable OPA architecture on lithium niobate integrated photonics. By synergistically combining ultra-high effective $\chi^{(2)}$ nonlinearity ($\sim$7,000\%/W-cm$^2$), high-order dispersion engineering, and local electro-thermal tuning of quasi-phase matching, our device achieves record gain spectral spanning more than an optical octave, from 770 to 1650 nm. This range covers key transitions of many photonic quantum systems and all telecommunication bands. Moreover, our approach eliminates the need for high-power, wavelength-tunable visible or UV pumps, delivering a peak on-chip gain of 23.67 dB with a single 1060 nm pump at 90 mW average on-chip power. This work opens new avenues for multi-functional, reconfigurable photonics unifying the visible and infrared regimes, with broad implications for quantum sensing and communications.

physics.optics

SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL

General-purpose Large Vision-Language Models (LVLMs), despite their massive scale, often falter in dermatology due to "diffuse attention" - the inability to disentangle subtle pathological lesions from background noise. In this paper, we challenge the assumption that parameter scaling is the only path to medical precision. We introduce SkinFlow, a framework that treats diagnosis as an optimization of visual information transmission efficiency. Our approach utilizes a Virtual-Width Dynamic Vision Encoder (DVE) to "unfold" complex pathological manifolds without physical parameter expansion, coupled with a two-stage Reinforcement Learning strategy. This strategy sequentially aligns explicit medical descriptions (Stage I) and reconstructs implicit diagnostic textures (Stage II) within a constrained semantic space. Furthermore, we propose a clinically grounded evaluation protocol that prioritizes diagnostic safety and hierarchical relevance over rigid label matching. Empirical results are compelling: our 7B model establishes a new state-of-the-art on the Fitzpatrick17k benchmark, achieving a +12.06% gain in Top-1 accuracy and a +28.57% boost in Top-6 accuracy over the massive general-purpose models (e.g., Qwen3VL-235B and GPT-5.2). These findings demonstrate that optimizing geometric capacity and information flow yields superior diagnostic reasoning compared to raw parameter scaling.

cs.CV

JWST NIRSpec finds no clear signs of an atmosphere on TOI-1685 b

Determining the prevalence of atmospheres on terrestrial planets is a core objective in exoplanetary science. While M dwarf systems offer a promising opportunity, conclusive observations of terrestrial atmospheres have remained elusive, with many yielding flat transmission spectra. We observe four transits of the hot terrestrial planet TOI-1685 b using JWST's NIRSpec G395H instrument. Combining this with the transit from the previously-observed phase curve of the planet with the same instrument, we perform a detailed analysis to determine the possibility of an atmosphere on TOI-1685 b. From our retrievals, the Bayesian evidence favours a simple flat line model, indicating no evidence for an atmosphere on TOI-1685 b, in line with results from the phase curve analysis. Our results show that hydrogen-dominated atmospheres can be confidently ruled out. For heavier, secondary atmospheres we find a lower limit on the mean molecular weight of ~10, at a significance of ~5 sigma. Pure CO2, SO2, H2O, and CH4 atmospheres, or a mixed secondary atmosphere (CO+CO2+SO2) could explain the data (Delta lnZ < 3). However, pure CH4 atmospheres may be physically unlikely, and the pure H2O and CO2 cases require a high-altitude cloud, which could also be interpreted as a thin cloud-free atmosphere. We discuss the theoretical possibility for different types of atmosphere on this planet, and consider the effects of atmospheric escape and stellar activity on the system. Though we find that TOI-1685 b is likely a bare rock, this study also highlights the challenges of detecting secondary atmospheres on rocky planets with JWST.

astro-ph.EP

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems built on them therefore show limited perception and localization ability. To address these challenges, we introduce Percept-WAM, a perception-enhanced World-Awareness-Action Model that is the first to implicitly integrate 2D/3D scene understanding abilities within a single vision-language model (VLM). Instead of relying on QA-style spatial reasoning, Percept-WAM unifies 2D/3D perception tasks into World-PV and World-BEV tokens, which encode both spatial coordinates and confidence. We propose a grid-conditioned prediction mechanism for dense object perception, incorporating IoU-aware scoring and parallel autoregressive decoding, improving stability in long-tail, far-range, and small-object scenarios. Additionally, Percept-WAM leverages pretrained VLM parameters to retain general intelligence (e.g., logical reasoning) and can output perception results and trajectory control outputs directly. Experiments show that Percept-WAM matches or surpasses classical detectors and segmenters on downstream perception benchmarks, achieving 51.7/58.9 mAP on COCO 2D detection and nuScenes BEV 3D detection. When integrated with trajectory decoders, it further improves planning performance on nuScenes and NAVSIM, e.g., surpassing DiffusionDrive by 2.1 in PMDS on NAVSIM. Qualitative results further highlight its strong open-vocabulary and long-tail generalization.

cs.CV

Passive harmonic mode-locked laser on lithium niobate integrated photonics

Mode-locked lasers (MLLs) are essential for a wide range of photonic applications, such as frequency metrology, biological imaging, and high-bandwidth coherent communications. The growing demand for compact and scalable photonic systems is driving the development of MLLs on various integrated photonics material platforms. Along these lines, developing MLLs on the emerging thin-film lithium niobate (TFLN) platform holds the promise to greatly broaden the application space of MLLs by harnessing TFLN 's unique electro-optic (E-O) response and quadratic optical nonlinearity. Here, we demonstrate the first electrically pumped, self-starting passive MLL in lithium niobate integrated photonics based on its hybrid integration with a GaAs quantum-well gain medium and saturable absorber. Our demonstrated MLL generates 4.3-ps optical pulses centered around 1060 nm with on-chip peak power exceeding 44 mW. The pulse duration can be further compressed to 1.75 ps via linear dispersion compensation. Remarkably, passive mode-locking occurs exclusively at the second harmonic of the cavity free spectral range, exhibiting a high pulse repetition rate $\sim$20 GHz. We elucidate the temporal dynamics underlying this self-starting passive harmonic mode-locking behavior using a traveling-wave model. Our work offers new insights into the realization of compact, high-repetition-rate MLLs in the TFLN platform, with promising applications for monolithic ultrafast microwave waveform sampling and analog-to-digital conversion.

physics.optics

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To address these challenges, we propose StreamForest, a novel architecture specifically designed for streaming video understanding. Central to StreamForest is the Persistent Event Memory Forest, a memory mechanism that adaptively organizes video frames into multiple event-level tree structures. This process is guided by penalty functions based on temporal distance, content similarity, and merge frequency, enabling efficient long-term memory retention under limited computational resources. To enhance real-time perception, we introduce a Fine-grained Spatiotemporal Window, which captures detailed short-term visual cues to improve current scene perception. Additionally, we present OnlineIT, an instruction-tuning dataset tailored for streaming video tasks. OnlineIT significantly boosts MLLM performance in both real-time perception and future prediction. To evaluate generalization in practical applications, we introduce ODV-Bench, a new benchmark focused on real-time streaming video understanding in autonomous driving scenarios. Experimental results demonstrate that StreamForest achieves the state-of-the-art performance, with accuracies of 77.3% on StreamingBench, 60.5% on OVBench, and 55.6% on OVO-Bench. In particular, even under extreme visual token compression (limited to 1024 tokens), the model retains 96.8% of its average accuracy in eight benchmarks relative to the default setting. These results underscore the robustness, efficiency, and generalizability of StreamForest for streaming video understanding.

cs.CV

Diversity of low-mass planet atmospheres in the C-H-O-N-S-Cl system with interior dissolution, nonideality, and condensation: Application to TRAPPIST-1e and sub-Neptunes

A quantitative understanding of the nature and composition of low-mass rocky (exo)planet atmospheres during their evolution is needed to interpret observations. The magma ocean stage of terrestrial- and sub-Neptune planets permits mass exchange between their interiors and atmospheres, during which the mass and speciation of the atmosphere is dictated by the planet's volatile budget, chemical equilibria, and gas/fluid solubility in molten rock. As the atmosphere cools, it is modified by gas-phase reactions and condensation. We combine these processes into an open-source Python package built using JAX called Atmodeller, and perform calculations for planet sizes and conditions analogous to TRAPPIST-1e and K2-18b. For TRAPPIST-1e-like planets, our simulations indicate that CO-dominated atmospheres are prevalent during the magma ocean stage, which, upon isochemical cooling, predominantly evolve into CO2-rich atmospheres of a few hundred bar at 280 K. Around 40% of our simulations predict the coexistence of liquid water, graphite, alpha-sulfur, and ammonium chloride, which are key ingredients for surface habitability. For sub-Neptune gas dwarfs, pressures are sufficiently high (a few GPa) that gas fugacities deviate from ideality, thereby drastically enhancing solubilities. This buffers the total atmospheric pressure to lower values than for the ideal case. These effects conspire to produce CH4-rich sub-Neptune atmospheres for total pressures exceeding around 3.5 GPa, provided H/C is approximately 100x solar and fO2 moderately reducing (3 log10 units below the iron-wustite buffer). Otherwise, molecular hydrogen remains the predominant species at lower total pressures and/or higher H/C. For all planets at high temperature, solubility enriches C/H in the atmosphere relative to the initial composition.

astro-ph.EP

Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by relying heavily on history input information, achieving seemingly strong planning results without genuinely understanding the visual inputs; and (2) the chain-ofthought (COT) reasoning processes are always misaligned with the motion planning outcomes, and how to effectively leverage the complex reasoning capability to enhance planning remains largely underexplored. In this paper, we start from a small-scale domain-specific VLM and propose Drive-R1 designed to bridges the scenario reasoning and motion planning for AD. Drive-R1 first undergoes the supervised finetuning on a elaborate dataset containing both long and short COT data. Drive-R1 is encouraged to reason step-by-step from visual input to final planning decisions. Subsequently, Drive-R1 is trained within a reinforcement learning framework that incentivizes the discovery of reasoning paths that are more informative for planning, guided by rewards based on predicted trajectories and meta actions. Experimental evaluations on the nuScenes and DriveLM-nuScenes benchmarks demonstrate that Drive-R1 achieves superior performance compared to existing state-of-the-art VLMs. We believe that Drive-R1 presents a promising direction for bridging reasoning and planning in AD, offering methodological insights for future research and applications.

cs.CV

Interior redox state effects on the stability of secondary atmospheres and observational manifestations: LP 791-18 d as a case study for outgassing rocky exoplanets

Recent advances in space and ground-based facilities now enable atmospheric characterization of a selected sample of rocky exoplanets. These atmospheres offer key insights into planetary formation and evolution, but their interpretation requires models that couple atmospheric processes with both the planetary interior and the surrounding space environment. This work focuses on the Earth-size planet LP791 18d, which is estimated to receive continuous tidal heating due to the orbital configuration of the system; thus, it is expected to exhibit volcanic activity. We estimate the mantle temperature of 1680-1880 K. Our results show that the atmospheric mean molecular weight gradient is controlled by oxygen fugacity rather than bulk metallicity. Furthermore, we use the atmospheric steady-state solutions produced from the interior redox state versus surface pressure parameter space and explore their atmospheric stability. We find that stability is achieved only in highly oxidized scenarios while reduced interior states fall into the hydrodynamic escape regime with mass loss rates on the order of 10^5-10^8 kg/s. We argue that scenarios with reduced interior states are likely to have exhausted their volatile budget during the planets lifetime. Furthermore, we predict the atmospheric footprint of the planets interior based on its oxidation state and assess its detectability using current or forthcoming tools to constrain the internal and atmospheric composition. We show that the degeneracy between bare rock surfaces and thick atmospheres can be resolved by using three photometric bands to construct a color-color diagram that accounts for potential effects from photochemical hazes and clouds. Our modeling approach connects interior and atmospheric processes, providing a basis to explore volatile evolution and potential habitability.

astro-ph.EP

The Gradient of Mean Molecular Weight Across the Radius Valley

Photo-evaporation shapes the observed radii of small exoplanets and constrains the underlying distributions of atmospheric and core masses. However, the diversity of atmospheric chemistries corresponding to these distributions remains unelucidated. We develop a first-principles carbon-hydrogen-oxygen-sulfur-silicon (CHOSSi) outgassing model that accounts for non-ideal gas behavior (via fugacities) at high pressures, as well as the tendency for water and hydrogen to dissolve in melt (via solubility laws). We use data-driven radius valley constraints to establish the relationship between the atmospheric surface pressures and melt temperatures of sub-Neptunes. Sub-Neptunes with less massive rocky cores retain less of their primordial hydrogen envelopes, which leads to less heat retention and diminished melt temperatures at the surfaces of these cores. Lower melt temperatures lead thermodynamically to the dominance of carbon-, oxygen-, sulfur- and silicon-bearing molecules over molecular hydrogen, which naturally produce a diversity of mean molecular weights. Our geochemical outgassing calculations robustly predict a gradient of mean molecular weight across the radius valley, where the strength of this gradient is primarily driven by the oxygen fugacity of the molten cores and not by the carbon enrichment (or "metallicity") of the atmosphere. Smaller sub-Neptunes are predicted to have less hydrogen-dominated atmospheres. The precise relationship between the observed and outgassed chemistries requires an understanding of how convection near the core interacts with large-scale atmospheric circulation (driven by stellar heating) near the photosphere, as well as the influence of photochemistry.

astro-ph.EP

Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce $\textbf{VLADBench}$, a challenging and fine-grained dataset featuring close-form QAs that progress from static foundational knowledge and elements to advanced reasoning for dynamic on-road situations. The elaborate $\textbf{VLADBench}$ spans 5 key domains: Traffic Knowledge Understanding, General Element Recognition, Traffic Graph Generation, Target Attribute Comprehension, and Ego Decision-Making and Planning. These domains are further broken down into 11 secondary aspects and 29 tertiary tasks for a granular evaluation. A thorough assessment of general and domain-specific (DS) VLMs on this benchmark reveals both their strengths and critical limitations in AD contexts. To further exploit the cognitive and reasoning interactions among the 5 domains for AD understanding, we start from a small-scale VLM and train the DS models on individual domain datasets (collected from 1.4M DS QAs across public sources). The experimental results demonstrate that the proposed benchmark provides a crucial step toward a more comprehensive assessment of VLMs in AD, paving the way for the development of more cognitively sophisticated and reasoning-capable AD systems.

cs.CL

Electrically Reconfigurable Intelligent Optoelectronics in 2-D van der Waals Materials

In optoelectronics, achieving electrical reconfigurability is crucial as it enables the encoding, decoding, manipulating, and processing of information carried by light. In recent years, two-dimensional van der Waals (2-D vdW) materials have emerged as promising platforms for realizing reconfigurable optoelectronic devices. Compared to materials with bulk crystalline lattice, 2-D vdW materials offer superior electrical reconfigurability due to high surface-to-volume ratio, quantum confinement, reduced dielectric screening effect, and strong dipole resonances. Additionally, their unique band structures and associated topology and quantum geometry provide novel tuning capabilities. This review article seeks to establish a connection between the fundamental physics underlying reconfigurable optoelectronics in 2-D materials and their burgeoning applications in intelligent optoelectronics. We first survey various electrically reconfigurable properties of 2-D vdW materials and the underlying tuning mechanisms. Then we highlight the emerging applications of such devices, including dynamic intensity, phase and polarization control, and intelligent sensing. Finally, we discuss the opportunities for future advancements in this field.

physics.optics

Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Large Vision-Language Models (LVLMs) have received widespread attention for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated and quantifiable assessment for self-driving, let alone the severe road corner cases. In this work, we propose CODA-LM, the very first benchmark for the automatic evaluation of LVLMs for self-driving corner cases. We adopt a hierarchical data structure and prompt powerful LVLMs to analyze complex driving scenes and generate high-quality pre-annotations for the human annotators, while for LVLM evaluation, we show that using the text-only large language models (LLMs) as judges reveals even better alignment with human preferences than the LVLM judges. Moreover, with our CODA-LM, we build CODA-VLM, a new driving LVLM surpassing all open-sourced counterparts on CODA-LM. Our CODA-VLM performs comparably with GPT-4V, even surpassing GPT-4V by +21.42% on the regional perception task. We hope CODA-LM can become the catalyst to promote interpretable self-driving empowered by LVLMs.

cs.CV

Knowledge Augmented Relation Inference for Group Activity Recognition

Most existing group activity recognition methods construct spatial-temporal relations merely based on visual representation. Some methods introduce extra knowledge, such as action labels, to build semantic relations and use them to refine the visual presentation. However, the knowledge they explored just stay at the semantic-level, which is insufficient for pursing notable accuracy. In this paper, we propose to exploit knowledge concretization for the group activity recognition, and develop a novel Knowledge Augmented Relation Inference framework that can effectively use the concretized knowledge to improve the individual representations. Specifically, the framework consists of a Visual Representation Module to extract individual appearance features, a Knowledge Augmented Semantic Relation Module explore semantic representations of individual actions, and a Knowledge-Semantic-Visual Interaction Module aims to integrate visual and semantic information by the knowledge. Benefiting from these modules, the proposed framework can utilize knowledge to enhance the relation inference process and the individual representations, thus improving the performance of group activity recognition. Experimental results on two public datasets show that the proposed framework achieves competitive performance compared with state-of-the-art methods.

cs.CV

Atmospheric Chemistry of Secondary and Hybrid Atmospheres of Super Earths and Sub-Neptunes

The atmospheres of small exoplanets likely derive from a combination of geochemical outgassing and primordial gases left over from formation. Secondary atmospheres, such as those of Earth, Mars and Venus, are sourced by outgassing. Persistent outgassing into long-lived, primordial, hydrogen-helium envelopes produces hybrid atmospheres of which there are no examples in the Solar System. We construct a unified theoretical framework for calculating the outgassing chemistry of both secondary and hybrid atmospheres, where the input parameters are the surface pressure, oxidation and sulfidation states of the mantle, as well as the primordial atmospheric hydrogen, helium and nitrogen content. Non-ideal gases (quantified by the fugacity coefficient) and non-ideal mixing of gaseous components (quantified by the activity coefficient) are considered. Both secondary and hybrid atmospheres exhibit a rich diversity of chemistries, including hydrogen-dominated atmospheres. The abundance ratio of carbon dioxide to carbon monoxide serves as a powerful diagnostic for the oxygen fugacity of the mantle, which may conceivably be constrained by James Webb Space Telescope spectra in the near future. Methane-dominated atmospheres are difficult to produce and require specific conditions: atmospheric surface pressures exceeding $\sim 10$ bar, a reduced (poorly oxidised) mantle and diminished magma temperatures (compared to modern Earth). Future work should include photochemistry in these calculations and clarify the general role of atmospheric escape. Exoplanet science should quantify the relationship between the mass and oxygen fugacity for a sample of super Earths and sub-Neptunes; such an empirical relationship already exists for Solar System bodies.

astro-ph.EP