SearcharxivSearch

arXiv subjects

Yue Ling

Publications and source records attributed to Yue Ling.

At least 19 recordsLinked to original sources

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

cs.AI

Life After Benchmark Saturation: A Case Study of CORE-Bench

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such as shortcuts, out-of-distribution generalizability, efficiency, reliability, the relative importance of the model versus the scaffold, and uplift from human-agent collaboration. We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimensions yields meaningful insights into agent performance even after accuracy saturates. First, we surface threats to construct validity in CORE-Bench Hard that are difficult to anticipate with less capable agents. We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD. Second, we find that despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency, reliability, model performance, and scaffold performance. Finally, we conduct a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks. We find a statistically significant speedup by about a factor of two -- likely underestimated due to one-fifth of human-only reproductions reaching the time limit before completing -- and describe various other findings. Together, our contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.

cs.AI

Seedance 2.0: Advancing Video Generation for World Complexity

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.

cs.CV

Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization

We study how different Chain-of-Thought (CoT) designs affect the acquisition of the generalizable visual reasoning ability in vision-language models (VLMs). While CoT data, especially long or visual CoT such as "think with image", has been widely used to supervise intermediate reasoning, it remains unclear why specific CoT designs help and which ones truly support generalizable reasoning. To systematically evaluate this, we focus on a controlled maze-solving benchmark where reasoning rules are fully visual, difficulty can be tuned by grid size, and all the intermediate steps can be automatically generated. Using Qwen2.5-VL-7B under a standard SFT-then-RL pipeline, we compare three representative CoT formats: Language CoT, Grounding CoT (with spatial coordinate trajectories), and Visual CoT (with image manipulations). Our experiments reveal that visual and longer CoT mainly accelerate convergence but do not lift the final performance ceiling; concise CoT containing only essential grounding steps outperforms longer traces; and, strikingly, CoT retaining only the minimal grounding results generalizes best across different maze sizes. We further validate these insights on other vision-centric tasks. These findings highlight a "short is long" effect and provide practical guidance for constructing more generalizable SFT datasets for visual reasoning.

cs.CV

CellStream: Dynamical Optimal Transport Informed Embeddings for Reconstructing Cellular Trajectories from Snapshots Data

Single-cell RNA sequencing (scRNA-seq), especially temporally resolved datasets, enables genome-wide profiling of gene expression dynamics at single-cell resolution across discrete time points. However, current technologies provide only sparse, static snapshots of cell states and are inherently influenced by technical noise, complicating the inference and representation of continuous transcriptional dynamics. Although embedding methods can reduce dimensionality and mitigate technical noise, the majority of existing approaches typically treat trajectory inference separately from embedding construction, often neglecting temporal structure. To address this challenge, here we introduce CellStream, a novel deep learning framework that jointly learns embedding and cellular dynamics from single-cell snapshot data by integrating an autoencoder with unbalanced dynamical optimal transport. Compared to existing methods, CellStream generates dynamics-informed embeddings that robustly capture temporal developmental processes while maintaining high consistency with the underlying data manifold. We demonstrate CellStream's effectiveness on both simulated datasets and real scRNA-seq data, including spatial transcriptomics. Our experiments indicate significant quantitative improvements over state-of-the-art methods in representing cellular trajectories with enhanced temporal coherence and reduced noise sensitivity. Overall, CellStream provides a new tool for learning and representing continuous streams from the noisy, static snapshots of single-cell gene expression.

q-bio.GN

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reinforcement learning with verifiable rewards (RLVR), which typically adopts Pass@1 as the reward, has faced the issues in balancing exploration and exploitation, causing policies to prefer conservative actions, converging to a local optimum. Identifying an appropriate reward metric is therefore crucial. Regarding the prior work, although Pass@k has been used in evaluation, its connection to LLM exploration ability in RLVR remains largely overlooked. To investigate this, we first use Pass@k as the reward to train the policy model (i.e., $\textbf{Pass@k Training}$), and observe the improvement on its exploration ability. Next, we derive an analytical solution for the advantage of Pass@k Training, leading to an efficient and effective process. Building on this, our analysis reveals that exploration and exploitation are not inherently conflicting objectives, while they can mutually enhance each other. Moreover, Pass@k Training with analytical derivation essentially involves directly designing the advantage function. Inspired by this, we preliminarily explore the advantage design for RLVR, showing promising results and highlighting a potential future direction.

cs.LG

Seed1.5-VL Technical Report

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B active parameters. Despite its relatively compact architecture, it delivers strong performance across a wide spectrum of public VLM benchmarks and internal evaluation suites, achieving the state-of-the-art performance on 38 out of 60 public benchmarks. Moreover, in agent-centric tasks such as GUI control and gameplay, Seed1.5-VL outperforms leading multimodal systems, including OpenAI CUA and Claude 3.7. Beyond visual and video understanding, it also demonstrates strong reasoning abilities, making it particularly effective for multimodal reasoning challenges such as visual puzzles. We believe these capabilities will empower broader applications across diverse tasks. In this report, we mainly provide a comprehensive review of our experiences in building Seed1.5-VL across model design, data construction, and training at various stages, hoping that this report can inspire further research. Seed1.5-VL is now accessible at https://www.volcengine.com/ (Volcano Engine Model ID: doubao-1-5-thinking-vision-pro-250428)

cs.CV

Exact computation of the color function for triangular element interfaces

The calculation of the volume enclosed by curved surfaces discretized into triangular elements, and a cube is of great importance in different domains, such as computer graphics and multiphase flow simulations. We propose a robust algorithm, the Front2VOF (F2V) algorithm, to address this problem. The F2V algorithm consists of two main steps. First, it identifies the polygons within the cube by segmenting the triangular elements on the surface, retaining only the portions inside the cube boundaries. Second, it computes the volume enclosed by these polygons in combination with the cube faces. To validate the algorithm's accuracy and robustness, we tested it using a range of synthetic configurations with known analytical solutions.

cs.GR

Impact of vaporization on drop aerobreakup

Aerodynamic breakup of vaporizing drops is commonly seen in many spray applications. While it is well known that vaporization can modulate interfacial instabilities, the impact of vaporization on drop aerobreakup is poorly understood. Detailed interface-resolved simulations were performed to systematically study the effect of vaporization, characterized by the Stefan number, on the drop breakup and acceleration for different Weber numbers and density ratios. It is observed that the resulting asymmetric vaporization rates and strengths of Stefan flow on the windward and leeward sides of the drop hinder bag development and prevent drop breakup. The critical Weber number thus generally increases with the Stefan number. The modulation of the boundary layer also contributes to a significant increase of drag coefficient. Numerical experiments were performed to affirm that the drop volume reduction plays a negligible role and the Stefan flow is the dominant reason for the breakup suppression and drag enhancement observed.

physics.flu-dyn

Effect of gas viscosity on the interfacial instability development in a two-phase mixing layer

The interfacial instability in a two-phase mixing layers between parallel gas and liquid streams is important to two-phase atomization. Depending on the inflow conditions and fluid properties, interfacial instability can be convective or absolute. The goal of the present study is to investigate the impact of gas viscosity on the interfacial instability. Both interface-resolved simulations and linear stability analysis (LSA) have been conducted. In LSA, the Orr-Sommerfeld equation is solved to analyze the spatio-temporal viscous modes. When the gas viscosity decreases, the Reynold number ($\text{Re}$) increases accordingly. The LSA demonstrates that when $\text{Re}$ is higher than a critical threshold, the instability transitions from the absolute to the convective (A/C) regimes. Such a $\text{Re}$-induced A/C transition is also observed in the numerical simulations, though the critical Re observed in simulations is significantly lower than that predicted by LSA. The LSA results indicate that the temporal growth rate decreases with Re. When the growth rate reaches zero, the A/C transition will occur. The $\text{Re}$-induced A/C transition is observed in both confined and unconfined mixing layers and also in cases with low and high gas-to-liquid density ratios. In the transition from typical absolute and convective regimes, a weak absolute regime is identified in the simulations, for which \tcr{the spectrograms} show both the absolute and convective modes. The dominant frequency in the weak absolute regime can be influenced by the perturbation introduced at the inlet. The simulation results also show that the wave propagation speed can vary in space. In the absolute instability regime, the wave propagation speed agrees well with the absolute mode celerity near the inlet and increases to the Dimotakis speed further downstream.

physics.flu-dyn

G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition

Millimeter wave radar is gaining traction recently as a promising modality for enabling pervasive and privacy-preserving gesture recognition. However, the lack of rich and fine-grained radar datasets hinders progress in developing generalized deep learning models for gesture recognition across various user postures (e.g., standing, sitting), positions, and scenes. To remedy this, we resort to designing a software pipeline that exploits wealthy 2D videos to generate realistic radar data, but it needs to address the challenge of simulating diversified and fine-grained reflection properties of user gestures. To this end, we design G3R with three key components: (i) a gesture reflection point generator expands the arm's skeleton points to form human reflection points; (ii) a signal simulation model simulates the multipath reflection and attenuation of radar signals to output the human intensity map; (iii) an encoder-decoder model combines a sampling module and a fitting module to address the differences in number and distribution of points between generated and real-world radar data for generating realistic radar data. We implement and evaluate G3R using 2D videos from public data sources and self-collected real-world radar data, demonstrating its superiority over other state-of-the-art approaches for gesture recognition.

cs.MM

A detailed numerical investigation of two-phase flows inside a planar flow-blurring atomizer

Flow-blurring atomization is an innovative twin-fluid atomization approach that has demonstrated superior effectiveness in producing fine sprays compared to traditional airblast atomization methods. In flow-blurring atomizers, the high-speed gas flow is directed perpendicular to the liquid jet. Under specific geometric and physical conditions, the gas penetrates back into the liquid nozzle, resulting in a highly unsteady bubbly two-phase mixing zone. Despite the remarkable atomization performance of flow-blurring atomizers, the underlying dynamics of the two-phase flows and breakup mechanisms within the liquid nozzle remain poorly understood, primarily due to the challenges in experimental measurements of flow details. In this study, detailed interface-resolved numerical simulations are conducted to investigate the two-phase flows generated by a planar flow-blurring atomizer. By varying key dimensionless parameters, including the dynamic-pressure ratio, density ratio, and Weber number, over wide ranges, we aim to comprehensively characterize their effects on the two-phase flow regimes and breakup dynamics.

physics.flu-dyn

Direct numerical simulation of compressible interfacial multiphase flows using a mass-momentum-energy consistent volume-of-fluid method

Compressible interfacial multiphase flows (CIMF) are essential to different applications, such as liquid fuel injection in supersonic propulsion systems. Since high-level details in CIMF are often difficult to measure in experiments, numerical simulation is an important alternative to shed light on the unclear physics. A direct numerical simulation (DNS) of CIMF will need to rigorously resolve the shock waves, the interfaces, and the interaction between the two. A novel numerical method has been developed and implemented in the present study. The geometric volume-of-fluid (VOF) method is employed to resolve the sharp interfaces between the two phases. The advection of the density, momentum, and energy is carried out consistently with VOF advection. To suppress spurious oscillations near shocks, numerical diffusion is introduced based on the Kurganov-Tadmor method in the region away from the interface. The contribution of pressure is incorporated using the projection method and the pressure is obtained by solving the Poisson-Helmholtz equation, which allows the present method to handle flows with all Mach numbers. The present method is tested by a sequence of CIMF problems. The simulation results are validated against theories, experiments, and other simulations, and excellent agreement has been achieved. In particular, the linear single-mode Richtmyer-Meshkov instabilities with finite Weber and Reynolds numbers are simulated. The simulation results agree very well with the linear stability theory, which affirms the capability of the present method in capturing the viscous and capillary effects on shock-interface interaction.

physics.flu-dyn

A model to predict the oscillation frequency for drops pinned on a vertical planar surface

Accurate prediction of the natural frequency for the lateral oscillation of a liquid drop pinned on a vertical planar surface is important to many drop applications. The natural oscillation frequency, normalized by the capillary frequency, is {mainly} a function of the equilibrium contact angle and the Bond number (Bo), when the contact lines remain pinned. Parametric numerical and experimental studies have been performed to establish a comprehensive understanding of oscillation dynamics. {An} inviscid model has been developed to predict the oscillation frequency for wide ranges of Bo and contact angles. The model reveals the scaling relation between the normalized frequency and Bo, which is validated by the numerical simulation results. For a given equilibrium contact angle, the lateral oscillation frequency decreases with Bo, implying that resonance frequencies will be magnified if the drop oscillations occur in a reduced gravity environment.

physics.flu-dyn

Impact of inlet gas turbulence on the formation, development and breakup of interfacial waves in a two-phase mixing layer

Understanding the development and breakup of interfacial waves in a two-phase mixing layer between the gas and liquid streams is paramount to atomization. Due to the velocity difference between the two streams, the shear on the interface triggers a longitudinal instability, which develops to interfacial waves that propagate downstream. As the interfacial waves grow spatially, transverse modulations arise, turning the interfacial waves from quasi-2D to fully 3D. The inlet gas turbulence intensity has a strong impact on the interfacial instability. Therefore, parametric direct numerical simulations are performed in the present study to systematically investigate the effect of the inlet gas turbulence on the formation, development, and breakup of the interfacial waves. The open-source multiphase flow solver, PARIS, is used for the simulations and the mass-momentum consistent volume-of-fluid method is used to capture the sharp gas-liquid interfaces. Two computational domain widths are considered and the wide domain will allow a detailed study of the transverse development of the interfacial waves. The dominant frequency and spatial growth rate of the longitudinal instability are found to increase with the inlet gas turbulence intensity. The dominant transverse wavenumber, determined by the Rayleigh-Taylor instability, scales with the longitudinal frequency, so it also increases with the inlet gas turbulence intensity. The holes formed in the liquid sheet is important to the disintegration of the interfacial waves. The holes formation is influenced by the inlet gas turbulence. As a result, the sheet breakup dynamics and the statistics of the droplets formed also change accordingly.

physics.flu-dyn

Numerical study of natural oscillations of supported drops with free and pinned contact lines

In the present study, the axisymmetric natural oscillations of a liquid drop supported by a flat surface is investigated by direct numerical simulation. The liquid-gas interface is captured using a geometric volume-of-fluid (VOF) method. A parametric study is carried out by varying the equilibrium contact angle and the gravitational Bond number (Bo). Both positive and negative gravities are considered, and thus the results cover both pendant and sessile drops. To incorporate the effect of contact line mobility, the two asymptotic limits, namely the pinned contact line (PCL) and free contact line (FCL) conditions, are considered and their effects on the drop oscillation features are characterized. The predicted oscillation frequencies for PCL and FCL serve as the upper and lower bounds for general situations. The drop oscillation is initiated by increasing the gravity magnitude for a short time. The first mode due to the drop centroid translation dominates the excited oscillation. The oscillation frequency scales with the capillary frequency, and the normalized frequency monotonically decreases with the equilibrium contact angle. For zero gravity, the computed frequencies for all contact angles agree remarkably well with the inviscid theory for both the PCL and FCL conditions. The kinetic energy correction factor is introduced to account for the additional contribution of the oscillation-induced internal flow to the overall kinetic energy of the drop. Both the frequency and the kinetic energy correction factor increase with Bo, decrease with the contact angle, and increase when the contact line condition changed from FCL to PCL. The variation of oscillation frequency due to the change of Bo is particularly significant when the contact angle is large.

physics.flu-dyn

Simulation and modeling of the vaporization of a freely moving and deforming drop at low to moderate Weber numbers

The vaporization of a freely moving drop in a uniform, high-temperature gas stream is investigated through direct numerical simulation. The incompressible Navier-Stokes equations with surface tension and phase change are solved in conjunction with the energy equations of each phase. The sharp liquid-gas interface is tracked using the geometric Volume-of-Fluid (VOF) method and an immersed Dirichlet boundary condition for temperature is imposed at the interface. The simulation approach is validated by simulating water and acetone drops at nearly zero Weber numbers, and the simulation results agree very well with the empirical relation for spherical drops. Parametric simulations were conducted to investigate the aerodynamic breakup of vaporizing drops at low to moderate Weber and Reynolds numbers. The range of Weber numbers considered has covered the vibrational and bag breakup regimes. Through the simulation results, we have characterized the impact of drop deformation and breakup on the drop vaporization rate. When the drop Weber number increases, the windward surface area increases more rapidly over time. As a result, the rate of drop volume reduction also increases. The correlation between the vaporization rate and the windward surface area is examined for different Weber and Reynolds numbers. Using the approximate correlation between the drop vaporization rate and the windward surface area and the TAB model for drop deformation, a new time-dependent drop vaporization model is proposed. The present model agrees well with the simulation results and shows a significant improvement over the conventional model for spherical drops.

physics.flu-dyn

Bio+Clinical BERT, BERT Base, and CNN Performance Comparison for Predicting Drug-Review Satisfaction

The objective of this study is to develop natural language processing (NLP) models that can analyze patients' drug reviews and accurately classify their satisfaction levels as positive, neutral, or negative. Such models would reduce the workload of healthcare professionals and provide greater insight into patients' quality of life, which is a critical indicator of treatment effectiveness. To achieve this, we implemented and evaluated several classification models, including a BERT base model, Bio+Clinical BERT, and a simpler CNN. Results indicate that the medical domain-specific Bio+Clinical BERT model significantly outperformed the general domain base BERT model, achieving macro f1 and recall score improvement of 11%, as shown in Table 2. Future research could explore how to capitalize on the specific strengths of each model. Bio+Clinical BERT excels in overall performance, particularly with medical jargon, while the simpler CNN demonstrates the ability to identify crucial words and accurately classify sentiment in texts with conflicting sentiments.

cs.CL