SearcharxivSearch

arXiv subjects

Yao Tang

Publications and source records attributed to Yao Tang.

At least 19 recordsLinked to original sources

GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation

We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2026. Given synchronized egocentric stereo views, the task is to predict a 3D vector from each of 21 hand joints to the closest point on the manipulated object. GOLF combines dense global context, locally sampled hand/object evidence, and common-frame Pl\"ucker-ray geometry. We adapt DINOv3 ViT-H+/16 with LoRA and trainable LayerNorm parameters, then jointly decode both interaction fields. Our primary model achieves an official score of 27.61 and a mean ADE of 27.96 mm on the hidden test set. An equal-weight ensemble with a complementary directly fine-tuned variant improves these results to an official score of 27.47 and a mean ADE of 27.82 mm, securing first place.

cs.CV

EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation

Diffusion-based visual generative models deliver strong image and video synthesis quality but incur high inference costs because sequential samplers repeatedly evaluate large networks. Caching-based methods reduce inference latency by reusing intermediate computations across adjacent timesteps. However, existing cache controllers rely primarily on local temporal variation and overlook the trajectory-level consequences of cache reuse. We introduce Error-Propagation-Aware Cache (EpaCache), a training-free caching policy that adaptively allocates the reuse budget on timesteps with lower downstream impact. Experiments on image and video synthesis models demonstrate that EpaCache consistently improves the latency--fidelity trade-off over existing caching methods. On FLUX.1-dev, EpaCache outperforms the prior state-of-the-art caching method in both latency and fidelity, reducing inference time from $11.7$ s to $11.3$ s while improving PSNR from $21.4$ to $22.8$. On HunyuanVideo, EpaCache achieves a $2.63\times$ speedup over uncached inference and improves SSIM from $0.891$ to $0.905$ over the prior state-of-the-art method at matched latency.

cs.AI

An Evolving Cosmic Shoreline and Sandbar Bounding the Rocky Airless Valley

Recent JWST observations challenge the traditional 'cosmic shoreline' from both sides, revealing thick volatile atmospheres on the hottest close-in 'lava worlds,' where irradiation should drive the most extreme escape, and bare rocky surfaces on cooler terrestrial planets around M dwarfs, where atmospheres would be expected to survive. Using a coupled atmosphere-interior evolution model, we show that atmosphere retention is governed not by a single escape boundary but by two: a hot, outgassing-regulated 'cosmic sandbar' and a cooler, escape-regulated 'cosmic shoreline,' separated by an 'airless valley' that may mark a graveyard of stripped sub-Neptune cores. The sandbar arises because long-lived magma oceans, sustained further by tidal heating from secular eccentricity excitation in multi-planet systems, keep most volatiles dissolved and expose only a small atmospheric reservoir to escape, whereas cooler planets solidify, sequestering volatiles in the deep solid mantle while overexposing the rest to loss. This two-regime structure recasts the single cosmic shoreline as two boundaries set by distinct physics: outgassing and escape. We provide time-evolving fits for both boundaries across G, K, and M stellar types as a function of volatile inventory, planetary mass, age, and tidal heating. Lava worlds with thick atmospheres are unlikely around stars cooler than K-type unless sustained by extreme tidal and/or other interior heating. This framework links atmosphere survival from USP lava worlds to habitable zone planets, informing target selection and interpretation for TRAPPIST-1 and JWST DDT characterization.

astro-ph.EP

SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching

Denoising diffusion transformers achieve strong generation quality but converge slowly during training. Regularizing their internal representations has emerged as an effective accelerator, yet existing methods split into two families with complementary costs. Target-based methods strengthen representations by aligning them to external features, which requires an external encoder and a learnable projection head to bridge feature spaces. Target-free methods hold no reference at all, and can only repel the model's own features across samples or layers, discarding whatever structure the data contains. Prior work suggests that spatial structure, rather than global semantics, drives the gains of alignment. We therefore ask whether such structure can serve as a target directly, and whether it exists not only within an image but across images. Our key insight is that the clean data latent already carries this structure in the relations among its tokens, where a relation is the similarity between two tokens, a single scalar comparable across feature spaces without a projection head. We propose Structural Parameter-free Affinity Regularization (SPARE), a regularizer that matches the pairwise affinities of intermediate tokens to those of the clean latents. To exploit this structure fully, SPARE extends the matching to token pairs across images, precisely the pairs that prior target-free methods repel by default, and calibrates both relation types with a single learning objective. On ImageNet $256 \times 256$ with SiT backbones under matched 400K-iteration budgets, SPARE adds no encoder, head, or parameters and only 0.08 GB of training memory, yet attains the lowest FID among parameter-free regularizers in every tested setting, recovers 37 to 54\% of REPA's FID reduction, and improves over REPA when combined with it, reaching FID 1.90 under classifier-free guidance at 1M iterations.

cs.CV

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often miss within-stage heterogeneity and the clinical context in electronic medical records (EMRs). We present HCC-STAR (Hepatocellular Carcinoma Staging, Treatment And pRognosis), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score-based staging, ranked guideline-consistent treatments with evidence-based rationales, and individualized survival estimates. We curated about 30,000 HCC cases from SEER and expanded them into EMR-style narrative training data using a clinician-validated, prompt-based augmentation workflow. On this corpus, we developed a knowledge-aligned reasoning framework optimized with a step-verifiable composite reward, moving beyond text-level memorization of clinical guidelines. In a multi-center cohort of 6,668 patients from 12 hospitals in China, HCC-STAR achieved state-of-the-art performance in treatment recommendation and risk stratification compared with clinical guidelines and competitive models, including GPT-5 and Gemini-2.5 Pro. Hypothetical overall-survival analysis showed a median survival of 51 months under adherence to HCC-STAR recommendations, compared with 29 and 32 months under BCLC and CNLC. In clinician-centric evaluations, blinded hepatobiliary specialists rated HCC-STAR's reasoning and evidence-based justifications as trustworthy. The model surpassed resident and attending physicians in treatment accuracy and helped physicians make more accurate decisions faster when used as an assistant. These findings support HCC-STAR as a reliable and verifiable decision-support system for risk stratification and precision therapy in HCC.

cs.AI

Generative Site-Specific Beamforming for UPAs via Decoupled Channel Sensing

A cross-fused generative site-specific beamforming (GenSSBF) framework is proposed for low-overhead beam alignment in uniform planar array (UPA) systems. A decoupled channel sensing strategy is developed, where the azimuth and elevation domains of the UPA are probed independently, and the online sweeping overhead is reduced from multiplicative to linear complexity compared to exhaustive two-dimensional codebook sweeping. However, the resulting reference signal received power (RSRP) observations only contain marginal angular power information. The explicit azimuth-elevation coupling of the UPA channel is therefore lost. Beam generation from these separate observations becomes highly ambiguous. To address this issue, a bidirectional cross-attention encoder is designed to extract and fuse the latent dependency between the azimuth and elevation sensing branches. Conditioned on the fused feature, a conditional normalizing flow generator is proposed to generate a compact set of high-fidelity beam candidates. These candidates are further verified through lightweight pilot measurements for final beam selection. A task-oriented training objective is also introduced to encourage the generated candidate set to contain at least one high-gain beam, rather than fitting the full conditional beam distribution. Simulation results based on DeepMIMO scenarios show that the proposed framework consistently outperforms deterministic beam prediction and conventional discrete Fourier transform (DFT) codebook search. Compared with the full 1024-beam two-dimensional DFT search, normalized beamforming gain improvements of 83.6%, 74.6%, and 38.1% are achieved in the I2_28, O1B_28, and Boston5G_28 scenarios, respectively, while the sweeping overhead is reduced by 93.8%.

eess.SP

OneBar: An End-to-End Content-Grounded Generative Query Recommendation Framework for E-Commerce Video Feeds

Short-video platforms now expose clickable search entries beneath the video player, enabling users to easily express content-induced search intent. However, conventional query recommendation systems on short-video platforms suffer from latency constraints and objective misalignment, while recent generative approaches struggle with noisy content-side metadata and preference drift. To address these issues, we propose OneBar, an end-to-end generative framework for real-time query recommendation for E-Commerce video feeds. OneBar features three key innovations: (1) a collaborative-multimodal intent grounding module that fuses multimodal video understanding and behavior-derived collaborative anchors; (2) a Unified End-to-End architecture equipped with a prompt-compression mechanism for efficient online serving; and (3) a progressive preference learning strategy for efficient preference-internalization, which internalizes hierarchical behavior preferences into the generative policy, eliminating the need for a separately trained reward model. Compared with online base, OneBar increases Query Exposure by 16.91\% and Query Click by 18.68\%, while maintaining a slight Query CTR gain of 0.19\%. The additional search traffic further contributes to 20.36\% more guided orders and 21.67\% higher GMV.

cs.IR

Hydrogen and Helium Dissolution, Outgassing, and Loss in Evolving Sub-Neptune Magma Oceans: Examining Demographic Features and Radius Evolution

Sub-Neptunes' molten interiors are expected to accommodate large quantities of volatiles, potentially altering their radius evolution. Previous studies have examined this effect in isolation with simplified evolution modeling, often assuming idealized interior and atmospheric conditions. To address this limitation, we introduce SEAMIST, a unified evolution model for sub-Neptunes and super-Earths that self-consistently combines, for the first time, interior structure, cooling, rock/iron solidification, boil-off, photoevaporation, H/He dissolution, and atmospheric composition. SEAMIST considers both a partially soluble case, in which hydrogen partitions into the magma ocean with a fraction set by mantle-envelope boundary conditions, and a fully miscible case, in which hydrogen may fully dissolve into the magma ocean at high temperatures. We identify a novel catastrophic boil-off mechanism, triggered by a positive feedback between hydrogen outgassing and mass loss that can operate billions of years after disk dispersal. In partially soluble models, the impact of H/He dissolution on radius evolution is modest. This contrasts with previous expectations, as we find that increasing hydrogen abundance from outgassing enhances mass-loss efficiency, counterbalancing volatile replenishment from the rock/iron interior. Fully miscible hydrogen, in contrast, significantly enhances envelope survival especially around low-mass stars. Overall, at intermediate to low masses, mass-radius curves from partially soluble models match observed distributions. The fully miscible case predicts a pronounced radius peak and excess planets around low stellar masses that appear inconsistent with current observations, although it reproduces the observed radius ``cliff" near 4$R_\oplus$ at higher masses. Our results suggest that high metallicity may explain the cliff, although alternatives cannot be entirely ruled out.

astro-ph.EP

Latent Reasoning with Normalizing Flows

Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only partially formed. Latent reasoning offers a higher-bandwidth alternative by performing intermediate computation in compact continuous states before committing to text. Yet existing latent-reasoning methods often sacrifice key advantages that make CoT effective in autoregressive language models, including native left-to-right generation, probabilistic sampling, compatibility with KV-cache decoding, and tractable likelihood estimation. We propose NF-CoT, a latent reasoning framework that preserves these advantages by modeling continuous thoughts with normalizing flows. NF-CoT instantiates a TARFlow-style normalizing flow inside the LLM backbone, defining a tractable probability model over compact continuous thoughts distilled from explicit CoT. Continuous-thought positions are generated by an NF head, while text positions are generated by the standard LM head within the same causal stream. This design provides exact likelihoods for latent thoughts, enables probabilistic left-to-right decoding with the original KV cache, and supports direct policy-gradient optimization in the latent reasoning space. On code-generation benchmarks, NF-CoT improves pass rates over explicit-CoT and prior latent-reasoning baselines while substantially reducing intermediate-reasoning cost.

cs.CL

The Velocity Deficit: Initial Energy Injection for Flow Matching

While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity Deficit. We show that the MSE objective systematically underestimates velocity magnitude, causing generated samples to fail to reach the data manifold-a phenomenon we term Integration Lag. To rectify this, we propose Initial Energy Injection, instantiated via two complementary methods: the training-based Magnitude-Aware Flow Matching (MAFM) and the training-free Scale Schedule Corrector (SSC). Both are grounded in our discovery of a crucial asymmetry: velocity contraction causes harmful kinetic stagnation at the trajectory's start, yet acts as a beneficial denoising mechanism at its end. Empirically, SSC yields significant efficiency gains with zero retraining and just one line of code. On ImageNet-1k (256x256), it improves FID by 44.6% (from 13.68 to 7.58) and achieves a 5x speedup, enabling a 50-step generator (FID 7.58) to beat a 250-step baseline (FID 8.65). Furthermore, our methods generalize to Text-to-Image tasks and high-resolution generation, improving FID on MS-COCO by ~22%.

cs.CV

Horizontal transport as a source of disequilibrium chemistry on the nightside of a hot exoplanet

Hot Jupiters have temperature gradients of several hundreds of degrees between their permanent day and nightsides. In equilibrium, the primary carbon reservoir is expected to transition from CO on the dayside to CH4 on the nightside. Theory predicts that the atmospheric circulation, characterised by km/s winds, can advect chemical species from the dayside to the nightside faster than the time needed for the CO-to-CH4 chemical reaction to reach equilibrium. However direct evidence of this process has, so far, remained elusive, partly because it is often degenerate with other processes, such as vertical mixing or non-stellar elemental abundances. Here, we present observational evidence for such day-to-night transport of chemical species by observing both the dayside and the nightside of the hot Jupiter NGTS-10A b with the JWST/NIRSpec instrument. We constrain the presence of H2O and CO with similar abundances on both the dayside and nightside. Our observations are compatible with a solar-composition atmosphere at chemical equilibrium on the dayside, but indicative of disequilibrium chemistry for the nightside as it is significantly depleted in CH4 compared to equilibrium chemistry predictions. We further show that the lack of CH4 on the planet's nightside cannot be attributed to non-solar elemental abundances or to vertical mixing mechanisms and must therefore be due to horizontal chemical quenching. Our study shows the fundamental role atmospheric transport plays in shaping the distribution of chemical species on exoplanet atmospheres.

astro-ph.EP

Diagnosing and Repairing Citation Failures in Generative Engine Optimization

Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution-how much a document influences a response-rather than citation, the mechanism that actually drives traffic back to creators. Also, these methods apply generic rewriting rules uniformly, failing to diagnose why individual document are not cited. This paper introduces a diagnostic approach to GEO that asks why a document fails to be cited and intervenes accordingly. We develop a unified framework comprising: (1) the first taxonomy of citation failure modes spanning different stages of a citation pipeline; (2) AgentGEO, an agentic system that diagnoses failures using this taxonomy, selects targeted repairs from a corresponding tool library, and iterates until citation is achieved; and (3) a document-centric benchmark evaluating whether optimizations generalize across held-out queries. AgentGEO achieves over 40% relative improvement in citation rates while modifying only 5% of content, compared to 25% for baselines. Our analysis reveals that generic optimization can harm long-tail content and some documents face challenges that optimization alone cannot fully address-findings with implications for equitable visibility in AI-mediated information access.

cs.IR

Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge

Large language models often solve complex reasoning tasks more effectively with Chain-of-Thought (CoT), but at the cost of long, low-bandwidth token sequences. Humans, by contrast, often reason softly by maintaining a distribution over plausible next steps. Motivated by this, we propose Multiplex Thinking, a stochastic soft reasoning mechanism that, at each thinking step, samples K candidate tokens and aggregates their embeddings into a single continuous multiplex token. This preserves the vocabulary embedding prior and the sampling dynamics of standard discrete generation, while inducing a tractable probability distribution over multiplex rollouts. Consequently, multiplex trajectories can be directly optimized with on-policy reinforcement learning (RL). Importantly, Multiplex Thinking is self-adaptive: when the model is confident, the multiplex token is nearly discrete and behaves like standard CoT; when it is uncertain, it compactly represents multiple plausible next steps without increasing sequence length. Across challenging math reasoning benchmarks, Multiplex Thinking consistently outperforms strong discrete CoT and RL baselines from Pass@1 through Pass@1024, while producing shorter sequences. The code and checkpoints are available at https://github.com/GMLR-Penn/Multiplex-Thinking.

cs.CL

Sub-Neptune Memories I: Implications of Inefficient Mantle Cooling and Silicate Rain

We explore the evolution of sub-Neptune (radii between $\sim$1.5 and 4 R$_\oplus$) exoplanet interior structures using our upgraded evolution code, \texttt{APPLE}, which self-consistently couples the thermal and compositional evolution of the whole structure. We incorporate stably stratified regions with convective mixing and, for the first time, ab initio results on the phase separation of silicate-hydrogen mixtures to model silicate rain in sub-Neptune envelopes. We demonstrate that inefficient mantle cooling can retain sufficient heat to Gyr ages: inefficient heat transport from mantle to envelope alone keeps radii $\sim$10\% larger than predicted by adiabatic models at late times. Silicate rain can contribute an additional $\sim$5\% to the radius, depending on envelope mass and initial metal abundance. The silicate-hydrogen immiscibility region may lie in the middle or even upper envelope, far above the envelope-mantle boundary layer, and bifurcates the envelope into two an upper, hydrogen-rich region and a lower, metal-rich region above the mantle. If silicate rain occurs, atmospheres should appear depleted of silicates while radii remain inflated at late ages. To demonstrate the effects of inefficient mantle cooling, we present interior evolution models for GJ 1214 b, K2-18 b, TOI-270 d, and TOI-1801 b, showing that hot, liquid silicate mantles with thin envelopes reproduce their radii and mean densities, providing an alternative to water-world interpretations. These results imply that bulk compositions inferred from mean density must account for the mantle thermal state and the envelope mixing/phase-separation history; such thermal ``memories'' may constrain formation entropies and temperatures when metallicities are more precisely measured.

astro-ph.EP

VeCoR -- Velocity Contrastive Regularization for Flow Matching

Flow Matching (FM) has recently emerged as a principled and efficient alternative to diffusion models. Standard FM encourages the learned velocity field to follow a target direction; however, it may accumulate errors along the trajectory and drive samples off the data manifold, leading to perceptual degradation, especially in lightweight or low-step configurations. To enhance stability and generalization, we extend FM into a balanced attract-repel scheme that provides explicit guidance on both "where to go" and "where not to go." To be formal, we propose \textbf{Velocity Contrastive Regularization (VeCoR)}, a complementary training scheme for flow-based generative modeling that augments the standard FM objective with contrastive, two-sided supervision. VeCoR not only aligns the predicted velocity with a stable reference direction (positive supervision) but also pushes it away from inconsistent, off-manifold directions (negative supervision). This contrastive formulation transforms FM from a purely attractive, one-sided objective into a two-sided training signal, regularizing trajectory evolution and improving perceptual fidelity across datasets and backbones. On ImageNet-1K 256$\times$256, VeCoR yields 22\% and 35\% relative FID reductions on SiT-XL/2 and REPA-SiT-XL/2 backbones, respectively, and achieves further FID gains (32\% relative) on MS-COCO text-to-image generation, demonstrating consistent improvements in stability, convergence, and image quality, particularly in low-step and lightweight settings. Project page: https://p458732.github.io/VeCoR_Project_Page/

cs.CV

Understanding the Origins of Super-Puff Planets: A New Mass-Loss Regime Coupled to Planetary Evolution

Super-puffs are a class of low-mass, large-radius planets that have challenged planet formation and evolution models. Their high inferred H/He mass fractions, required to explain their physical sizes, would lead to rapid atmospheric escape, raising questions about their long-term retention. Recent modeling work indicates that low-mass planets typically require 50\% less H/He mass to match their observed radius, due to significant roles of the radiative atmosphere and interior heating from the rock/iron core. Here, through a new quantitative analysis of XUV-driven escape in sub-Neptunes, we find that previous studies overestimated mass loss, as scaling laws in low-gravity regimes deviate greatly from the widely used energy-limited regime. We define a new regime, thermal-energy-mediated photoevaporation (TEMP), in which thermal energy conversion critically sets the mass-loss rate. These effects make super-puffs more resilient to mass loss than previously thought. We develop a coupled evolution model integrating this updated thermal evolution framework with a 1D hydrodynamic photoevaporation model. Applying this novel, joint model to observed super-puffs and young low-density planets, we find that their masses, radii and transit pressures align with predictions assuming either a clear or hazy atmosphere. This indicates that super-puffs have undergone a combination of boil-off and photoevaporative mass loss, with boil-off dominating the process. Our results indicate that low-density planets typically possess both a thick convective envelope and substantial radiative atmosphere, which contribute to their large radii. For this to occur, these planets must have intermediate masses of 5-10$M_\oplus$ and receive stellar insolation $\lesssim 30F_\oplus$, favoring FG-type stars over M-dwarfs.

astro-ph.EP

NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management

Note-taking is a critical practice for capturing, organizing, and reflecting on information in both academic and professional settings. The recent success of large language models has accelerated the development of AI-assisted tools, yet existing solutions often struggle with efficiency. We present NoteBar, an AI-assisted note-taking tool that leverages persona information and efficient language models to automatically organize notes into multiple categories and better support user workflows. To support research and evaluation in this space, we further introduce a novel persona-conditioned dataset of 3,173 notes and 8,494 annotated concepts across 16 MBTI personas, offering both diversity and semantic richness for downstream tasks. Finally, we demonstrate that NoteBar can be deployed in a practical and cost-effective manner, enabling interactive use without reliance on heavy infrastructure. Together, NoteBar and its accompanying dataset provide a scalable and extensible foundation for advancing AI-assisted personal knowledge management.

cs.CL

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

Pre-trained vision-language models have exhibited remarkable abilities in detecting out-of-distribution (OOD) samples. However, some challenging OOD samples, which lie close to in-distribution (InD) data in image feature space, can still lead to misclassification. The emergence of foundation models like diffusion models and multimodal large language models (MLLMs) offers a potential solution to this issue. In this work, we propose SynOOD, a novel approach that harnesses foundation models to generate synthetic, challenging OOD data for fine-tuning CLIP models, thereby enhancing boundary-level discrimination between InD and OOD samples. Our method uses an iterative in-painting process guided by contextual prompts from MLLMs to produce nuanced, boundary-aligned OOD samples. These samples are refined through noise adjustments based on gradients from OOD scores like the energy score, effectively sampling from the InD/OOD boundary. With these carefully synthesized images, we fine-tune the CLIP image encoder and negative label features derived from the text encoder to strengthen connections between near-boundary OOD samples and a set of negative labels. Finally, SynOOD achieves state-of-the-art performance on the large-scale ImageNet benchmark, with minimal increases in parameters and runtime. Our approach significantly surpasses existing methods, and the code is available at https://github.com/Jarvisgivemeasuit/SynOOD.

cs.CV