SearcharxivSearch

arXiv subjects

Yipeng Wang

Publications and source records attributed to Yipeng Wang.

At least 19 recordsLinked to original sources

KuaiRP Series Role-playing Models Technical Report

This paper introduces the complete technical solution for the KuaiRP series of role-playing models. We aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, and high-efficiency deployment with a small parameter size. However, effectively injecting deep domain knowledge often leads to a severe catastrophic forgetting of the model's general agent capabilities. To overcome this trade-off, we propose a multi-stage training pipeline. First, we design a standardized character template and construct an SFT data pipeline based on user behavior simulation and reverse profile filtering. Next, we utilize a rule-based composite reward function during the Reinforcement Learning (RL) phase to eliminate common degradation phenomena like length expansion and repetitive generation. Finally, to recover the general capabilities compromised during SFT and RL, we propose a novel self-distillation paradigm using Two-stage On-Policy Distillation (OPD) equipped with Cumulative-Divergence Decay (CDD). By using the domain-adapted model as the teacher and the original base model as the student, we effectively balance deep domain knowledge injection with the preservation of general agent capabilities. Experimental results demonstrate that the KuaiRP models not only match the current state-of-the-art proprietary models in role-playing fidelity within our target domains, but also successfully recover general agent capabilities, maintaining extremely low deployment costs.

cs.AI

Rigidity and flexibility under positive isotropic curvature

For every $n\ge4$ and $L>0$, we construct a smooth $4$-PIC metric on $S^n$ with Urysohn $1$-width at least $L$ and an embedded stable minimal disk of intrinsic inradius at least $L$. These examples disprove the proposed width and stable-disk radius bounds under a positive lower bound for isotropic curvature. On closed even-dimensional manifolds, we prove the sharp estimate $\lambda_1^{(2)}\ge(n-1)\sigma/2$ under $\sigma$-PIC and show that equality forces roundness if a closed eigenform attaining the bound has rank at least four at some point.

math.DG

Minimal Two-Spheres and Manifolds with Positive Isotropic Curvature

We improve the Micallef--Moore index estimate for harmonic two-spheres in $n$-manifolds with positive isotropic curvature to the sharp bound $n-3$. Combining with recent work, this completes the diffeomorphism classification of closed manifolds with positive isotropic curvature in the remaining dimensions five and six.

math.DG

Whose Posts Get Ranked: Identical-Text Exposure Gaps in Bluesky Custom Feeds

Bluesky lets users deploy custom feeds, independently operated recommendation algorithms that the platform serves alongside thousands of others. This paper investigates how evenly these feeds treat posts with the same text. To measure this, we take repeated snapshots of the Top-50 lists that 1,366 public feeds return, and we group posts with identical text, from different authors, that were created before the same list response and closely matched in age. Exposure diverges widely inside these matched sets, which span 250 feeds: in 33% of sets, one copy appears on the list while another does not. Fixed-effects regressions show that this divergence is associated with the author's history on the specific feed. Authors new to a feed receive less exposure for the same text (-0.061 in reciprocal-rank weight), while authors whose posts the feed has returned before receive more. A new author with more followers than the competing author still loses 74\% of head-to-head comparisons. Media and post-type features show no detectable association after multiple-comparison correction. These results are early evidence that access to many independent feeds is not enough to give identical texts equal exposure.

cs.SI

Predicting Custom-Feed Returns for New Bluesky Posts: A Prospective Study

The conventional approach to cold-start recommendation addresses new users or newly introduced items. Bluesky custom feeds create a different setting: independently operated feeds filter content from a shared public stream. In this setting, newly published posts are the cold-start objects, while the feeds serve as candidates. We propose a cold-start routing task in which a newly ingested public post is the query and all rankable feeds in the monitored panel are ranked according to whether each will subsequently return it. We build a still-evolving collect-first, label-later benchmark dataset. The collected dataset covers a fixed panel of 5,000 monitored feeds and contains 17.804 million public posts, 1.865 million observable post--feed return records, and 625,083 valid feed polls. The labels record whether a post is observed among a feed's AppView Top-50 results in at least one poll during the 24 hours after publication. The current experiments use two disjoint 24-hour test folds, each paired with a 24-hour training window and separated by a 24-hour outcome-availability gap. Evaluation is conditional on the 602,186 test posts that have at least one positive observed label and satisfy the metric eligibility criteria; these posts account for 9.04% of all 6,661,658 test posts. Across the two folds, LambdaRank achieves the best equal-fold mean values among the evaluated models: 0.7361 for capped Recall@10, 0.6127 for NDCG@10, and 0.7749 for Hit@10.

cs.IR

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.

cs.CL

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

We present Flex4DHuman, a multi-view video diffusion model that transforms a monocular or sparse multi-view video of a dynamic subject into synchronized dense multi-view videos using only relative camera-pose conditioning. Unlike prior human-centric methods that rely on skeletons, depth maps, normals, or rendered target-view geometry, Flex4DHuman requires no explicit geometry priors and instead conditions generation through relative camera-pose positional encoding. The generated videos can be directly ingested by downstream reconstruction pipelines to create dynamic 4D Gaussian splats. Built on the Wan 2.1 1.3B text-to-video model, Flex4DHuman preserves the backbone architecture and encodes camera and view information through a five-axis positional encoding that extends spatio-temporal RoPE with view indices and continuous SE(3) relative camera geometry. A three-stage curriculum progressively trains the model for pose following, flexible reference-to-target view generation, and temporal rollout. To support temporal rollout, we train with clean historical target-view tokens. We also add multi-view captions to enable test-time text control. Combined with an off-the-shelf 4D Gaussian Splatting stage, our framework lifts monocular static-camera videos into dynamic 4D Gaussian splats. Experiments on DNA-Rendering and ActorsHQ show that Flex4DHuman surpasses prior state-of-the-art methods, while the same formulation generalizes to animal categories after mixed human-animal training. These capabilities make Flex4DHuman a practical step toward scalable 4D content creation from casual monocular videos for simulation, gaming, AR/VR, and video re-shooting.

cs.CV

EgoPush: Learning End-to-End Egocentric Multi-Object Rearrangement for Mobile Robots

Humans can rearrange objects in cluttered environments using egocentric perception, navigating occlusions without global coordinates. Inspired by this capability, we study long-horizon multi-object non-prehensile rearrangement for mobile robots using a single egocentric camera. We introduce EgoPush, a policy learning framework that enables egocentric, perception-driven rearrangement without relying on explicit global state estimation that often fails in dynamic scenes. EgoPush designs an object-centric latent space to encode relative spatial relations among objects, rather than absolute poses. This design enables a privileged reinforcement-learning (RL) teacher to jointly learn latent states and mobile actions from sparse keypoints, which is then distilled into a purely visual student policy. To reduce the supervision gap between the omniscient teacher and the partially observed student, we restrict the teacher's observations to visually accessible cues. This induces active perception behaviors that are recoverable from the student's viewpoint. To address long-horizon credit assignment, we decompose rearrangement into stage-level subproblems using temporally decayed, stage-local completion rewards. Extensive simulation experiments demonstrate that EgoPush significantly outperforms end-to-end RL baselines in success rate, with ablation studies validating each design choice. We further demonstrate zero-shot sim-to-real transfer on a mobile platform in the real world. Code and videos are available at https://ai4ce.github.io/EgoPush/.

cs.RO

Area-minimizing capillary cones

We construct non-flat minimal capillary cones with bi-orthogonal symmetry groups for any dimension and contact angle. These cones interpolate between rescalings of a singular solution to the one-phase problem and the free-boundary cone obtained by halving a Lawson cone along a hyperplane of symmetry. The existence and uniqueness of such cones is proved by solving a nonlinear free boundary equation parametrized by the contact angle and obtaining monotonicity properties for the solutions. The constructed cones are minimizing in ambient dimension $8$ or higher, for appropriate contact angles, demonstrating that the regularity theory for minimizing capillary hypersurfaces can have singularities in codimension $7$ and completing the capillary regularity theory for contact angles near $\pi/2$. We further develop the connection between capillary hypersurfaces and solutions of the one-phase problem, consequently producing new examples of singular minimizing free boundaries for the Alt-Caffarelli functional.

math.DG

Stability inequalities for one-phase cones

We obtain strict stability inequalities for homogeneous solutions of the one-phase Bernoulli problem. We prove that in dimension $7$ and above, cohomogeneity one solutions with bi-orthogonal symmetry are strictly stable. As a consequence, we obtain a bound on the first eigenvalue and the decay rates of Jacobi fields, with applications to the generic regularity of the one-phase problem.

math.AP

The Tag is the Signal: URL-Agnostic Credibility Scoring for Messages on Telegram

Telegram has become one of the leading platforms for disseminating misinformational messages. However, many existing pipelines still classify each message's credibility based on the reputation of its associated domain names or its lexical features. Such methods work well on traditional long-form news articles published by well-known sources, but high-risk posts on Telegram are short and URL-sparse, leading to failures for link-based and standard TF-IDF models. To this end, we propose the TAG2CRED pipeline, a method designed for such short, convoluted messages. Our model will directly score each post based on the tags assigned to the text. We designed a concise label system that covers the dimensions of theme, claim type, call to action, and evidence. The fine-tuned large language model (LLM) assigns tags to messages and then maps these tags to calibrated risk scores in the [0,1] interval through L2-regularized logistic regression. We evaluated 87,936 Telegram messages associated with Media Bias/Fact Check (MBFC), using URL masking and domain disjoint splits. The results showed that the ROC-AUC of the TAG2CRED model reached 0.871, the macro-F1 value was 0.787, and the Brier score was 0.167, outperforming the baseline TF-IDF (macro-F1 value 0.737, Brier score 0.248); at the same time, the number of features used in this model is much smaller, and the generalization ability on infrequent domains is stronger. The performance of the stacked ensemble model (TF-IDF + TAG2CRED + SBERT) was further improved over the baseline SBERT. ROC-AUC reached 0.901, and the macro-F1 value was 0.813 (Brier score 0.114). This indicates that style labels and lexical features may capture different but complementary dimensions of information risk.

cs.SI

Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training

3D Gaussian Splatting (3DGS) has recently emerged as a foundational technique for real-time neural rendering, 3D scene generation, volumetric video (4D) capture. However, its rendering and training impose massive computation, making real-time rendering on edge devices and real-time 4D reconstruction on workstations currently infeasible. Given its fixed-function nature and similarity with traditional rasterization, 3DGS presents a strong case for dedicated hardware in the graphics pipeline of next-generation GPUs. This work, Vorion, presents the first GPGPU prototype with hardware-accelerated 3DGS rendering and training. Vorion features scalable architecture, minimal hardware change to traditional rasterizers, z-tiling to increase parallelism, and Gaussian/pixel-centric hybrid dataflow. We prototype the minimal system (8 SIMT cores, 2 Gaussian rasterizer) using TSMC 16nm FinFET technology, which achieves 19 FPS for rendering. The scaled design with 16 rasterizers achieves 38.6 iterations/s for training.

cs.AR

DynamicRTL: RTL Representation Learning for Dynamic Circuit Behavior

There is a growing body of work on using Graph Neural Networks (GNNs) to learn representations of circuits, focusing primarily on their static characteristics. However, these models fail to capture circuit runtime behavior, which is crucial for tasks like circuit verification and optimization. To address this limitation, we introduce DR-GNN (DynamicRTL-GNN), a novel approach that learns RTL circuit representations by incorporating both static structures and multi-cycle execution behaviors. DR-GNN leverages an operator-level Control Data Flow Graph (CDFG) to represent Register Transfer Level (RTL) circuits, enabling the model to capture dynamic dependencies and runtime execution. To train and evaluate DR-GNN, we build the first comprehensive dynamic circuit dataset, comprising over 6,300 Verilog designs and 63,000 simulation traces. Our results demonstrate that DR-GNN outperforms existing models in branch hit prediction and toggle rate prediction. Furthermore, its learned representations transfer effectively to related dynamic circuit tasks, achieving strong performance in power estimation and assertion prediction.

cs.LG

On fill-ins with scalar curvature bounded from below and an inequality of Hijazi-Montiel-Rold\'an

We consider fill-ins of spin manifolds with scalar curvature bounded by $-n(n-1)$. Gromov proposed a conjecture relating the infimum of the mean curvature of such a fill-in to the hyperspherical radius. We observe that the inequality conjectured by Gromov follows by combining an inequality of Hijazi-Montiel-Rold\'an for the first Dirac eigenvalue with a recent theorem of B\"ar. Moreover, we give an alternative proof of the Hijazi-Montiel-Rold\'an inequality based on the work of B\"ar and B\"ar-Ballmann.

math.DG

Observation of Iron Oxide to Nitride Conversion via Liquid Liquid Phase Separation in High pressure Borate Melt

High pressure chemistry provides a powerful route to materials that are inaccessible or difficult to synthesize under ambient conditions. However, high pressure chemical reaction processes and mechanisms remain largely unexplored because of the challenges associated with in situ characterization under high pressure and high temperature, particularly within the deeply enclosed sample environment of a large volume press. Here, we employ the state of the art real time synchrotron X ray radiography to image a high pressure chemical reaction at 5.4 GPa and 1700 K within a large volume press. Using Fe2O3 and BN as precursors, we capture the complete dynamic metal oxide to nitride conversion and show that it differs fundamentally from conventional solid state diffusion controlled nitridation. Synchrotron X ray radiography clearly revealed that this conversion involves a two stage liquid liquid separation process (LLPS), including fluid nitrogen and fluid Fe N alloy. On the basis of these observations, we propose a nitrogen driven mechanism for LLPS in borate melts. Specifically, nitrogen reduces metal cations in the borate network, altering their coordination environments and triggering a substantial reorganization of the melt structure. This coordination induced restructuring destabilizes the borate melt and promotes the LLPS of fluid Fe N alloy. Our in-situ observations suggest a general pathway for high pressure metal oxide to nitride conversion. This study provides a direct visualization of a pres-sure-enabled chemical reaction that is inaccessible under ambient conditions, offering fundamental in-sight into how high pressure reshapes chemical reaction pathways and enables the synthesis of metal nitrides.

physics.chem-ph

Uniqueness of Cylindrical Tangent Cones $C_{p,q} \times \mathbb{R}$

We show the uniqueness of the cylindrical tangent cone $C(\mathbb{S}^2 \times \mathbb{S}^4) \times \mathbb{R}$ for area-minimizing hypersurfaces in $\mathbb{R}^9$, completing the uniqueness of all tangent cones of the form $C_{p,q} \times \mathbb{R}$ proved by Simon for dimensions at least 10 and Sz\'ekelyhidi for the Simons cone.

math.DG

Time-Lapse Video-Based Embryo Grading via Complementary Spatial-Temporal Pattern Mining

Artificial intelligence has recently shown promise in automated embryo selection for In-Vitro Fertilization (IVF). However, current approaches either address partial embryo evaluation lacking holistic quality assessment or target clinical outcomes inevitably confounded by extra-embryonic factors, both limiting clinical utility. To bridge this gap, we propose a new task called Video-Based Embryo Grading - the first paradigm that directly utilizes full-length time-lapse monitoring (TLM) videos to predict embryologists' overall quality assessments. To support this task, we curate a real-world clinical dataset comprising over 2,500 TLM videos, each annotated with a grading label indicating the overall quality of embryos. Grounded in clinical decision-making principles, we propose a Complementary Spatial-Temporal Pattern Mining (CoSTeM) framework that conceptually replicates embryologists' evaluation process. The CoSTeM comprises two branches: (1) a morphological branch using a Mixture of Cross-Attentive Experts layer and a Temporal Selection Block to select discriminative local structural features, and (2) a morphokinetic branch employing a Temporal Transformer to model global developmental trajectories, synergistically integrating static and dynamic determinants for grading embryos. Extensive experimental results demonstrate the superiority of our design. This work provides a valuable methodological framework for AI-assisted embryo selection. The dataset and source code will be publicly available upon acceptance.

cs.CV