SearcharxivSearch

arXiv subjects

Cheng Chen

Publications and source records attributed to Cheng Chen.

At least 19 recordsLinked to original sources

CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization

Generalized Category Discovery (GCD) assigns unlabeled instances, mixed with labeled data, to known or novel categories, requiring human-like compositional reasoning: reusing primitives learned from known classes and deciding when new combinations imply new categories. Existing GCD methods operate on unstructured token features and struggle to extrapolate to novel compositions. We propose CoGe-GCD, which rethinks GCD through compositional generalization with two coupled stages. (i) Compositional Perception structures patch tokens by mapping them to a small vocabulary of primitives and refining token embeddings via competitive token-primitive assignment and information passing, yielding coherent groups for discovery. (ii) Generalizing Induction exploits the induced geometric structure and applies a structure-preserving calibration over spatial relations, maintaining probabilistic semantics while improving extrapolation to unseen primitive combinations. CoGe-GCD is implemented as an inductive-bias module between backbone and projection head, without modifying heads or losses, and can be plugged into diverse GCD frameworks. On standard benchmarks, it consistently improves all-class accuracy, unknown-class number estimation, and geometric quality, with marginal computational overhead. Code is available at https://github.com/lytang63/CoGe-GCD.

cs.LG

Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance

Reliable video generation requires more than high-quality frames to form a coherent story: a model must maintain a persistent state, transporting visual attributes such as identity, scene layout, motion, and fine details across time. Existing training-free methods mainly strengthen cross-frame attention or analyze local attention entropy, but these views do not reveal whether temporal interactions stay in a healthy transport regime. In this work, we study video generation through the perspective of Temporal State Transport. We introduce Spectral Tension, a signed diagnostic that compares local attention diffuseness with global spectral diversity, and use it to identify two opposite temporal failures: fragmented transport and over-mixing hotspots. Based on this diagnosis, we propose Spectral Transport Homeostasis, a training-free regulator that softly corrects pathological temporal states while largely preserving balanced ones. Experiments on pretrained video generation models show that the original model often occupies imbalanced temporal regimes, whereas our method selectively applies larger corrections to the worst temporal hotspots and improves temporal consistency and visual quality without finetuning. Code: https://github.com/lytang63/temporal-state-transport

cs.CV

Some constructions of restricted Kakeya sets

In this paper, we consider Kakeya sets with the additional restriction that centers of the unit line segments belong to a given set. In particular, for every uncountable Borel set $A\subset\mathbb{R}^d$, $d\geq 2$, we construct a compact subset of $\mathbb{R}^d$ of Lebesgue measure zero that contains, in every direction, a unit line segment whose center lies in $A$. Notice that every such Kakeya set (not even necessarily compact) must have positive Lebesgue measure if $A$ is countable. So our result shows that the countability is in fact the only obstruction.

math.CA

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irregular, and sparsely supervised, and in computation asymmetry, where each request scores many candidates against one shared user history under tight latency budgets. We propose ReST, a recommendation-native Transformer scaling framework. For signal quality, it introduces a sequence encoder with dual-gated attention, rotary positional and temporal embedding, stabilized residual normalization, and training-only auxiliary objectives. For computation asymmetry, it factorizes ranking into a heavy reusable encoder and a lightweight cross decoder with projection-free KV attention and token-specific parameterization, coupling user-level shared-prefix training with shared-prefix serving for compute-once, decode-many-times ranking. Across industrial and public benchmarks, ReST achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate. A one-week online A/B test on a production advertising platform improves online AUC by 1.31% and lifts a core revenue metric by 11.93% within a 50 ms P99 budget; ReST has since been fully deployed in production, showing that behavior-sequence scaling remains a promising, under-exploited axis for production ranking.

cs.IR

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.

cs.RO

On the formation of retrograde S-type planets in binaries with a polar circumbinary disk

Retrograde S-type planets have been observed in several binary systems, yet their formation pathway remains poorly understood. With high-resolution hydrodynamic simulations, we demonstrate that a polar circumbinary disk around an eccentric, unequal-mass binary can form and sustain a retrograde mini disk around the primary star. This provides a direct in-situ formation channel for retrograde S-type planets. The mini disk forms via a sub-Keplerian accretion stream that is slightly misaligned from the polar disk. The mini disk initially undergoes von Zeipel-Kozai-Lidov (ZKL) oscillations, driving coupled eccentricity and inclination evolution. Rather than oscillating indefinitely, the inner mini disk evolves past the critical ZKL inclination, decouples from the outer disk, and settles into a stable retrograde orbit. This evolution is sensitive to numerical resolution: the retrograde configuration is absent in previous lower-resolution simulations, where the mini disk accretion timescale is too short to sustain ZKL-driven evolution. For a higher disk viscosity, the mini disk remains near-polar due to a shorter accretion timescale. Since protoplanetary disks typically have low viscosity, our results suggest that retrograde S-type planets can form in-situ from retrograde mini disks around polar circumbinary disks, and their occurrence rate may be higher than currently estimated.

astro-ph.EP

Observation of electron spin interactions between Rydberg atoms

We report the observation of electron spin interactions between Rydberg atoms, which are driven by spin-orbit coupling through second-order dipole perturbation and exhibit a spatial anisotropy governed by the atomic configuration. Specifically, we observe coherent electron spin exchange dynamics, with the measured coupling strength agreeing well with both numerical calculations and theoretical models. Furthermore, we show that global microwave dressing enables active engineering and dynamical freezing of the spin exchange by introducing a differential AC Stark shift between the participating states. Additionally, we achieve tunability of the interaction by applying a stronger magnetic field, which effectively modifies the energy contributions of the underlying spin-orbit coupling channels. Finally, measuring spin dynamics in one-dimensional multi-atom chains aligned parallel or perpendicular to the magnetic field provides a self-consistent validation of the anisotropic XXZ framework. This electron spin interaction natively features spin-position coupling and enables a natural mapping onto the Heisenberg-Kitaev model. Our findings reveal a new class of electron spin-spin interactions among Rydberg atoms, expanding the scope of quantum simulation with Rydberg atom arrays.

cond-mat.quant-gas

Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less labor consumption on large-scale datasets. Predominant methods rely on strong prior assumptions to recover the missing semantics from partial annotations. However, these statistic priors suffer from unstable semantic mistakes and thus lead to catastrophic overfitting. Toward this end, we propose a Language-driven Dense Semantic Adaptor (LDSA) that excavates prior-adaptive relationships from multimodal pretrained CLIP models. In our approach, the densely contrastive adaptor is first proposed to construct dense visual contrastive constraints, transferring the task-specific knowledge to visual domains. We then propose a language-driven interactive decoder with the help of class-specific prompt tuning, which adapts language proxies with visual domains. With the collaborative learning of proposed modules, experimental results demonstrate our proposed LDSA achieves a new state of the art on public multi-label classification benchmarks, and interpretable analyses reveal that our LDSA discovers implicit semantic relationships with the prior-adaptive learning scheme.

cs.CV

OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects

Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-language detectors recover long-tail categories and small, fine-grained objects overlooked by closed-set models, they also tend to fragment large surfaces and merge small objects into larger neighboring objects, compromising instance-level consistency and undermining mapping fidelity. Moreover, existing methods struggle to retrieve previously unmapped targets or determine whether a queried object is absent, hindering robust embodied open-world navigation and exploration. We present OVIP-SG, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval. OVIP-SG uses a vision-language model (VLM) to enumerate scene-specific categories for robust open-world detection. Symmetric 3D Intersection over Union (IoU) association and area-weighted feature fusion preserve small independent instances, while VLM-inferred object functions partition scenes into compact functional search regions. A four-stage cascaded retrieval pipeline further incorporates voxel voting and determines target absence from exploration coverage. Under a unified evaluation protocol on Replica, OVIP-SG outperforms ConceptGraphs by 6.31 points in class-mean accuracy (mAcc) and 5.15 points in frequency-weighted mIoU (F-mIoU) while achieving a class-agnostic native-instance Panoptic Quality (PQ) of 0.398. It reduces the search area to 21.8% of the indoor floor space and reaches 0.773 balanced accuracy for object-presence classification. Real-world robotic experiments further demonstrate its practical effectiveness. Code is available at https://github.com/Agibot-Spatial-Intelligence/OVIP-SG.

cs.RO

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at https://github.com/ShareLab-SII/VLZip.

cs.CV

Perturbed Dyadic Cubes and Quantitative Estimates for Schr\"odinger Operators with Potentials in $RH^{n/2}$

Let $L:=-\Delta+V$ be a Schr\"odinger operator on the Euclidean space $\mathbb{R}^n$ with potential $V$ in the reverse H\"older class $RH^{n/2}$ satisfying some mild assumptions that $V$ neither decays too rapidly nor oscillates violently at infinity. In this paper, the authors construct a new system of dyadic cubes $\mathcal{D}^V$ that reflects the intrinsic geometry perturbed by $V$. Then using the quantitative geometric information of $\mathcal{D}^V$, the authors characterize the $L^p$ operator norm of the Riesz potential $L^{-\alpha/2}$ for all $\alpha\in (0,2]$ and $p\in (1,\infty)$. As applications, some quantitative spectral estimates for $L$ are given.

math.AP

Beyond Retrieval: Analytic Memory for Multimodal Agents

Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 6.9\%, respectively.

cs.AI

OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment

Deploying diabetic retinopathy (DR) screening models in primary care requires edge-efficient systems that remain accurate, safe, and reliable under domain shift. Multi-teacher knowledge distillation (KD) is a natural compression strategy, but existing approaches largely assume that all teachers provide equally trustworthy supervision. In our setting, this assumption fails: a strong CNN teacher (EfficientNet-B3, 0.876 QWK) and a weaker Transformer teacher (Swin-Base, 0.830 QWK) are complementary, yet the Transformer's logits can still mislead the student. We therefore propose OrthKD, a selective-trust distillation framework that transfers full supervision from the strong CNN, uses feature-only distillation from the weak ViT, and enforces orthogonality between teacher-specific student projections to encourage complementary rather than redundant evidence. This design preserves local lesion precision, injects global structural context, and improves robustness to distribution shift. On 132,049 retinal images, a 5.4M-parameter MobileNetV3 student reaches 0.885 QWK on EyePACS and improves zero-shot Messidor-2 performance from 0.507 to 0.728 QWK, while also achieving strong referral AUC and calibration. These results show that selectively distilling heterogeneous teachers can enable practical DR screening on resource-constrained devices.

cs.LG

Electromagnetic form factors of singly charmed baryons $\Sigma_c$ and $\Lambda_c$ in a covariant quark-diquark model

We present a systematic study of the spacelike electromagnetic form factors of the ground-state singly charmed baryons, $\Sigma_c$ ($\Sigma_c^{++},\Sigma_c^+,\Sigma_c^0$) and $\Lambda_c^+$, within a covariant quark-diquark model. Based on this framework, we obtain their magnetic moments as well as electric charge and magnetic moment radii. Such observables are important to understand the internal structure and the inner dynamics of these heavy baryon states. Our theoretical calculations are in qualitative agreement with available lattice QCD results of $\Sigma_c^{++}$ and $\Sigma_c^0$ baryons. We also discuss the mechanism behind the dependence of the numerical results on the charm quark and light diquark. A key finding is that the electric form factors of the singly charmed baryons fall off much more slowly with momentum transfer $Q^2$ than that of the proton, indicating a more compact electric charge distribution, which is attributed to the heavy charm quark acting as a localized core. More importantly, we observe a striking difference in the magnetic structure: while the magnetic form factors of $\Sigma_c$ are dominated by the light axial-vector diquark, those of $\Lambda_c^+$ are unexpectedly governed by the charm quark due to the vanishing contribution of the scalar diquark. This highlights the decisive role of the light-diquark spin configuration in determining the magnetic properties. Finally, using an empirical asymptotic relation connecting the spacelike and timelike regions, we predict the total cross section for $e^+e^-\to\Sigma_c\bar{\Sigma}_c$. Our findings can be tested at existing facilities including the BESIII, Belle II and LHCb, as well as the proposed Super Tau-Charm Facility.

hep-ph

LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation

Large Language Models (LLMs) have shown great promise in recommendation due to superior reasoning abilities. However, existing methods mainly rely on explicit Chain-of-Thought (CoT), resulting in verbose reasoning texts and inefficient response times. latent reasoning aims to balance efficiency by thinking within a continuous latent space, yet it faces two major challenges: (1) Lack of Fine-grained Supervision: Latent reasoning relies solely on feedback from the final labels, providing sparse supervisory signals that struggle to effectively guide the optimization of multiple hidden reasoning steps. (2) Single Reasoning Path: The deterministic nature of latent reasoning impedes the exploration of users' diverse interests and preferences, thereby limiting the recommendation capabilities of LLMs. To address these issues, we propose \textbf{$LaRec$}, an efficient generative recommendation framework designed to unleash the potential of latent reasoning in LLMs. $LaRec$ consists of two core stages: First, we design Latent Pre-training that empowers LLMs with latent reasoning capabilities by providing rich supervisory signals to the latent space reasoning via step-level alignment and process direction alignment. Second, we introduce Personalized RL-tuning. Specifically, we construct a personalized Gaussian Mixture Distribution for each user based on their historical interests. By randomly sampling distinct reasoning starting points from this distribution during training, we guide the LLMs to traverse diverse reasoning paths within the latent space, enabling efficient exploration of user's multi-faceted interests. Experiments on multiple datasets show that $LaRec$ significantly outperforms existing baselines with comparable efficiency.

cs.IR

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/

cs.CV

Many-Body Physics with Rydberg Atoms: Quantum Simulation and Non-equilibrium Dynamics

Rydberg atoms, characterized by their strong and long-range dipole-dipole interactions, provide a versatile platform for exploring intriguing collective and many-body effects. Recently, the experimental realization of these effects in dense ensembles and reconfigurable atomic arrays has attracted significant interest, particularly for applications in quantum simulations and non-equilibrium physics. This review focuses on such recent development, discussing the theoretical foundations of the interactions between Rydberg atoms and the ensuing many-body physics, while providing a critical survey of experimental techniques for their precise manipulation and observation. We further discuss recent breakthroughs in leveraging Rydberg collective effects to probe novel many-body phases and non-equilibrium dynamics of these systems. By synthesizing theoretical insights with experimental milestones, we provide a comprehensive perspective on this rapidly evolving field and its transformative potential for future quantum technologies.

quant-ph