SearcharxivSearch

arXiv subjects

Ziwen Chen

Publications and source records attributed to Ziwen Chen.

7 recordsLinked to original sources

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.

cs.CV

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Dermatologic diseases impose a large and growing global burden, affecting billions and substantially reducing quality of life. While modern therapies can rapidly control acute symptoms, long-term outcomes are often limited by single-target paradigms, recurrent courses, and insufficient attention to systemic comorbidities. Traditional Chinese medicine (TCM) provides a complementary holistic approach via syndrome differentiation and individualized treatment, but practice is hindered by non-standardized knowledge, incomplete multimodal records, and poor scalability of expert reasoning. We propose DERM-3R, a resource-efficient multimodal agent framework to model TCM dermatologic diagnosis and treatment under limited data and compute. Based on real-world workflows, we reformulate decision-making into three core issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents: DERM-Rec, DERM-Rep, and DERM-Reason, each targeting one component of this pipeline. Built on a lightweight multimodal LLM and partially fine-tuned on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across dermatologic reasoning tasks. Evaluations using automatic metrics, LLM-as-a-judge, and physician assessment show that despite minimal data and parameter updates, DERM-3R matches or surpasses large general-purpose multimodal models. These results suggest structured, domain-aware multi-agent modeling can be a practical alternative to brute-force scaling for complex clinical tasks in dermatology and integrative medicine.

cs.AI

Generating 360° Video is What You Need For a 3D Scene

Generating 3D scenes is still a challenging task due to the lack of readily available scene data. Most existing methods only produce partial scenes and provide limited navigational freedom. We introduce a practical and scalable solution that uses 360° video as an intermediate scene representation, capturing the full-scene context and ensuring consistent visual content throughout the generation. We propose WorldPrompter, a generative pipeline that synthesizes traversable 3D scenes from text prompts. WorldPrompter incorporates a conditional 360° panoramic video generator, capable of producing a 128-frame video that simulates a person walking through and capturing a virtual environment. The resulting video is then reconstructed as Gaussian splats by a fast feedforward 3D reconstructor, enabling a true walkable experience within the 3D scene. Experiments demonstrate that our panoramic video generation model, trained with a mix of image and video data, achieves convincing spatial and temporal consistency for static scenes. This is validated by an average COLMAP matching rate of 94.6\%, allowing for high-quality panoramic Gaussian splat reconstruction and improved navigation throughout the scene. Qualitative and quantitative results also show it outperforms the state-of-the-art 360° video generators and 3D scene generation models.

cs.GR

FERA: A Pose-Based Framework for Rule-Grounded Multimedia Decision Support with a Foil Fencing Case Study

Multimedia decision support requires more than recognition; it requires explicit state estimates that can be checked against rules, audited by humans, and consumed by downstream decision logic. We present the FEncing Referee Assistant (FERA), a pose-based framework for this setting, and study it through foil fencing, where decisions depend on fast bilateral motion and right-of-way rules. The framework separates canonical participant tracking, kinematic tokenization, calibrated temporal perception, a compact structured decision layer, and an explanation-oriented retrieval interface. We also release an audited benchmark with adjudicated labels and fixed folds for reproducible evaluation. Under a shared protocol, a lightweight lifted-depth sidecar strengthens the best graph-based perception model, while a compact structured classifier on the fixed two-dimensional token stream reaches 0.624 accuracy and a 0.632 macro-averaged F1 score on the final Left / Right / None decision. The case study supports a broader design lesson: keep the boundary between perception and rule application explicit, preserve uncertainty, and choose the perception front end according to the downstream operating point.

cs.AI

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth's properties for enhancing model capability, training stability, and generalization, as well as application to other tasks.

cs.CV

Modularity, Architectural Innovation, and New Venture Success

Startups face a classic dilemma in innovation strategy: should they pursue cumulative, low-risk improvements or disruptive, high-risk breakthroughs? The Henderson and Clark framework suggests that architectural innovation, which reconfigures existing economic modules in novel ways, tends to be disruptive and risky for established organizations, but the success of this strategy for entrepreneurs remains less well understood, largely based on methodological constraints. Building on a complex-economics perspective and advanced computational models, we distinguish architectural innovation from modular innovation, which incrementally updates economic modules, and modular invention, which forges new ones, within the entrepreneurship context. Then we examine how each strategy influences startup performance. We analyze 298,915 U.S. venture-funded start-ups from 1976-2020, embedding company descriptions within a dynamic semantic space constructed from business and patent discourse to measure innovation structure across the entire economy. Event history models reveal that architectural innovation leads to successful IPOs and high-value acquisitions, while both modular innovation and invention increase the risk of failure. By comparing the outcomes of architectural and modular innovation and invention, this paper reveals that what is typically seen as the riskiest form of innovation can, for startups, be the safest route to success. This reconceptualization inverts the trade-off between exploration-exploitation typically assumed in organizational learning with critical implications for entrepreneurial strategy and innovation policy.

econ.EM

Destructive Creation, Creative Destruction, and the Paradox of Innovation Science

Innovation or the creation and diffusion of new material, social and cultural things in society has been widely studied in sociology and across the social sciences, with investigations sufficiently diverse and dispersed to make them unnavigable. This complexity results from innovation's importance for society, but also the fundamental paradox underlying innovation science: When innovation becomes predictable, it ceases to be an engine of novelty and change. Here we review innovation studies and show that innovations emerge from contexts of discord and disorder, breaches in the structure of prior success, through a process we term destructive creation. This often leads to a complementary process of creative destruction whereby local structures protect and channel the diffusion of successful innovations, rendering alternatives obsolete. We find that social scientists naturally focus far more on how social and cultural contexts influence material innovations than the converse. We highlight computational tools that open new possibilities for the analysis of novel content and context in interaction, and show how this brings us empirically toward the broader range of possibilities that complex systems and science studies have theorized-and science fiction has imagined-the social, cultural and material structures of innovation conditioning each other's change through cycles of disruption and development.

physics.soc-ph