SearcharxivSearch

arXiv subjects

Xuan Wei

Publications and source records attributed to Xuan Wei.

11 recordsLinked to original sources

Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heterogeneous boundaries pose a deployment challenge, but also create an opportunity: strategically coordinating multiple LLMs may unlock collective intelligence exceeding any single model. Existing approaches fix how models are combined in advance, overlooking the dynamic, state-dependent role of complementarity in complex problem solving. Drawing on the wisdom-of-crowds paradigm, we reconceptualize collective LLM intelligence as relay-style complementarity: a sequential process in which each successor model is selected to address the specific bottleneck identified in its predecessor's output. To operationalize this, we propose WILC (Wisdom Integration of LLM Crowds), a framework grounded in two design principles. First, iterative reflection-and-refinement establishes a state-preserving workflow through which models diagnose and refine prior outputs. Second, complementarity-driven model selection governs transitions via a dual-gate mechanism: prospective complementarity fit (PCF) identifies the worker most suited to the current bottleneck, while posterior complementarity gain (PCG) evaluates whether the selected transition improves the evolving solution. Experiments across four diverse benchmarks show that WILC outperforms existing approaches, including single-model self-refinement, ensemble methods, and query-routing methods. Under standardized pricing assumptions, WILC matches the average benchmark performance of GPT-5.2 at roughly 7 times lower estimated per-query cost, while facilitating data sovereignty through self-hosted deployment. This study extends wisdom-of-crowds theory from static aggregation to sequential AI complementarity and provides transferable design principles for multi-AI coordination.

cs.AI

Memento: Reconstruct to Remember for Consistent Long Video Generation

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten. In this paper, we propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. Specifically, Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, recovering target appearances using historical memory and global story captions. To disentangle long-range subject evidence from short-range cues, Memento introduces a dual-query memory mechanism, where one query retrieves identity-relevant memory and the other selects short-context keyframes for coherent continuation. Additionally, a subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.

cs.CV

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applications in talking-head synthesis, co-speech gesture generation, and dynamic presentations. Moving beyond conventional keypoint-based methods that often struggle to capture subtle motion dynamics, We propose a novel implicit-motion framework for generating realistic and temporally coherent human motion videos from a single static image and audio. Our approach uses a two-stage pipeline that decouples motion prediction from rendering. The first stage integrates appearance priors and hierarchical depth cues into a region-aware attention mechanism to model latent motion features. The second stage employs a Mamba-enhanced diffusion model to directly predict these features from audio and the source image, enabling unsupervised learning of fine-grained motion patterns. This decoupled architecture enhances flexibility and efficiency. Trained on a new 380-hour high-quality dataset, our method outperforms prior work across multiple public benchmarks and our collected data in accuracy, naturalness, and temporal coherence, setting a new state-of-the-art.

cs.CV

Native Audio-Visual Alignment for Generation

Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on either dual-tower designs with posterior alignment or fully unified tri-modal designs that mix textual context, audio and video in one shared space. The former weakens fine-grained audio-video co-evolution, while the latter couples semantic conditioning with low-level synchronization. To address these limitations, we propose NAVA, a Native Audio-Visual Alignment framework for joint audio-video generation. NAVA is built upon context-conditioned native audio-visual alignment: it first establishes audio-video correspondence in a dedicated interaction space, and then uses external context to condition the joint denoising process. Specifically, NAVA is instantiated with an Align-then-Fuse MMDiT architecture, which transitions from modality-aware audio-video alignment to modality-shared joint denoising. Furthermore, we introduce Timbre-in-Context Conditioning to associate reference timbre cues with corresponding speech spans to achieve controllable speech timbre. Experiments on Verse-Bench and Seed-TTS, together with a user study, demonstrate that NAVA achieves superior video quality, precise audio-visual synchronization, competitive audio quality, and stronger reference-timbre controllability using only 6.3B parameters.

cs.CV

Research on the central region of quasars based on variability and structure function

Quasars,asextremelyluminousanddistantspecialcelestialbodiesintheuniverse,aredrivenbyacomplexsystemcomposedof supermassiveblackholesandsurroundingaccretiondisks.Thispaperadoptsatime-domainobservationstrategyandcombines the analysis of light curves with the construction of structure functions to indirectly reveal the physical essence of the central regionofquasarsfromtheperspectiveofvariability.Theresearchdataarederivedfromthelargesampleobservationdataofthe SloanDigitalSkySurvey(SDSS).Throughextensivedatastatisticsandcorrelationanalysis,aseriesofimportantfindingshave been obtained: the characteristic parameters of the structure function of quasars show significant correlations with luminosity, black hole mass, and Eddington ratio. That is, quasars with higher luminosity, larger black hole mass, and larger Eddington ratiohavelargerstructurefunctions.Forquasarsofthesameluminosity,thelargertheEddingtonratio,thesmallerthestructure function. However, the correlation between the structure function and redshift or rest wavelength is not significant, indicating that the variabilitycharacteristicsofquasars aremainly determined bytheir own physical propertiesandareminimallyaffected by the cosmologicalredshifteffect.

astro-ph.GA

The characteristics of variability of AGNs based on the structure function

Variability is one of the classic features of active galactic nuclei (AGNs). The normalized structure function was applied to distinguish variability samples from OVRO, ASAS-SN and Fermi. A power-law function model was selected to fit the structure functions of samples of three bands. We present the available samples of three bands, and by integrating two parameters, we obtain ideal discrimination results for three bands. Meanwhile, the differences between BL Lacs and FSRQs of Fermi and non-Fermi samples are well verified. The results show that the improved structure function can effectively distinguish samples of radio, optical, and gamma-ray. Additionally, BL Lacs and FSRQs in both Fermi and non-Fermi samples can be distinguished. The conclusion obtained through the distinction of structural functions in different bands supports that the variability in the three bands are caused by different physical mechanisms respectively: the samples in the optical band are radio quiet AGNs, and their variability is mainly caused by the fluctuations of the accretion disk, and the samples of radio band and gamma-ray band are radio loud AGNs whose variability is mainly caused by relativistic jet radiation. This conclusion conforms to the unified standard interpretation of variability about AGNs. Using these two parameters, we verify that there is no fundamental difference between Fermi and non-Fermi BL Lacs, while significant differences exist between FSRQs. However, the power exponent of the two can well distinguish BL Lacs.

astro-ph.HE

State-dependent broadband X-ray Timing Reconfiguration in the Changing-look AGN NGC 1566

NGC 1566 has shown dramatic X-ray spectral changes during its recent changing-look outburst, but the evolution of its broadband X-ray timing properties remains poorly constrained. We combine long-term Swift/XRT monitoring with high-time-resolution XMM-Newton observations to construct one pre-outburst Dim broadband PSD and two outburst broadband reconstructions associated with the O1 peak and O2 decay observations. In the outburst reconstructions, the same Swift/XRT up-state monitoring segment provides the low-frequency constraint, while the O1 and O2 XMM-Newton observations provide phase-specific high-frequency constraints. Using PSRESP forward modelling with the observed sampling windows, we test bending-power-law PSD models in the soft (0.3-2 keV) and hard (2-10 keV) bands. The Dim and O2-associated reconstructions are acceptably described by bending-power-law solutions, whereas the O1 peak observation does not yield a robust bend-frequency measurement. For the accepted Dim and O2-associated solutions, the preferred bend frequency shifts from about 2.0 x 10^-5 to about 2.7 x 10^-7 Hz in the soft band, and from about 2.1 x 10^-5 to about 2.7 x 10^-7 Hz in the hard band, implying a substantially longer characteristic variability timescale in the O2-associated reconstruction. This consistent shift in both energy bands suggests that the timing evolution is not confined to the soft-excess component alone, but reflects a broader change in the X-ray variability structure. Together with previous spectral studies, these results point to a transient reconfiguration of the disc-corona variability timescale during the changing-look transition in NGC 1566.

astro-ph.HE

Timescale-dependent Optical Variability of Turn-on Changing-look Active Galactic Nuclei

Changing-look active galactic nuclei (CL AGNs) provide a valuable opportunity to study optical variability associated with changes in accretion state. We investigate the optical variability of turn-on CL AGNs, selected through the emergence or strengthening of broad emission lines, using g-band light curves from the Zwicky Transient Facility. For a final sample of 106 objects, we measure structure-function amplitudes at three fixed rest-frame timescales, SF30, SF150 and SF300, and examine their dependence on optical luminosity, Eddington ratio, black-hole mass and rest-frame wavelength using Spearman-rank correlations and multivariate regressions. The strongest trends are found on monthly timescales. SF30 is significantly anti-correlated with both optical luminosity and Eddington ratio, with Spearman coefficients of rho = -0.39 and rho = -0.33, respectively. These negative trends remain in multivariate regressions after accounting for black-hole mass and rest-frame wavelength. The dependence weakens at longer timescales: SF150 shows only weak evidence for luminosity or Eddington-ratio dependence, while SF300 shows no robust dependence. Black-hole mass shows no robust single-parameter correlation, and its multivariate coefficients depend on the model form, indicating that its apparent role is affected by covariance with luminosity and Eddington ratio. The rest-frame wavelength term is generally negative, but cannot be separated from redshift in this single-band analysis. These results suggest that monthly optical variability in turn-on CL AGNs is more closely linked to bright-state accretion properties than variability on half-year to year-long timescales.

astro-ph.GA

Optimizing Prompts for Large Language Models: A Causal Approach

Large Language Models (LLMs) are increasingly embedded in enterprise workflows, yet their performance remains highly sensitive to prompt design. Automatic Prompt Optimization (APO) seeks to mitigate this instability, but existing approaches face two persistent challenges. First, commonly used prompt strategies rely on static instructions that perform well on average but fail to adapt to heterogeneous queries. Second, more dynamic approaches depend on offline reward models that are fundamentally correlational, confounding prompt effectiveness with query characteristics. We propose Causal Prompt Optimization (CPO), a framework that reframes prompt design as a problem of causal estimation. CPO operates in two stages. First, it learns an offline causal reward model by applying Double Machine Learning (DML) to semantic embeddings of prompts and queries, isolating the causal effect of prompt variations from confounding query attributes. Second, it utilizes this unbiased reward signal to guide a resource-efficient search for query-specific prompts without relying on costly online evaluation. We evaluate CPO across benchmarks in mathematical reasoning, visualization, and data analytics. CPO consistently outperforms human-engineered prompts and state-of-the-art automated optimizers. The gains are driven primarily by improved robustness on hard queries, where existing methods tend to deteriorate. Beyond performance, CPO fundamentally reshapes the economics of prompt optimization: by shifting evaluation from real-time model execution to an offline causal model, it enables high-precision, per-query customization at a fraction of the inference cost required by online methods. Together, these results establish causal inference as a scalable foundation for reliable and cost-efficient prompt optimization in enterprise LLM deployments.

cs.AI

Bi-level Mixed-Integer Nonlinear Optimization for Pelagic Island Microgrid Group Energy Management Considering Uncertainty

To realize the safe, economical and low-carbon operation of the pelagic island microgrid group, this paper develops a bi-level energy management framework in a joint energy-reserve market where the microgrid group (MG) operator and renewable and storage aggregators (RSA) are independent stakeholders with their own interests. In the upper level, MG operator determines the optimal transaction prices with aggregators to minimize MG operation cost while ensuring all safety constraints are satisfied under uncertainty. In the lower level, aggregators utilize vessels for batteries swapping and transmission among islands in addition to energy arbitrage by participating in energy and reserve market to maximize their own revenue. An upper bound tightening iterative algorithm is proposed for the formulated problem with nonlinear terms and integer variables in the lower level to improve the efficiency and reduce the gap between upper bound and lower bound compared with existing reformulation and decomposition algorithm. Case studies validate the effectiveness of the proposed approach and demonstrate its advantage of the proposed approach in terms of optimality and computation efficiency, compared with other methods.

math.OC

Multi-Label Annotation Aggregation in Crowdsourcing

As a means of human-based computation, crowdsourcing has been widely used to annotate large-scale unlabeled datasets. One of the obvious challenges is how to aggregate these possibly noisy labels provided by a set of heterogeneous annotators. Another challenge stems from the difficulty in evaluating the annotator reliability without even knowing the ground truth, which can be used to build incentive mechanisms in crowdsourcing platforms. When each instance is associated with many possible labels simultaneously, the problem becomes even harder because of its combinatorial nature. In this paper, we present new flexible Bayesian models and efficient inference algorithms for multi-label annotation aggregation by taking both annotator reliability and label dependency into account. Extensive experiments on real-world datasets confirm that the proposed methods outperform other competitive alternatives, and the model can recover the type of the annotators with high accuracy.

cs.LG