SearcharxivSearch

arXiv subjects

Dan Li

Publications and source records attributed to Dan Li.

At least 37 records · Page 2Linked to original sources

Minimally $(k,k)$-edge-connected graphs via spectral radius

For $l > 1$, the $l$-edge-connectivity $\kappa'_l(G)$ of a connected graph $G$ is defined as the minimum number of edges whose removal leaves a graph with at least $l$ components. A graph is minimally $(k,l)$-edge-connected if $\kappa'_l(G)\geq k$ but for any edge $e\in E(G)$ satisfies that $\kappa'_l(G-e)< k$. Motivated by two foundational extremal problems: Brualdi and Solheid's problem [SIAM J. Algebra Discrete Methods (1986)] for graphs of fixed order: determine sharp upper bounds for the spectral radius over graph families and characterize extremal graphs; and its fixed size analogue proposed by Brualdi and Hoffman [Linear Algebra Appl. (1985)], we resolve both problems for minimally $(k,k)$-edge-connected graphs. Building on the structural framework of Hennayake, Lai, Li, and Mao [J. Graph Theory (2003)], we combine edge-switching method and double eigenvectors skill to characterize the graphs maximizing the spectral radius among all minimally $(k,k)$-edge-connected graphs of prescribed order or size. Our results generalize the $k=2$ cases established by Lou, Min, and Huang [Electron. J. Comb. (2023)] and Chen and Guo [Discrete Math. (2019)].

math.SP

Unconventional Magnetism: Symmetry Classification, Hybrid-parity and Unconstrained-parity Classes

Unconventional magnetism has emerged as a transformative frontier in condensed matter physics. Such phases are characterized by substantial non-relativistic spin splitting (NSS) in symmetry-compensated magnets. They have been classified by the parity of their spin textures under momentum inversion, leading to the paradigms of altermagnets (even-parity) and odd-parity magnets. However, the full symmetry landscape remains largely unexplored. In this Letter, we present a systematic classification framework for unconventional magnetism based on the representation theory of the spin textures and the associated parity properties. Within this framework, we predict two previously unidentified classes beyond the established pure-parity categories: hybrid-parity magnets (HPMs) and unconstrained-parity magnets (UPMs), where the spin textures exhibit contrasting parities among their Cartesian components and the parity of the spin textures is ill-defined, respectively. We derive universal symmetry criteria that categorize HPMs into three distinct types. Importantly, by combining the spin splitting characteristics of altermagnets and odd-parity magnets, HPMs can enable the coexistence of the spin current and Edelstein effects. Taking FePO4 as an example, we perform first-principles calculations to demonstrate this coexistence. Finally, we discuss the potential applications of HPMs in spintronic devices. Our work provides a comprehensive symmetry classification of unconventional magnetism and establishes HPMs as a promising platform for multi-functional spintronics.

cond-mat.mtrl-sci

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning

In this paper, we propose the first VL\underline{\textbf{M}} \underline{\textbf{a}}gentic \underline{\textbf{r}}easoning framework for few-\underline{\textbf{s}}hot multimodal \underline{\textbf{T}}ime \underline{\textbf{S}}eries \underline{\textbf{C}}lassification (\textsc{MarsTSC}), which introduces a self-evolving knowledge bank as a dynamic context iteratively refined via reflective agentic reasoning. The framework comprises three collaborative roles: i) Generator conducts reliable classification via reasoning; ii) Reflector diagnoses the root causes of reasoning errors to yield discriminative insights targeting the temporal features overlooked by Generator; iii) Modifier applies verified updates to the knowledge bank to prevent context collapse. We further introduce a test-time update strategy to enable cautious, continuous knowledge bank refinement to mitigate few-shot bias and distribution shift. Extensive experiments across 12 mainstream time series benchmark datasets demonstrate that \method{} delivers substantial and consistent performance gains across 5 VLM backbones, outperforming both classical and foundation model-based time series baselines under few-shot conditions, while producing interpretable rationales that ground each classification decision in human-readable feature evidence. Code is available at https://github.com/HuangJW0821/MarsTSC.

cs.AI

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion Transformer backbone with a fine-grained Mixture-of-Experts (MoE) design (128 experts, Top-8 routing), yielding a 25B-parameter model that activates only 3B parameters, significantly reducing training costs while scaling up the model capacity. Mamoda2.5 achieves top-tier generation performance on VBench 2.0 and sets a new record in video editing quality, surpassing evaluated open-source models and matching the performance of current top-tier proprietary models, including the Kling O1 on OpenVE-Bench. Furthermore, we introduce a joint few-step distillation and reinforcement learning framework that compresses the 30-step editing model into a 4-step model and greatly accelerates model inference. Compared to open-source baselines, Mamoda2.5 achieves up to $95.9\times$ faster video editing inference. In real-world applications, Mamoda2.5 has been successfully deployed for content moderation and creative restoration tasks in advertising scenarios, achieving a 98% success rate in internal advertising video editing scenario.

cs.CV

Human-Less LLM Serving: Quantifying the Human Tax on Throughput

Every major LLM serving system is designed to meet TTFT and TPOT SLOs. These metrics capture latency as a human user perceives it, and the mechanisms built to satisfy them are now standard infrastructure. We observe that long-horizon AI tasks call LLMs programmatically in tight loops where no human observes TTFT or TPOT. We ask: how much throughput do serving systems sacrifice to meet TTFT and TPOT SLAs that these workloads never need? We conduct a systematic measurement study across chunk sizes, SLO settings, context lengths, and concurrency levels. We find that the human tax on throughput grows substantially with context length and lands in the 60-93% range. At 64K token contexts, tightening the TTFT SLO to production-typical settings costs a large fraction of throughput versus the human-less baseline. The human tax is larger at higher concurrency and is qualitatively similar across SGLang and Sarathi-Serve. We term the unconstrained optimum human-less serving and provide a prototype demonstrating that it is practical on real workloads. Our findings argue that serving systems should expose workload-class-aware SLA configurations rather than silently applying the human tax uniformly to all traffic.

cs.NI

Reshaping the inner shadow of a Kerr black hole by a torn accretion disk

When an accretion flow extends to the event horizon, their intersection defines the contour of the inner shadow. However, the morphological evolution of this critical feature remains largely unexplored within a torn accretion disk system, a configuration comprising distinct sub-disks formed when a tilted disk is disrupted by frame-dragging. To address this, we phenomenologically construct a torn accretion disk model and numerically simulate the inner shadow of a Kerr black hole using relativistic backward ray-tracing. We discover that the torn disk geometry profoundly alters the black hole's observational signatures, inducing severe erosion of the inner shadow and generating novel features such as bifurcated shadows, crescent-like structures, and multiple orders of shadow rings. These exotic morphologies, which are predominantly governed by the spatial discontinuity between the sub-disks and the tilt angle of the outer sub-disk, are exceedingly difficult to replicate within standard equatorial accretion paradigms. Our findings demonstrate that these distinctive shadow structures hold significant potential to serve as robust diagnostic probes for torn accretion environments, simultaneously implying that relying solely on the inner shadow to test gravity theories is fundamentally insufficient.

gr-qc

Gravitational emissions and light curves of quasi-periodic orbits in Schwarzschild spacetime embedded in a Dehnen-type dark matter halo

Timelike orbits in curved spacetimes encode intrinsic information about the background geometry and serve as critical probes for investigating gravitational theories and source distributions. In this study, we investigate strictly closed timelike orbits within a Schwarzschild spacetime embedded in a Dehnen-type dark matter halo. By solving the geodesic equations, we identify various configurations of these closed orbits and simulate their corresponding gravitational waves and electromagnetic light curves. Our findings reveal that the morphology of closed orbits is primarily governed by the ratio of the azimuthal period to the radial period. Notably, dark matter halo parameters such as the core scale and density parameters exert a significant amplification effect on the orbital scale, which further induces a discernible phase lag in the gravitational wave signals. Furthermore, within a specific parameter space, we discover a linear relationship between the number of peaks in the light curves and the number of orbital leaves. From a theoretical perspective, these findings reveal the multimessenger signatures of closed orbits, which may provide a potential theoretical foundation for establishing a connection between orbital dynamics and the surrounding dark matter environment.

gr-qc

WaveMoE: A Wavelet-Enhanced Mixture-of-Experts Foundation Model for Time Series Forecasting

Time series foundation models (TSFMs) have recently achieved remarkable success in universal forecasting by leveraging large-scale pretraining on diverse time series data. Complementing this progress, incorporating frequency-domain information yields promising performance in enhancing the modeling of complex temporal patterns, such as periodicity and localized high-frequency dynamics, which are prevalent in real-world time series. To advance this direction, we propose a new perspective that integrates explicit frequency-domain representations into scalable foundation models, and introduce WaveMoE, a wavelet-enhanced mixture-of-experts foundation model for time series forecasting. WaveMoE adopts a dual-path architecture that jointly processes time series tokens and wavelet tokens aligned along a unified temporal axis, and coordinates them through a shared expert routing mechanism that enables consistent expert specialization while efficiently scaling model capacity. Preliminary experimental results on 16 diverse benchmark datasets indicate that WaveMoE has the potential to further improve forecasting performance by incorporating wavelet-domain corpora.

cs.LG

Democratizing Federated Learning with Blockchain and Multi-Task Peer Prediction

The synergy between Federated Learning and blockchain has been considered promising; however, the computationally intensive nature of contribution measurement conflicts with the strict computation and storage limits of blockchain systems. We propose a novel concept to decentralize the AI training process using blockchain technology and Multi-task Peer Prediction. By leveraging smart contracts and cryptocurrencies to incentivize contributions to the training process, we aim to harness the mutual benefits of AI and blockchain. We discuss the advantages and limitations of our design.

cs.CR

MsFormer: Enabling Robust Predictive Maintenance Services for Industrial Devices

Providing reliable predictive maintenance is a critical industrial AI service essential for ensuring the high availability of manufacturing devices. Existing deep-learning methods present competitive results on such tasks but lack a general service-oriented framework to capture complex dependencies in industrial IoT sensor data. While Transformer-based models show strong sequence modeling capabilities, their direct deployment as robust AI services faces significant bottlenecks. Specifically, streaming sensor data collected in real-world service environments often exhibits multi-scale temporal correlations driven by machine working principles. Besides, the datasets available for training time-to-failure predictive services are typically limited in size. These issues pose significant challenges for directly applying existing models as robust predictive services. To address these challenges, we propose MsFormer, a lightweight Multi-scale Transformer designed as a unified AI service model for reliable industrial predictive maintenance. MsFormer incorporates a Multi-scale Sampling (MS) module and a tailored position encoding mechanism to capture sequential correlations across multi-streaming service data. Additionally, to accommodate data-scarce service environments, MsFormer adopts a lightweight attention mechanism with straightforward pooling operations instead of self-attention. Extensive experiments on real-world datasets demonstrate that the proposed framework achieves significant performance improvements over state-of-the-art methods. Furthermore, MsFormer outperforms across industrial devices and operating conditions, demonstrating strong generalizability while maintaining a highly reliable Quality of Service (QoS).

cs.LG

The Steiner Tree Problem: Novel QUBO Formulation and Quantum Annealing Implementation

The Steiner Tree Problem (STP) is a well-known NP-hard combinatorial optimization problem, which has wide applications in network design, integrated circuit layout, bioinformatics, and other fields. However, traditional algorithms often struggle to balance efficiency and solution quality when dealing with large-scale STP instances. In this paper, we propose a new quantum annealing-based algorithm for solving the STP: we first model the STP into a quadratic unconstrained binary optimization (QUBO) form suitable for quantum annealing, then design a corresponding encoding strategy, and finally verify the algorithm through experimental tests. The results show that our quantum annealing-based method can obtain high-quality solutions with relatively low computational overhead for moderate-scale STP instances, providing a new feasible path for handling this intractable combinatorial optimization problem.

quant-ph

Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-Forwarding

Packet-level discrete-event simulation (PLDES) is a prevalent tool for evaluating detailed performance of large model training. Although PLDES offers high fidelity and generality, its slow performance has plagued networking practitioners. Existing optimization techniques either simplify the network model, resulting in large errors; or execute it in parallel using multiple processors, with an upper bound on speedup. This paper explores an alternative optimization direction that reduces the computational loads of PLDES while maintaining high fidelity. Our key insight is that, in distributed LLM training, packet-level traffic behaviors often exhibit repetitive contention patterns and steady-states where flow rates stabilize, ignoring these redundant discrete events speeds up the simulation considerably and the error is negligible. We realize this idea by proposing Wormhole, a user-transparent PLDES kernel capable of automatically memoization for unsteady-states and skipping for steady-states. Wormhole adopts network partitioning, state memoization and reuse, and rate-based steady-state identification to accurately determine the periods of each flow's steady-state, while maintaining simulation consistency after fast-forwarding. Experiments demonstrate that Wormhole can achieve a 744x speedup over the original ns-3 (510x for MoE workload), with a bounded error of <1%. Applying current multithreading parallel techniques and Wormhole together allows a 1012x speedup, reducing the simulation time for one GPT-13B training under 128 GPUs from 9 hours to 5 minutes.

cs.NI

Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning

Time series is a pervasive data type across various application domains, rendering the reasonable solving of diverse time series tasks a long-standing goal. Recent advances in large language models (LLMs), especially their reasoning abilities unlocked through reinforcement learning (RL), have opened new opportunities for tackling tasks with long Chain-of-Thought (CoT) reasoning. However, leveraging LLM reasoning for time series remains in its infancy, hindered by the absence of carefully curated time series CoT data for training, limited data efficiency caused by underexplored data scheduling, and the lack of RL algorithms tailored for exploiting such time series CoT data. In this paper, we introduce VeriTime, a framework that tailors LLMs for time series reasoning through data synthesis, data scheduling, and RL training. First, we propose a data synthesis pipeline that constructs a TS-text multimodal dataset with process-verifiable annotations. Second, we design a data scheduling mechanism that arranges training samples according to a principled hierarchy of difficulty and task taxonomy. Third, we develop a two-stage reinforcement finetuning featuring fine-grained, multi-objective rewards that leverage verifiable process-level CoT data. Extensive experiments show that VeriTime substantially boosts LLM performance across diverse time series reasoning tasks. Notably, it enables compact 3B, 4B models to achieve reasoning capabilities on par with or exceeding those of larger proprietary LLMs.

cs.AI

EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents

Recent advances in Multimodal Large Language Models (MLLMs) have enabled agents to operate in open-ended web and operating system environments. However, existing benchmarks predominantly target consumer-oriented scenarios (e.g., e-commerce and travel booking), failing to capture the complexity and rigor of professional enterprise workflows. Enterprise systems pose distinct challenges, including high-density user interfaces, strict business logic constraints, and a strong reliance on precise, state-consistent information retrieval-settings in which current generalist agents often struggle. To address this gap, we introduce EntWorld, a large-scale benchmark consisting of 1,756 tasks across six representative enterprise domains, including customer relationship management (CRM), information technology infrastructure library (ITIL), and enterprise resource planning (ERP) systems. Unlike previous datasets that depend on fragile execution traces or extensive manual annotation, EntWorld adopts a schema-grounded task generation framework that directly reverse-engineers business logic from underlying database schemas, enabling the synthesis of realistic, long-horizon workflows. Moreover, we propose a SQL-based deterministic verification mechanism in building datasets that replaces ambiguous visual matching with rigorous state-transition validation. Experimental results demonstrate that state-of-the-art models (e.g., GPT-4.1) achieve 47.61% success rate on EntWorld, substantially lower than the human performance, highlighting a pronounced enterprise gap in current agentic capabilities and the necessity of developing domain-specific agents. We release EntWorld as a rigorous testbed to facilitate the development and evaluation of the next generation of enterprise-ready digital agents.

cs.AI

Understanding and Preserving Safety in Fine-Tuned LLMs

Fine-tuning is an essential and pervasive functionality for applying large language models (LLMs) to downstream tasks. However, it has the potential to substantially degrade safety alignment, e.g., by greatly increasing susceptibility to jailbreak attacks, even when the fine-tuning data is entirely harmless. Despite garnering growing attention in defense efforts during the fine-tuning stage, existing methods struggle with a persistent safety-utility dilemma: emphasizing safety compromises task performance, whereas prioritizing utility typically requires deep fine-tuning that inevitably leads to steep safety declination. In this work, we address this dilemma by shedding new light on the geometric interaction between safety- and utility-oriented gradients in safety-aligned LLMs. Through systematic empirical analysis, we uncover three key insights: (I) safety gradients lie in a low-rank subspace, while utility gradients span a broader high-dimensional space; (II) these subspaces are often negatively correlated, causing directional conflicts during fine-tuning; and (III) the dominant safety direction can be efficiently estimated from a single sample. Building upon these novel insights, we propose safety-preserving fine-tuning (SPF), a lightweight approach that explicitly removes gradient components conflicting with the low-rank safety subspace. Theoretically, we show that SPF guarantees utility convergence while bounding safety drift. Empirically, SPF consistently maintains downstream task performance and recovers nearly all pre-trained safety alignment, even under adversarial fine-tuning scenarios. Furthermore, SPF exhibits robust resistance to both deep fine-tuning and dynamic jailbreak attacks. Together, our findings provide new mechanistic understanding and practical guidance toward always-aligned LLM fine-tuning.

cs.LG

Manipulating Anomalous Transport via Crystal Symmetry in 2D Altermagnets

Anomalous transports, including the anomalous Hall effect (AHE) and anomalous Nernst effect (ANE), are typical manifestations of time-reversal-symmetry-breaking responses in materials. In general, the two Hall states with opposite Hall conductivities can be regarded as time-reversal pairs coupled to magnetic order, and switching between them relies on reversing the magnetization via an external magnetic field or electric current. Here, we introduce a approach for manipulating anomalous transport through crystal symmetry engineering in two-dimensional (2D) altermagnetic systems. Based on symmetry analysis, we demonstrate that 2D altermagnets (AM) with out-of-plane N\'eel vectors will not host any anomalous Hall transport. Remarkably, breaking the symmetry connecting the two magnetic sublattices, an anomalous Hall response can emerge immediately, and the signs of the anomalous Hall and anomalous Nernst conductivities can be flexibly controlled by the symmetry-breaking term, thereby realizing tunable sign-reversible anomalous transport. Furthermore, the feasibility of the theoretical scheme is further verified by explicit lattice-model construction. Using first-principles calculations, we investigate the realization of crystal symmetry-controlled anomalous transport in a 2D AM material Cr$_{2}$O$_{2}$. The results indicate that Cr$_{2}$O$_{2}$ with out-of-plane N\'eel vectors can sequentially exhibit the AHE and quantum anomalous Hall effect (QAHE) under continuous uniaxial strain. Interestingly, the sign reversal between these two effects can be achieved by simply rotating the strain direction by C$_{4z}$ symmetry. The corresponding ANE and its sign reversal are also revealed. Our findings provide a new strategy to manipulate anomalous transport, and should have significant potential applications.

cond-mat.mes-hall

Effect of superconductivity by Nb and V substitution in kagome CaPd5

Materials featuring kagome lattices have attracted significant research interest due to their unique geometric frustration, which gives rise to rich physical phenomena such as non-trivial topology, spin fluctuations, and superconductivity. In this work, using CaPd5 as the prototype structure, we discover and systematically investigate a new class of kagome superconductors, CaMxPd5-x (M = Nb and V) alloys. First-principles calculations confirm that these compounds are non-magnetic metals, among which four are dynamically stable: CaNb5, CaV5, CaNb2Pd3, and CaV2Pd3. CaNb5 is identified as a strong electron-phonon coupling (EPC) superconductor with the highest superconducting transition temperature (Tc) of 10.1 K, which can be further increased to 12.8 K under external pressure. In contrast, CaV5, CaNb2Pd3, and CaV2Pd3 exhibit weaker EPC and correspondingly lower Tc values. Furthermore, by applying the method of symmetry indicators, we systematically classify the topological and nodal characteristics of CaNb5, providing valuable insights for determining its superconducting pairing symmetry. Our findings demonstrate that Nb and V substitution in kagome CaPd5 provides an effective route for designing a new type of kagome superconductor with relatively high Tc. This study also offers new perspectives on topological superconductivity in kagome systems and establishes a useful guideline for discovering other superconducting materials with unique properties.

cond-mat.supr-con

Switchable Giant Spin Injection Current in Janus Altermagnet Fe$_2$SSeO

Generating and controlling spin current in miniaturized magnetic quantum devices remains a central objective of spintronics, due to its potential to enable future energy-efficient information technologies. Among the existing magnetic phases, altermagnetism have recently emerged as a highly promising platform for spin current generation and control, going beyond ferromagnetism and antiferromagnetism. Here, we propose a symmetry-allowed spin photovoltaic effect in two-dimensional (2D) altermagnetic semiconductors that enables predictable control of giant spin injection currents. Distinct from parity-time ($\mathcal{PT}$)-antiferromagnets, Janus altermagnetic semiconductors generate not only shift current but also a unique injection current with spin momentum locked in a specific direction under linearly polarized light -- a mechanism absent in $\mathcal{PT}$-antiferromagnets. Through symmetry analysis and first-principles calculations, we identify Janus Fe$_2$SSeO as a promising candidate. Specifically, the monolayer Fe$_2$SSeO exhibits a polarization-dependent injection conductivity reaching $\sim$1,200~$\mu$A/V$^{2}\!\cdot\!\hbar/2e$, and the giant spin injection current can be effectively switched by rotating the magnetization direction and engineering strains. These findings underscore the potential of 2D altermagnets in spin photovoltaics and open avenues for innovative quantum devices.

cond-mat.mtrl-sci