SearcharxivSearch

arXiv subjects

Ziyang Lin

Publications and source records attributed to Ziyang Lin.

3 recordsLinked to original sources

Accelerating Multi-modal LLM Gaming Performance via Input Prediction and Mishit Correction

Real-time sequential control agents are often bottlenecked by inference latency. Even modest per-step planning delays can destabilize control and degrade overall performance. We propose a speculation-and-correction framework that adapts the predict-then-verify philosophy of speculative execution to model-based control with TD-MPC2. At each step, a pretrained world model and latent-space MPC planner generate a short-horizon action queue together with predicted latent rollouts, allowing the agent to execute multiple planned actions without immediate replanning. When a new observation arrives, the system measures the mismatch between the encoded real latent state and the queued predicted latent. For small to moderate mismatch, a lightweight learned corrector applies a residual update to the speculative action, distilled offline from a replanning teacher. For large mismatch, the agent safely falls back to full replanning and clears stale action queues. We study both a gated two-tower MLP corrector and a temporal Transformer corrector to address local errors and systematic drift. Experiments on the DMC Humanoid-Walk task show that our method reduces the number of planning inferences from 500 to 282, improves end-to-end step latency by 25 percent, and maintains strong control performance with only a 7.1 percent return reduction. Ablation results demonstrate that speculative execution without correction is unreliable over longer horizons, highlighting the necessity of mismatch-aware correction for robust latency reduction.

cs.AI

Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI

With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as these models operate in resource constrained settings and directly interact with users. Yet, modern AI models routinely exceed the resource capacity of individual devices, necessitating distributed execution across heterogeneous devices over variable and contention prone networks. Existing planners for hybrid (e.g., data and pipeline) parallelism largely optimize for throughput or device utilization, overlooking QoE, leading to severe resource inefficiency (e.g., unnecessary energy drain) or QoE violations under runtime dynamics. We present Dora, a framework for QoE aware hybrid parallelism in distributed edge AI training and inference. Dora jointly optimizes heterogeneous computation, contention prone networks, and multi dimensional QoE objectives via three key mechanisms: (i) a heterogeneity aware model partitioner that determines and assigns model partitions across devices, forming a compact set of QoE compliant plans; (ii) a contention aware network scheduler that further refines these candidate plans by maximizing compute communication overlap; and (iii) a runtime adapter that adaptively composes multiple plans to maximize global efficiency while respecting overall QoEs. Across representative edge deployments, including smart homes, traffic analytics, and small edge clusters, Dora achieves 1.1--6.3 times faster execution and, alternatively, reduces energy consumption by 21--82 percent, all while maintaining QoE under runtime dynamics.

cs.DC

Pion to photon transition form factor: Beyond valence quarks

We investigate the singly virtual transition form factor (TFF) for the $\pi^0\to\gamma^*\gamma$ process in the space-like region using the hard-scattering formalism within the Basis Light-Front Quantization (BLFQ) framework. This form factor is expressed in terms of the perturbatively calculable hard-scattering amplitudes (HSAs) and the light-front wave functions (LFWFs) of the pion. We obtain the pion LFWFs by diagonalizing the light-front QCD Hamiltonian, which is determined for its constituent quark-antiquark and quark-antiquark-gluon Fock sectors with a three-dimensional confinement. We employ the HSAs up to next-to-leading order (NLO) in the quark-antiquark Fock sector and leading order (LO) in the quark-antiquark-gluon Fock sector. The NLO correction to the TFF in the quark-antiquark Fock sector is of the same order as the LO contribution to the TFF in the quark-antiquark-gluon Fock sector. We find that while the quark-antiquark-gluon Fock sector has minimal effect in the large momentum transfer ($Q^2$) region, it has a noteworthy impact in the low-$Q^2$ region. Our results show that, after accounting for both Fock sectors, the TFF within the BLFQ framework aligns well with existing experimental data, particularly in the low $Q^2$ region.

hep-ph